Performance monitoring and performance optimization method and device for network-on-chip

By configuring and monitoring the data stream transmission request frequency parameters in the AI ​​chip bus system, combined with three threshold judgment methods, accurately locate the cause of transmission congestion in the bus system and optimize the performance, the problem of inability to accurately locate the cause of congestion and optimize the performance in the prior art is solved, and the computing power of the AI ​​chip is improved.

CN120186099AActive Publication Date: 2025-06-20WUHAN LINGJIU MICROELECTRONICS CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510649156.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately locate the causes of local transmission congestion in the AI ​​chip bus system, which leads to the inability to effectively optimize the performance of the bus system, affecting the real-time computing capabilities of the AI ​​chip.

Method used

By configuring the data streaming request frequency parameters for master component, slave component and NoC routes, monitoring the real-time data transmission performance parameters, setting three thresholds to determine the cause of transmission congestion in the bus system, and formulating an optimization plan.

Benefits of technology

It has achieved accurate positioning of the causes of transmission congestion of bus systems, and alleviated the transmission performance pressure of bus systems and improved the high-performance computing capabilities of AI chips without reducing the performance of AI chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186099A_ABST
    Figure CN120186099A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of AI chips, and provides a performance monitoring and performance optimizing method and device for a network-on-chip. According to the technical scheme, three threshold values are set for the master component, the slave component and the NoC route, the corresponding transmission performance parameters are obtained by analyzing the real-time data transmission state of the bus system, and then the corresponding transmission performance parameters are subjected to logical judgment and analysis and are compared with the set threshold values in sequence, so that the real-time data transmission performance of the bus system is obtained. According to the method, whether the transmission congestion of the bus system is from a master component, a slave component or an NOC local route can be accurately positioned, meanwhile, a reasonable transmission performance optimization scheme is formulated, and the transmission performance pressure of the bus system is relieved on the premise that the system performance is not reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of AI chips, and particularly relates to a method and device for monitoring and optimizing the performance of a network on chip. Background Art

[0002] In recent years, with the rapid development of semiconductor processes and technologies, the application scenarios related to artificial intelligence (AI) technology have been continuously enriched, and the demand for the high-performance computing capabilities of AI chips has been increasing day by day. A network on chip (NoC) is the core of information interaction in an AI chip, and data interaction between chip components needs to be transmitted and forwarded through the NoC. Therefore, a high-performance NoC bus system is the key to realizing high-bandwidth and low-latency data processing in an AI chip, and can significantly improve the high-performance computing capabilities of the AI chip. Usually, when an AI chip performs high-performance computing, a large number of data processing requests are sent by multiple core components inside it, and the NoC needs to transmit and forward the data request operations. The instantaneous high-density data requests often easily cause excessive bus data transmission pressure in the local routing of the NoC, resulting in an increase in the data transmission delay and a decrease in the bandwidth of the NoC, causing the problem of data transmission congestion on the bus system, and ultimately seriously affecting the real-time computing performance of the AI chip.

[0003] Currently, the mainstream method for monitoring and optimizing the performance of the bus system is to monitor the number of outstanding commands, real-time bandwidth, and latency on the interface between the master component (main device component) in the system and the system bus, and compare the bandwidth or latency with the transmission threshold set by the system. After exceeding the threshold, backpressure processing is performed on the master component, and the data processing pressure of the bus system is reduced by restricting the send request operations of the master component. While this operation reduces the data transmission pressure of the bus system, it also reduces the computing capabilities of the AI chip. This solution does not deeply analyze the problem of local transmission congestion in the bus system and determine the cause of the local transmission congestion problem in the bus system.

[0004] In summary, there is an urgent need for a new network-on-chip performance monitoring and performance optimization solution to monitor the transmission pressure of the bus system, accurately locate whether the transmission congestion in the bus system comes from the system master component, the NoC local routing, or the system slave component (slave device component), and at the same time, formulate a reasonable bus performance optimization solution to relieve the transmission pressure of the bus system without reducing the performance of the AI chip. Summary of the Invention

[0005] In view of the above problems, the object of the present invention is to provide a method and device for performance monitoring and performance optimization of a network-on-chip, aiming to solve the technical problem in the existing technical solution that the bus system of the AI chip cannot accurately locate the local transmission congestion of the bus system.

[0006] The present invention adopts the following technical solutions: On the one hand, the method for performance monitoring and performance optimization of the network-on-chip includes the following steps: Step S1: Configure the data stream operation request sent by the master component to the specified slave component according to the current application scenario of the bus system; Step S2: Configure the priority parameter of the data stream operation request sent by the master component on the NoC route according to the current application scenario of the bus system; Step S3: Configure the data stream transmission request frequency parameters of the master component, slave component and NoC route according to the current application scenario of the bus system, and start the simulation; Step S4: Monitor the real-time data stream on the bus interfaces of the master component and the slave component, and obtain the real-time transmission performance parameters of the master component and the slave component; Step S5: Analyze and judge whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters. If it meets, continue to monitor; Step S6: If the first threshold is not met, judge whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario; Step S7: If the second threshold of the slave component is still not met, determine that one of the reasons for the bus system transmission congestion comes from the slave component, adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then return to Step S4 to continue monitoring; Step S8: If the second threshold of the slave component is met, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameters of the current data request operation; Step S9: Judge whether the path transmission performance parameter meets the third threshold of the current application scenario; Step S10: If the third threshold is not met, determine that one of the reasons for the bus system transmission congestion comes from the NoC route, configure to improve the data transmission priority of the master component in the network-on-chip, and then return to Step S4 to continue monitoring; Step S11: If the third threshold is met, it is determined that one of the reasons for the transmission congestion in the bus system comes from the master component. Configure the data stream transmission request frequency parameters of the master component, slave component, and NoC routing, increase the transmission clock frequency of the data request on the chip bus system, and achieve overclocking transmission of the data stream in the on-chip network.

[0007] On the other hand, the on-chip network performance monitoring and performance optimization device is applied to the bus system. The bus system has NoC routing connected by multiple paths, and several master components and several slave components are mounted on the bus system. The device includes: A parameter configuration module for configuring the data stream operation request sent by the master component to the specified slave component, configuring the priority parameter of the data stream operation request sent by the master component on the NoC routing, and configuring the data stream transmission request frequency parameters of the master component, slave component, and NoC routing, and starting the simulation; A performance monitoring module for monitoring the real-time data stream on the bus interfaces of the master component and slave component, and obtaining the real-time transmission performance parameters of the master component and slave component; An analysis and optimization module for analyzing and determining whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario based on the monitored real-time transmission performance parameters. If it meets, continue monitoring; if it does not meet the first threshold, then determine whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario; if it still does not meet the second threshold of the slave component, it is determined that one of the reasons for the transmission congestion in the bus system comes from the slave component, adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then return to continue monitoring; if it meets the second threshold of the slave component, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameters of the current data request operation; determine whether the path transmission performance parameter meets the third threshold of the current application scenario; if it does not meet the third threshold, it is determined that one of the reasons for the transmission congestion in the bus system comes from the NoC routing, configure to increase the data transmission priority of the master component in the on-chip network, and then return to continue monitoring; if it meets the third threshold, it is determined that one of the reasons for the transmission congestion in the bus system comes from the master component, configure the data stream transmission request frequency parameters of the master component, slave component, and NoC routing, increase the transmission clock frequency of the data request on the chip bus system, and achieve overclocking transmission of the data stream in the on-chip network.

[0008] The beneficial effects of the present invention are as follows: By setting three thresholds for the master component, slave component, and NoC router, and analyzing the real-time data transmission status of the bus system to obtain corresponding transmission performance parameters, and then through logical judgment and analysis, comparing with various set thresholds in sequence, it can accurately locate whether the transmission congestion of the bus system comes from the master component, slave component, or the NOC local router. At the same time, a reasonable transmission performance optimization scheme is formulated to relieve the transmission performance pressure of the bus system without reducing the system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a flowchart of a method for performance monitoring and performance optimization of a network-on-chip provided by an embodiment of the present invention; Figure 2 is a topological diagram of an application scenario example provided by an embodiment of the present invention; Figure 3 is a bus request and response timing diagram; Figure 4 is a data request and response timing diagram; Figure 5 is a structural schematic diagram of a network-on-chip performance monitoring and performance optimization device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0011] In order to illustrate the technical solutions described in the present invention, the following will be illustrated through specific embodiments.

[0012] Embodiment 1: As Figure 1 shown, the method for performance monitoring and performance optimization of a network-on-chip provided in this embodiment includes the following steps: Step S1: Configure the data stream operation request sent by the master component to a specified slave component according to the current application scenario of the bus system.

[0013] Step S2: Configure the priority parameter of the data stream operation request sent by the master component on the NoC router according to the current application scenario of the bus system.

[0014] Step S3: Configure the data stream transmission request frequency parameters of the master component, slave component, and NoC router according to the current application scenario of the bus system, and start the simulation.

[0015] The above steps S1 - S3 are the parameter configuration phase.

[0016] The application scenario described in this embodiment refers to a high - performance application scenario in which the chip bus system conducts performance limit testing, including several master components, slave components, and NoC routers, etc. Each master component is sequentially connected to each slave component through several NoC routers, and a NoC path is formed between the master component and the slave component. By configuring the working parameters of relevant master components, NoC routers, and slave components, the limit performance testing of the chip core components (master components and slave components) and local NoC routers is realized to verify the performance monitoring and performance optimization functions of the chip bus system.

[0017] Such as Figure 2 shown in an example of an application scenario, which shows a specific connection relationship between the core components and local NoC routers.

[0018] In the figure, M0, M1, and M2 are master components, S0, S1, S2, and S3 are slave components, and Router0, Router1, Router2, Router3, Router4, and Router5 are local NoC routers. In the high - performance testing of the chip, the master component will send a large number of data request operations to the slave component through the NoC router. For example, if most of the data request operations are concentrated on S1 and S2, as the number of data request operations of M0, M1, and M2 increases, the path delays of the total line paths P0->P3->P6->P9, P1->P4->P10, P0->P3->P7->P9, and P2->P5 - P10 will gradually increase, and at the same time, the response delays of S1 and S2 will also gradually increase. If the chip does not perform performance optimization of the bus system, with the increase of the chip performance testing time, it is very easy to have a transmission congestion problem in the local bus system, and ultimately lead to a serious decline in the chip performance.

[0019] Therefore, when this embodiment realizes monitoring and optimization, it is first necessary to configure the request relationship between the data operations of the master component and the corresponding target slave component, then configure the priority of the data flow operation, and then configure the transmission request frequency.

[0020] Step S4, monitor the real - time data flow on the bus interfaces of the master component and the slave component, and obtain the real - time transmission performance parameters of the master component and the slave component.

[0021] This step is the performance monitoring stage, which monitors the real-time transmission performance parameters on the bus-related interfaces. The main parameters include real-time bandwidth, the number of outstanding commands, and transmission latency, etc. These can be used as the criteria to measure the speed of data transmission in the on-chip network. The calculation methods of each performance parameter are as follows: The calculation method of the number of outstanding commands Outstanding is as follows: Outstanding is the number of data request sent but not yet received a return command. That is, for each data request sent, Outstanding is incremented by 1, and for each data response received, Outstanding is decremented by 1. When the data request and data response are sent and received simultaneously, Outstanding remains unchanged. Outstanding is dynamically calculated.

[0022] The calculation method of the bus data transmission latency Latency is: Latency = request_interval_time + transaction_response_time. Here, request_interval_time is the request interval time, and transaction_response_time is the request transaction response time. As Figure 3 shown in the bus request and response timing diagram, taking the calculation of the Latency of bus request ID1 as an example. In each data request sent on the AXI bus interface, when calculating the Latency of bus request ID1, the completion time of the ID0 request is used as the start time of the ID1 request interval time. The start time of the ID1 request is used as the end time of the ID1 request interval time and the start time of the ID1 request transaction response time. The response received by the bus for ID1 is used as the end time of the ID1 request transaction response time. Then, the sum of the request interval time and the request transaction response time is calculated as the transmission latency Latency of the ID1 data request transaction.

[0023] The calculation method of the real-time bandwidth Bandwidth is: Bandwidth = transaction_data / Latency. Here, transaction_data is the number of bits of each data request, and transaction_data = (len + 1) * 2 size, len is the length of each data request, and size is the data bit width of each data request. In the bus system, the data request len sent by the master component is usually not a fixed value, which will cause the calculated real-time bandwidth to fluctuate greatly due to the change of the length len, and it is impossible to quantitatively judge whether the bus system meets the performance requirements through the real-time bandwidth parameter. Therefore, in this embodiment, the concept of real-time average bandwidth is added on the basis of the real-time bandwidth. The real-time average bandwidth Average_Bandwidth = Bandwidth / (len + 1), that is, 2 size / Latency, which eliminates the influence of the length len in the data request operation on the real-time bandwidth.

[0024] Step S5: According to the monitored real-time transmission performance parameters, analyze and judge whether the bus transmission performance of the master component in the current data request operation meets the first threshold of the current application scenario. If it meets, continue to monitor.

[0025] In this step, the judgment criterion for whether the first threshold is met is: whether both the transaction latency and the real-time average bandwidth from the master component sending the request to the bus meet the corresponding first thresholds, that is, specifically, these two performance parameters of transaction latency and real-time average bandwidth are used for judgment.

[0026] Here, the first analysis is that the master component sends a data request to the bus and the bus responds. The transaction latency and real-time average bandwidth corresponding to the master component are calculated according to the calculation method shown in step S4, which will not be elaborated here. There are two first thresholds, corresponding to the transaction latency and the real-time average bandwidth respectively. When as long as one of these two monitored performance parameters does not meet the first threshold condition, it is determined that the first threshold of the current application scenario is not met. If both performance parameters meet the corresponding first threshold conditions, it is determined that the first threshold of the current application scenario is met.

[0027] Step S6: If the first threshold is not met, judge whether the bus transmission performance of the slave component in the current data request operation meets the second threshold of the current application scenario.

[0028] If the first threshold is not met, further judge whether the performance parameters of the slave component meet the second threshold of the previous application scenario. Similarly, after the slave component receives the bus response, the slave component returns the content corresponding to the data request to the bus, and it is also calculated according to the calculation method shown in step S4, which will not be elaborated here, and the transaction latency and real-time average bandwidth corresponding to the slave component can also be obtained. Compare these two performance parameters with the two set second thresholds.

[0029] Step S7: If the second threshold of the slave component is still not met, it is determined that one of the reasons for the bus system transmission congestion comes from the slave component. Adjust the operation request of the current master component to the slave component with a bus transmission performance less than the set second threshold, and then return to step S4 to continue monitoring.

[0030] If at least one of the two performance parameters does not meet the second threshold condition, it is determined that the second threshold of the current application scenario is not met. At this time, it can be located that one of the reasons for the bus system transmission congestion comes from the slave component. Then, during optimization, first, in the manner of step S1, adjust the operation request of the current master component to the slave component with a bus transmission performance less than the set second threshold, and then continue to return to step S4 for monitoring.

[0031] Step S8: If the second threshold of the slave component is met, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameter of the current data request operation.

[0032] If the condition of the second threshold is met, further calculate the path transmission performance parameter of the data request operation on the NoC path. As Figure 4 shown in the complete data request and response timing diagram, the data request is sent from the master component and forwarded through the NoC router to the slave component to obtain a data response.

[0033] In the figure, the request data of the master component has an ID number and an address. By using the ID number and the address, the ID number and address of the data response of the slave component bus interface can be matched, so that the path transmission performance parameter of the data response operation of the slave component corresponding to the same data request operation of the master component on the bus system can be obtained. In this embodiment, it includes the path delay NoC_latency and the delay average bandwidth NoCAverage_Bandwidth.

[0034] It can be seen from the figure that the transaction delay of the master component is Tmaster_latency, and the transaction delay of the slave component of the data response is Tslave_latency. The path delays of the data request transaction on the NoC are the NoC request delay and the NoC response delay respectively. Therefore, the path delay NoC_latency of the data request of the master component on the NoC router = Tmaster_latency - Tslave_latency. Referring to the above calculation method of the average real-time bandwidth, the delay average bandwidth NoCAverage_Bandwidth of the data request on the NoC router can be calculated as 2 size / NoC_latency.

[0035] Step S9: Determine whether the path transmission performance parameter meets the third threshold of the current application scenario.

[0036] Similarly, the criterion for determining whether it meets the third threshold is: whether both the path delay and the delay average bandwidth of the request data on the NoC router meet the corresponding third thresholds.

[0037] Step S10: If it does not meet the third threshold, it is determined that one of the reasons for the transmission congestion of the bus system comes from the NoC router. Configure to increase the data transmission priority of the master component in the on-chip network, and then return to step S4 to continue monitoring.

[0038] If the third threshold is not met, that is, at least one of the two performance parameters does not meet the corresponding third threshold, it can be located that one of the reasons for the transmission congestion of the bus system comes from the NoC router. When optimizing, first configure to increase the data transmission priority of the master component on the NoC router in the manner of step S2, and then continue to return to step S4 for monitoring.

[0039] Step S11: If it meets the third threshold, it is determined that one of the reasons for the transmission congestion of the bus system comes from the master component. Configure the data stream transmission request frequency parameters of the master component, slave component, and NoC router, increase the transmission clock frequency of the data request on the chip bus system, and achieve overclocked transmission of the data stream in the on-chip network.

[0040] If the condition of meeting the third threshold is met, it can be located that one of the reasons for the congestion on the bus system comes from the master component. When optimizing, first configure the data stream transmission request frequency parameters of the master component, slave component, and NoC in the manner of step S3, increase the transmission clock frequency of the data request on the bus system, achieve overclocked transmission of the data stream on the NoC, and relieve the transmission pressure of the bus system in high-performance application scenarios.

[0041] The above steps S5 - S11 are the analysis, judgment, and optimization adjustment stages based on the monitoring results. Through the three-level comparison results of the bus real-time transmission bandwidth and the set thresholds, the real-time data transmission status of the bus system can be truly reflected, and it can be accurately located whether the transmission congestion of the bus system comes from the master component, the NOC local router, or the slave component. At the same time, a reasonable transmission performance optimization plan is formulated to relieve the transmission performance pressure of the bus system without reducing the system performance. Additionally, it should be noted that the first threshold, second threshold, and third threshold here are the optimal performance parameter values measured when a single master component sends data requests under the empty load of the bus system, and 60% of the optimal performance parameter values are taken as the corresponding thresholds for performance judgment.

[0042] Embodiment 2: As Figure 5 shown, this embodiment provides a performance monitoring and performance optimization device for a network-on-chip, which is applied to a bus system. The bus system has multiple paths connected to NoC routers, and several master components (illustrated as M0 and M1 in the figure) and several slave components (illustrated as S0 - S3 in the figure) are mounted on the bus system. The device includes: A parameter configuration module 101, which is used to configure the data stream operation requests sent by the master components to the specified slave components according to the current application scenario of the bus system, configure the priority parameters of the data stream operation requests sent by the master components on the NoC routers, and configure the data stream transmission request frequency parameters of the master components, slave components and NoC routers, and start the simulation; A performance monitoring module 102, which is used to monitor the real-time data stream on the bus interfaces of the master components and slave components, and obtain the real-time transmission performance parameters of the master components and slave components; An analysis and optimization module 103, which is used to analyze and judge whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters. If it meets, continue to monitor; if it does not meet the first threshold, then judge whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario; if it still does not meet the second threshold of the slave component, it is determined that one of the reasons for the bus system transmission congestion comes from the slave component, and the operation request of the current master component is adjusted to the slave component whose bus transmission performance is less than the set second threshold, and then return to continue monitoring; if it meets the second threshold of the slave component, calculate the path transmission performance parameters of the data request operation on the NoC path according to the real-time transmission performance parameters of the current data request operation; judge whether the path transmission performance parameters meet the third threshold of the current application scenario; if they do not meet the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the NoC router, configure to increase the data transmission priority of the master component in the network-on-chip, and then return to continue monitoring; if it meets the third threshold, it is determined that one of the reasons for the bus system transmission congestion comes from the master component, configure the data stream transmission request frequency parameters of the master component, slave component and NoC router, increase the transmission clock frequency of the data request on the chip bus system, and realize the overclock transmission of the data stream in the network-on-chip.

[0043] The three functional modules 101-103 of the device of this embodiment correspond to the three stages in the first embodiment. Specifically, module 101 corresponds to steps S1-S3, module 102 corresponds to step S4, and module 103 corresponds to steps S5-S11.

[0044] The parameter configuration module 101 mainly configures the working parameters, and the performance monitoring module 102 monitors and obtains the real-time performance parameters of the master component and the slave component on the NoC bus interface respectively, and then passes the real-time performance parameters of the master component and the slave component to the bus performance analysis and optimization module 103. The analysis and optimization module 103 caches the real-time performance parameters, and matches the real-time performance parameters corresponding to the data request transactions of the master component and the slave component according to the ID number and address information in the data request, obtains the real-time bus performance parameters of the same data request transaction on the bus system, and calculates the real-time bus performance parameters of the data request transaction on the NoC path. The analysis and optimization module 103 then analyzes and judges the bus transmission performance, and judges in turn whether the delay and real-time average bandwidth on the master component, the slave component and the NoC route meet the threshold set by the bus system. For the data request transaction whose performance parameters do not meet the threshold set by the bus system, the cause of bus congestion is analyzed to determine whether the congestion comes from the NoC route, the slave component, or the master component.

[0045] After locating the specific congestion location, the corresponding configuration parameters are generated for the specified functional components or NoC routes through the parameter configuration module 101, and the simulation parameters are reconfigured to achieve the NoC transmission performance optimization operation, and then the aforementioned performance monitoring and performance optimization operations are continued until the transmission congestion problem on the chip bus system is finally solved, so as to alleviate the bus system transmission performance pressure without reducing the system performance, thereby ensuring the high-performance computing capability of the chip.

[0046] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for performance monitoring and performance optimization of a network on chip, characterized in that: The method comprises the following steps: Step S1: According to the current application scenario of the bus system, configure the data flow operation request sent by the master component to the specified slave component; Step S2: According to the current application scenario of the bus system, configure the priority parameters of the data flow operation request issued by the master component on the NoC routing; Step S3: according to the current application scenario of the bus system, configure the data flow transmission request frequency parameters of the master component, the slave component and the NoC routing, and start the simulation; Step S4: monitor the real-time data flow on the bus interface of the master component and the slave component, and obtain the real-time transmission performance parameters of the master component and the slave component; Step S5: Analyze and determine whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters, and continue monitoring if it meets the first threshold; Step S6: If the first threshold is not met, determine whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario; Step S7: If the second threshold of the slave component is still not met, it is determined that one of the causes of the bus system transmission congestion comes from the slave component, and the operation request of the current master component is adjusted to the slave component whose bus transmission performance is less than the set second threshold, and then return to step S4 to continue monitoring; Step S8: If the second threshold of the slave component is met, calculate the path transmission performance parameter of the data request operation on the NoC path according to the real-time transmission performance parameter of the current data request operation; Step S9: determining whether the path transmission performance parameter meets a third threshold of the current application scenario; Step S10: If the third threshold is not met, it is determined that one of the reasons for the bus system transmission congestion is from the NoC routing, and the master component is configured to increase the data transmission priority in the on-chip network, and then the process returns to step S4 to continue monitoring; Step S11: If the third threshold is met, it is determined that one of the reasons for the bus system transmission congestion comes from the master component, and the data stream transmission request frequency parameters of the master component, slave component and NoC routing are configured to increase the transmission clock frequency of the data request on the chip bus system to achieve overclocked transmission of the data stream in the on-chip network.

2. The method for performance monitoring and performance optimization of a network on chip as claimed in claim 1, characterized in that: The application scenario is a high-performance application scenario in which a chip bus system is undergoing performance limit testing. Each master component is connected to each slave component in turn through a number of NoC routes, and a NoC path is formed between the master component and the slave component.

3. The method for performance monitoring and performance optimization of a network on chip as claimed in claim 2, characterized in that: The real-time transmission performance parameters include real-time bandwidth, number of advanced commands and transmission delay, wherein: The calculation method of the number of outstanding commands is as follows: Outstanding is the number of commands that have not received a return among the number of data requests sent, that is, Outstanding increases by 1 for each data request sent, and decreases by 1 for each data response received. If both the data request and the data response sent reach Outstanding at the same time, it remains unchanged, and Outstanding is dynamically counted; The calculation method of transmission delay Latency is: Latency = request_interval_time + transaction_response_time, where request_interval_time is the request interval time and transaction_response_time is the request transaction response time; The calculation method of real-time bandwidth is: Bandwidth = transaction_data / Latency, where transaction_data is the number of bits for each data request, transaction_data = (len+1) * 2 size , len is the length of each data request, and size is the data bit width of each data request.

4. The method for performance monitoring and performance optimization of a network on chip as claimed in claim 3, characterized in that: The real-time transmission performance parameter also includes the real-time average bandwidth, which is Average_Bandwidth=Bandwidth / (len+1), which is 2 size / Latency.

5. The method for performance monitoring and performance optimization of a network on chip as claimed in claim 4, characterized in that: The path transmission performance parameters described in step S8 include path delay and delay average bandwidth, path delay NoC_latency=Tmaster_latency-Tslave_latency, where Tmaster_latency is the master component transaction delay, Tslave_latency is the slave component transaction delay; delay average bandwidth NoCAverage_Bandwidth=2 size / NoC_latency.

6. The method for performance monitoring and performance optimization of a network on chip as claimed in claim 5, characterized in that: In step S5, the criterion for determining whether the first threshold is met is whether the transaction delay and real-time average bandwidth of the master component sending the request to the bus both meet the corresponding first threshold; in step S6, the criterion for determining whether the second threshold is met is whether the transaction delay and real-time average bandwidth of the slave component responding to the request to the bus both meet the corresponding second threshold; in step S9, the criterion for determining whether the third threshold is met is whether the path delay and delay average bandwidth of the data on the NoC routing both meet the corresponding third threshold.

7. The method for performance monitoring and performance optimization of a network on chip as claimed in claim 6, characterized in that: The first threshold, the second threshold and the third threshold are the optimal performance parameter values ​​measured when a single master component sends a data request when the bus system is unloaded, and 60% of the optimal performance parameter value is taken as the corresponding threshold for performance judgment.

8. A performance monitoring and performance optimization device for a network on chip, characterized in that: The device is applied to a bus system, the bus system has NoC routing connected by multiple paths, the bus system is mounted with several master components and several slave components, and the device includes: A parameter configuration module is used to configure the data flow operation request issued by the master component to the specified slave component, configure the priority parameters of the data flow operation request issued by the master component on the NoC routing, and configure the data flow transmission request frequency parameters of the master component, the slave component and the NoC routing according to the current application scenario of the bus system, and start the simulation; Performance monitoring module, used to monitor the real-time data flow on the bus interface of the master component and the slave component, and obtain the real-time transmission performance parameters of the master component and the slave component; The analysis and optimization module is used to analyze and determine whether the bus transmission performance of the master component of the current data request operation meets the first threshold of the current application scenario according to the monitored real-time transmission performance parameters, and continue monitoring if it does; if it does not meet the first threshold, determine whether the bus transmission performance of the slave component of the current data request operation meets the second threshold of the current application scenario; if it still does not meet the second threshold of the slave component, determine that one of the causes of bus system transmission congestion comes from the slave component, adjust the operation request of the current master component to the slave component whose bus transmission performance is less than the set second threshold, and then return to continue monitoring; if the second threshold of the slave component is met, according to the current The real-time transmission performance parameters of the data request operation calculate the path transmission performance parameters of the data request operation on the NoC path; determine whether the path transmission performance parameters meet the third threshold of the current application scenario; if the third threshold is not met, it is determined that one of the causes of the bus system transmission congestion comes from the NoC routing, and the master component is configured to increase the data transmission priority of the on-chip network, and then return to continue monitoring; if the third threshold is met, it is determined that one of the causes of the bus system transmission congestion comes from the master component, and the data flow transmission request frequency parameters of the master component, slave component and NoC routing are configured to increase the transmission clock frequency of the data request on the chip bus system, so as to realize overclocked transmission of the data flow in the on-chip network.

Citation Information

Patent Citations

  • Performance verification method and device of system on chip

    CN115952074A

  • Network congestion control method and device, equipment and storage medium

    CN117014947A

  • AXI bus transmission bandwidth control method based on back pressure type monitoring

    CN117648231A

  • Conversion system from I2C to AXImaster

    CN216956942U

  • Heterogeneous multiprocessor network on chip devices, methods and operating systems for control thereof

    EP1574965A1