A method, device, equipment and medium for chip bottleneck analysis

By analyzing and calculating the files of the system-level chip, using file parsing, performance calculation and bottleneck analysis scripts, the performance bottleneck points of the chip are quickly positioned, solving the efficiency problem of performance optimization in chip development and reducing the risk of chipping.

CN114781293BActive Publication Date: 2025-07-22SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210445336.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-07-22
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

During the chip development process, how to quickly and efficiently locate performance bottlenecks to optimize chip performance, especially to improve chip competitiveness within a limited R&D cycle.

Method used

By analyzing and calculating the system-level chip files using preset file analysis scripts and performance calculation scripts, the device delay information, path delay information and throughput information are determined, and bottleneck analysis scripts are used for bottleneck analysis to generate a visual table to display performance bottleneck points.

Benefits of technology

It realizes rapid analysis of performance indicators of each module of the SOC system during the RTL simulation process, efficiently position performance bottlenecks, reduce the risk of chips, and improves the comprehensiveness and efficiency of chip optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114781293B_ABST
    Figure CN114781293B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and medium for chip bottleneck analysis, which relates to the field of chip development. The method is applied to a system-on-chip and includes: parsing a pre-obtained first file by using a preset file parsing script to determine target information for performance calculation based on the first file, and processing the target information by using a preset performance calculation script to obtain device delay information, path delay information and throughput information, and determining a second file for bottleneck analysis based on the device delay information, the path delay information and the throughput information; inputting the second file into a preset bottleneck analysis script so as to determine performance bottleneck points in the system-on-chip by using the second file through the bottleneck analysis script. In this way, the pre-configured first file can be utilized, and the performance bottleneck points in the system-on-chip can be quickly determined by using various performance parameters of the chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chip development, and particularly to a method, apparatus, device and medium for analyzing chip bottlenecks. Background Art

[0002] With the continuous increase in the design scale and complexity of integrated circuits, there are more and more factors affecting chip performance, and the requirements for chip engineers are also constantly improving.

[0003] In the entire chip development process, the optimization of chip performance can start from the architecture level and the front-end design level. The performance optimization at the architecture level can be evaluated by constructing a system C model + TLM model (i.e., Transaction Level Modeling), and reasonable optimization of the architecture in the early stage of chip development can point the way for subsequent RTL (i.e., Register Transfer Level) design, saving R & D time, but it has high requirements for the accuracy of model establishment; the optimization at the front-end design level requires constructing test cases, analyzing various performance indicators of the designed circuit through simulation, finding specific bottleneck modules, and improving performance by modifying the design. Currently, popular SOC (i.e., System on Chip) performance analysis tools on the market include Vivado, Verdi, etc. Compared with the system C + TLM model, the front-end RTL design can more truly reflect the current system performance status.

[0004] Chip performance is the key point reflecting the competitiveness of SOC chips. How to quickly optimize chip performance within a limited chip R & D cycle requires chip R & D personnel to efficiently locate the performance bottleneck points.

[0005] As can be seen from the above, in the process of chip performance optimization, how to use a more reasonable scheme to achieve rapid optimization of chip performance, and through rapid analysis of chip performance, efficiently locate the performance bottleneck points of the chip is a problem to be solved in this field. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a method, apparatus, device and medium for analyzing chip bottlenecks, which can quickly analyze various performance indicators in the chip during the RTL simulation process, and finally efficiently locate the performance bottleneck points and convert them into visual tables for display. The specific scheme is as follows:

[0007] In a first aspect, the present application discloses a method for analyzing chip bottlenecks, which is applied to a system-on-chip, and includes:

[0008] Parse the pre-obtained first file using a preset file parsing script to determine target information for performance calculation based on the first file, and save the target information into a preset first data structure;

[0009] Process the target information in the first data structure using a preset performance calculation script to obtain device delay information, path delay information, and throughput information of the system-on-chip, and determine a second file for bottleneck analysis based on the device delay information, path delay information, and throughput information;

[0010] Input the second file into a preset bottleneck analysis script so as to determine performance bottleneck points in the system-on-chip by using the second file through the bottleneck analysis script.

[0011] Optionally, before parsing the pre-obtained first file using a preset file parsing script, it further includes:

[0012] Obtain a first file using a preset file acquisition interface; the first file includes the name of the project and the register transfer level version, the path of the log file, the device name of the target device, the working mode of the target device, the bus protocol type, the standard reference value of device delay, the standard reference value of path delay, the standard reference value of throughput, the bus bit width, the clock frequency information, and the time unit of the log;

[0013] Input the first file into the preset file parsing script.

[0014] Optionally, parsing the pre-obtained first file using a preset file parsing script to determine target information for performance calculation based on the first file includes:

[0015] Parse the pre-obtained first file using a preset file parsing script to obtain the path of the log file in the first file, and determine a target log file according to the path of the log file;

[0016] Determine target data corresponding to a target data packet from the target log file, and determine target information for performance calculation from the target data by using a preset character matching method.

[0017] Optionally, before determining target information for delay calculation from the target data by using a preset character matching method, it further includes:

[0018] Create corresponding hash variables for storing data according to different data flow directions; the data flow directions include the data sending direction and the data receiving direction;

[0019] Correspondingly, determining the target information for delay calculation from the target log file includes:

[0020] Classify the target data according to different data flow directions by using a preset character matching method, and store the classified target data in corresponding hash variables;

[0021] Analyze the target data based on the hash variables to determine the target information for delay calculation from the target data.

[0022] Optionally, after processing the target information in the first data structure by using a preset performance calculation script to obtain the device delay information, path delay information, and throughput information of the system-on-chip, it further includes:

[0023] Save the device delay information and path delay information to a preset second data structure and third data structure respectively;

[0024] Correspondingly, determining the second file for bottleneck analysis based on the device delay information, path delay information, and throughput information includes:

[0025] Determine the storage paths of the second data structure and the third data structure, and determine the standard reference values of device delay, path delay, and throughput according to the first file;

[0026] Use the storage paths of the second data structure and the third data structure, the throughput information, the standard reference value of device delay, the standard reference value of path delay, and the standard reference value of throughput to determine the second file for bottleneck analysis.

[0027] Optionally, in the process of determining the performance bottleneck points in the system-on-chip by using the second file through the bottleneck analysis script, it includes:

[0028] Determine the device delay bottleneck points in the system-on-chip by using the first information in the second file through the bottleneck analysis script; the first information is determined based on the storage path of the second data structure and the standard reference value of device delay;

[0029] Determine the path delay bottleneck points in the system-on-chip by using the second information in the second file through the bottleneck analysis script; the second information is determined based on the storage path of the third data structure and the standard reference value of path delay;

[0030] Determine the throughput bottleneck points in the system-on-chip by using the third information in the second file through the bottleneck analysis script; the third information is determined based on the throughput information and the standard reference value of the throughput.

[0031] Optionally, the determining the device delay bottleneck points in the system-on-chip by using the first information in the second file includes:

[0032] Determine the storage path of the second data structure by using the second file through the preset bottleneck analysis script, and determine the second data structure based on the storage path;

[0033] Perform string splitting on each row of data in the second data structure, and determine the first target column and the second target column in each row of data in the second data structure, and then extract the first target data corresponding to the first target column and the second target data corresponding to the second target column; the first target column is the device name column in each row of data, and the second target column is other columns in each row of data except the device name column;

[0034] Convert the data type of the second target data to a preset data type to obtain the corresponding converted data, and form a first list based on the converted data;

[0035] Use the second file to determine the device delay information and the standard reference value of the device delay, and based on the device delay information and the standard reference value of the device delay, use a preset bottleneck calculation function to calculate each parameter in the first list to obtain replacement parameters corresponding to each parameter in the first list, and then replace each parameter with the replacement parameter to form a second list;

[0036] Use a preset data partitioning method to partition the data in the second list to obtain partitioned data, and generate a two-layer nested dictionary with the device name as the key by using the partitioned data and the first target data; in the two-layer nested dictionary, the device name, the data flow direction corresponding to the device name, and the parameters corresponding to the data flow direction are all key-value pairs with mapping relationships;

[0037] Use a preset screening formula to calculate the corresponding parameters in the two-layer nested dictionary. If the preset bottleneck condition is satisfied, retain the key-value pair corresponding to the corresponding device name. If the preset bottleneck condition is not satisfied, delete the key-value pair corresponding to the corresponding device name to generate an updated dictionary;

[0038] Determine the bottleneck device and the bottleneck data flow direction of the bottleneck device from the updated dictionary, and determine the device delay bottleneck points in the system-on-chip according to the bottleneck device and the bottleneck data flow direction of the bottleneck device.

[0039] Optionally, the chip bottleneck analysis method further includes:

[0040] Use a preset table conversion method to convert the device delay bottleneck points, path delay bottleneck points, and throughput bottleneck points respectively to generate corresponding device delay tables, path delay tables, and throughput tables; the device delay tables, path delay tables, and throughput tables all contain the bottleneck devices of each bottleneck point and the bottleneck data flow direction of the bottleneck device;

[0041] Output and display the device delay table, path delay table, and throughput table to a preset interface respectively.

[0042] In a second aspect, the present application discloses a chip bottleneck analysis device, including:

[0043] A file parsing module, configured to parse a pre-obtained first file by using a preset file parsing script, to determine target information for performance calculation based on the first file, and save the target information to a preset first data structure;

[0044] A performance calculation module, configured to process the target information in the first data structure by using a preset performance calculation script, to obtain device delay information, path delay information, and throughput information of the system-on-chip, and determine a second file for bottleneck analysis based on the device delay information, path delay information, and throughput information;

[0045] A bottleneck analysis module, configured to input the second file into a preset bottleneck analysis script, so as to determine the performance bottleneck points in the system-on-chip by using the second file through the bottleneck analysis script.

[0046] In a third aspect, the present application discloses an electronic device, including:

[0047] A memory, configured to store a computer program;

[0048] A processor, configured to execute the computer program to implement the foregoing chip bottleneck analysis method.

[0049] In a fourth aspect, the present application discloses a computer storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the foregoing disclosed chip bottleneck analysis method are implemented.

[0050] This application parses a pre-acquired first file by using a preset file parsing script to determine target information for performance calculation based on the first file, and saves the target information into a preset first data structure. Then, it processes the target information in the first data structure by using a preset performance calculation script to obtain device delay information, path delay information, and throughput information of the system-on-chip, and determines a second file for bottleneck analysis based on the device delay information, path delay information, and throughput information. Next, it inputs the second file into a preset bottleneck analysis script so that the bottleneck analysis script can use the second file to determine the performance bottleneck points in the system-on-chip. In this way, this solution performs a series of parsing and calculations on a pre-configured first file by using a preset file parsing script, performance calculation script, and bottleneck analysis script, can quickly analyze various performance indicators in the chip during the RTL simulation process, and finally efficiently locate the performance bottleneck points and convert them into a visual table for display. This solution can quickly analyze the performance indicators of each module of the entire SOC system during the RTL simulation process, and finally locate the performance bottleneck points and convert them into a visual table for display, making the performance indicators of the SOC chip converge before tape-out and reducing the tape-out risk to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0052] Figure 1 It is a flowchart of a chip bottleneck analysis method provided by this application;

[0053] Figure 2 It is a flowchart of a specific performance calculation method provided by this application;

[0054] Figure 3 It is a flowchart of a specific method for determining device delay bottleneck points provided by this application;

[0055] Figure 4 It is a schematic diagram of the device delay performance analysis process provided by this application;

[0056] Figure 5 It is an overall flowchart of a chip bottleneck analysis method provided by this application;

[0057] Figure 6 It is a schematic diagram of the structure of a chip bottleneck analysis device provided by this application;

[0058] Figure 7 A structural diagram of an electronic device provided for this application. Specific implementation manners

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0060] In the prior art, in the entire chip development process, the optimization of chip performance can start from the architecture level and the front-end design level. Among them, the front-end RTL design can more truly reflect the current system performance status. Popular SOC performance analysis tools on the market include Vivado, Verdi, etc. In this application, a new type of chip bottleneck analysis method for the chip development stage and based on RTL performance analysis is proposed, which can quickly analyze the performance indicators of each module of the entire SOC system and find out the performance bottleneck points in the system.

[0061] An embodiment of the present invention discloses a chip bottleneck analysis method, which is applied to a system-on-chip. Refer to Figure 1 As described, the method includes:

[0062] Step S11: Parse a pre-obtained first file by using a preset file parsing script, so as to determine target information for performing performance calculation based on the first file, and save the target information into a preset first data structure.

[0063] It can be understood that the preset file parsing script in this embodiment can parse the first file, and the first file is pre-obtained. After the preset file parsing script parses the first file, the target information for performing performance calculation can be determined according to the first file. It should be noted that the first file is a file custom-configured by the user according to different application scenarios.

[0064] In this embodiment, before parsing the pre-obtained first file by using the preset file parsing script, it may further include: obtaining the first file by using a preset file acquisition interface; the first file includes the name of the project and the register transfer level version, the path of the log file, the device name of the target device, the working mode of the target device, the type of bus protocol, the standard reference value of device delay, the standard reference value of path delay, the standard reference value of throughput, the bus bit width, the clock frequency information, and the time unit of the log; input the first file into the preset file parsing script.

[0065] It can be understood that the first file described in this embodiment may include: the name of the project and the RTL version (i.e., the register transfer level version), the path of the log file (i.e., the log file), the device name of the target device, the working mode of the target device, the bus protocol type, the standard reference value of the device delay, the standard reference value of the path delay, the standard reference value of the throughput, the bus bit width, the clock frequency information, and the time unit of the log. The working mode of the target device may include the master mode or the slave mode. The bus protocol type includes but is not limited to AXI (i.e., Advanced eXtensible Interface, advanced extensible interface), AHB (i.e., Advanced High Performance Bus, advanced high performance bus), PCIE (i.e., peripheral component interconnect express, high-speed serial computer expansion bus standard), etc. It should be noted that the first file can be customized according to the different needs of users in different application scenarios, and the path of the log file can be customized according to user needs. The file parsing script can be implemented based on the Python tool.

[0066] In this embodiment, the step of parsing a pre-obtained first file by using a preset file parsing script to determine target information for performance calculation may include: parsing the pre-obtained first file by using the preset file parsing script to obtain the log file path in the first file; determining a target log file according to the log file path, and invoking a preset log parsing script to parse the target log file to determine target data corresponding to a target data packet from the target log file, and determining the target information for performance calculation from the target data by using a preset character matching method. It can be understood that the preset file parsing script in this embodiment can parse the first file, and the file parsing script can parse the target log file. In a specific implementation manner, after parsing and obtaining the log file path, the file parsing script will invoke a preset log parsing script and use the log parsing script to perform Transaction parsing on the log file corresponding to the log file path to extract the target information. The specific implementation process is to extract the parameter information required for subsequent calculation of delay information by reading and parsing the Transaction content in the logs of different protocols. The parameter information may include: command type, burst type, burst size, burst length, ID, command start time, command end time, start and end times of the first data, start and end times of the last data, start time of the first resp, start and end times of the last resp.

[0067] It can be understood that the first file in this embodiment stores the log file path. The file parsing script can use the log file path to determine the target log file, and the file parsing script can invoke a preset log parsing script to determine the target data from the target log file by using the log parsing script. In a specific implementation manner, during the process of determining the target log file, if the bus protocol type in the first file is the PCIE type, the file parsing script will, after obtaining the PCIE log path (i.e., the above-mentioned log file path), obtain the log file (i.e., the above-mentioned target log file) through the path, and then determine the characters in the file that contain TLP (Transaction Layer Packet) information (at this time, the TLP information is the target data packet, and the characters containing the TLP information are the target data), and determine the target information according to the characters containing the TLP information.

[0068] It should be noted that the file parsing script in this embodiment is implemented based on the Python tool, and the log parsing script is implemented based on the Perl tool.

[0069] In this embodiment, before determining the target information for delay calculation from the target data by using a preset character matching method, the following steps are further included: creating corresponding hash variables for storing data according to different data flow directions; the data flow directions include a data sending direction and a data receiving direction; correspondingly, determining the target information for delay calculation from the target log file includes: classifying the target data according to different data flow directions by using a preset character matching method, and storing the classified target data into the corresponding hash variables; analyzing the target data based on the hash variables to determine the target information for delay calculation from the target data.

[0070] In this embodiment, after determining the target data corresponding to the target data packet from the target log file, the target information for performance calculation is determined from the target data. Before determining the target information, two hash variables for storing the target data are declared, and the hash variables are used to analyze the target data, where one of the two hash variables is used to store the data in the data sending direction and the other is used to store the data in the data receiving direction. It should be noted that the method for storing the target data includes, but is not limited to, using the method of hash variables. In some specific embodiments, a dictionary can also be used to save the target data.

[0071] In this embodiment, when classifying the target data according to different data flow directions, specifically, the TX / RX (i.e., Transmit / Receive, send / receive) direction, CMD (i.e., command, command prompt) type, CPLD (i.e., Complex Programming Logic Device, complex programmable logic device) type, 3DW / 4DW of the data can also be classified. The start time and end time of the command, request id, tag, length, size, burst, start time and end time of the first data, start time and end time of the last data, start time and end time of the first response, start time and end time of the last response, etc. are matched by using the character matching method. In the specific implementation process, one RC (i.e., Root Complex) device may correspond to multiple EP (i.e., EndPoint) devices. Therefore, a loop structure can be used to process the EP log. The process of declaring hash variables and processing the RC log and EP log is as follows:

[0072] / / rc hash:

[0073] my$rc_tx={};

[0074] my $rc_rx = {};

[0075] ($rc_tx, $rc_rx) = &rc_data_structure_proc($rc_log_file, $rc_tx, $rc_rx);

[0076] / / ep hash:

[0077] my $ep_num = 0;

[0078] foreach my $ep_file (@ep_log_file) {

[0079] ($ep_tx["$ep_num"], $ep_rx["$ep_num"]) = &ep_data_structure_proc($ep_log_file, $ep_tx["$ep_num"], $ep_rx["$ep_num"]);}

[0080] Among them, rc_data_structure_proc and ep_data_structure_proc are the log processing subroutines for RC and EP, rc_log_file is the RC log file, and ep_log_file is the EP log file.

[0081] The initially obtained information is as follows:

[0082] / / Extract information content:

[0083] $hash_ref->{"$cmd_id_num"}->{'begin_cycle'} = 0;

[0084] $hash_ref->{"$cmd_id_num"}->{'end_cycle'} = 0;

[0085] $hash_ref->{"$cmd_id_num"}->{'cmd'} = $cmd;

[0086] $hash_ref->{"$cmd_id_num"}->{'addr'} = addr;

[0087] $hash_ref->{"$cmd_id_num"}->{'request_id'} = $request_id;

[0088] $hash_ref->{"$cmd_id_num"}->{'tag'} = $tag;

[0089] $hash_ref->{"$cmd_id_num"}->{'length'} = length;

[0090] $hash_ref->{"$cmd_id_num"}->{'size'} = size;

[0091] $hash_ref->{"$cmd_id_num"}->{'burst'} = burst;

[0092] Next, the target data will be analyzed based on the hash variable to determine the target information for delay calculation from the target data. During the analysis of the target data, the RC device and the EP device will be matched. For the RC device, each command in the sending direction of the RC device is arranged in chronological order, and each command is traversed. When the command is a write command, all the information in the receiving direction of the collected EP devices is traversed and arranged in chronological order. A mark is set for the successfully matched command, and it will not be matched again during subsequent traversals.

[0093] Among them, when matching the RC with the EP, the following requirements need to be met: the tag information and address information of the command sent by the RC are consistent with those of the EP, the start time of the command sent by the RC is less than the start time of the command received by the EP, and the TLP packet length of the command sent by the RC is the same as the packet length received by the EP. The start time of the command received by the EP is assigned to the end time of the command sent by the RC. Specifically, the code example for end time mapping is as follows:

[0094] / / End time mapping:

[0095] $rc_tx->{$rc_key}->{'end_time'} = $all_ep_rx_db->{$bresp_key}->{'begin_time'};

[0096] $ep_tx[0]->{$rc_key}->{'end_time'} = $all_ep_rx_db->{$bresp_key}->{'begin_time'};

[0097] In this step, after determining the target information, the target information will be saved to a preset first data structure.

[0098] Step S12: Process the target information in the first data structure by using a preset performance calculation script to obtain the device delay information, path delay information, and throughput information of the system-on-chip, and determine a second file for bottleneck analysis based on the device delay information, path delay information, and throughput information.

[0099] It can be understood that in this embodiment, the preset performance calculation script can process the target information, and after the processing, obtain the device delay information, path delay information, and throughput information of the system-on-chip, and the second file is determined by the device delay information, path delay information, and throughput information.

[0100] In this solution, both the device delay information and the path delay information belong to delay information. Delay refers to the time required for a slave device to complete the request reply response after the current master device initiates a read / write request command. The smaller the delay, the faster the system response speed and the better the performance; throughput refers to the amount of data successfully transmitted by a device or port per unit time, and the larger the throughput, the better the performance.

[0101] In this embodiment, the process of using a preset performance calculation script to process the target information to obtain the device delay information, path delay information, and throughput information of the system-on-chip may include: using a preset first performance calculation script to process the target information to obtain the device delay information and path delay information of the system-on-chip; the device delay information includes the maximum value, minimum value, and average value of the device delay; the path delay information includes the maximum value, minimum value, and average value of the path delay; using a preset second performance calculation script to process the target information to obtain the throughput information of the system-on-chip; the throughput information includes the maximum value, minimum value, and average value of the throughput.

[0102] It can be understood that the device delay information, path delay information, and throughput information corresponding to the three parameters of device delay, path delay, and throughput all include their corresponding maximum values, minimum values, and average values. And in this embodiment, the device delay information and path delay information belong to delay information, and these delay information are all processed by a preset first performance calculation script, while the throughput information is processed by a preset second performance calculation script. That is to say, the first performance calculation script is used to process delay information, and the second performance calculation script is used to process throughput information.

[0103] In the specific implementation process of this step, the maximum, minimum, and average values of device latency and path latency can be calculated in both read and write directions using the first performance calculation script, and throughput information can also be calculated using the second performance calculation script that runs synchronously with the first performance calculation script. Among them, the process of calculating the maximum, minimum, and average values of device latency and path latency in both read and write directions using the first performance calculation script is as follows:

[0104] For device latency (i.e., the delay of the device), it is necessary to first distinguish whether the device is a master or a slave. However, whether it is a slave or a master, it is not necessary to distinguish which slave each command is sent to, but only to distinguish read and write operations, and then calculate the maximum, minimum, and average values of its latency. In a specific implementation, if the device is a master, the maximum latency is the maximum value of the difference between the response time corresponding to the command sent by the master and the start time of the command; the minimum latency is the minimum value of the difference between the response time corresponding to the command sent by the master and the start time of the command; the average latency is the difference between all commands sent by the master and the responses received for the corresponding commands, summed and averaged. If the device is a slave, the idea is similar.

[0105] For path latency (i.e., the delay of the path), only the master is analyzed. When analyzing, it is necessary to distinguish which slave each command is sent to and distinguish read and write operations, and then calculate the maximum, minimum, and average values of the latency on the path between each master and slave.

[0106] After calculating the maximum, minimum, and average values of device latency and path latency, corresponding device latency information and path latency information will be generated based on these values. After generating the device latency information and path latency information of the system-on-chip, the second file for bottleneck analysis will be determined in combination with the obtained throughput information.

[0107] Step S13: Input the second file into a preset bottleneck analysis script so that the bottleneck analysis script can use the second file to determine the performance bottleneck points in the system-on-chip.

[0108] It can be understood that in this embodiment, after determining the second file by using the device latency information, path latency information, and throughput information, the second file is input as the input content of a preset bottleneck analysis script. After the bottleneck analysis script obtains the second file, it will use the second file to determine the performance bottleneck points in the system-on-chip. It should be noted that the bottleneck analysis script in this embodiment is implemented based on a python tool.

[0109] In a specific implementation manner, the specific implementation process for determining the performance bottleneck points in the bottleneck analysis script may include: comparing the maximum value and the average value in the latency information with the standard reference value of the latency. When any one of the maximum value or the average value of the latency exceeds the standard reference value, a latency bottleneck point is given; respectively comparing the minimum value and the average value of the throughput with the standard throughput reference value. When the average value or the minimum value of the throughput is lower than the standard reference value, a throughput bottleneck point is given. It can be understood that the performance bottleneck points described in this embodiment include the latency bottleneck points and the throughput bottleneck points, and the latency bottleneck points include device latency bottleneck points and path latency bottleneck points.

[0110] In addition, it should be noted that the first file in this embodiment is configured with a register transfer level version. The chip bottleneck analysis method described in this solution can generate different performance bottleneck analysis results according to different versions of the register transfer level. Further, the bottleneck analysis results of the current version can be combined with the bottleneck analysis results of the historical version and presented in the form of a line chart, which can more intuitively reflect the repair results of the performance bottleneck points in the improvement solution and which version of the RTL has better performance, thereby providing reference guidance for chip design engineers.

[0111] In this embodiment, in the process of determining the performance bottleneck points in the system-on-chip by using the second file through the bottleneck analysis script, it includes: determining the device latency bottleneck points in the system-on-chip by using the first information in the second file through the bottleneck analysis script; the first information is determined based on the storage path of the second data structure and the standard reference value of the device latency; determining the path latency bottleneck points in the system-on-chip by using the second information in the second file through the bottleneck analysis script; the second information is determined based on the storage path of the third data structure and the standard reference value of the path latency; determining the throughput bottleneck points in the system-on-chip by using the third information in the second file through the bottleneck analysis script; the third information is determined based on the throughput information and the standard reference value of the throughput. Specifically, the bottleneck analysis script in this embodiment has three independent modules inside, which are respectively used for performing performance bottleneck analysis on three parameters: path latency, device latency, and throughput.

[0112] In this embodiment, the chip bottleneck analysis method may further include: using a preset table conversion method to convert the device delay bottleneck points, path delay bottleneck points, and throughput bottleneck points respectively to generate corresponding device delay tables, path delay tables, and throughput tables; the device delay tables, path delay tables, and throughput tables all contain the bottleneck devices of each bottleneck point and the bottleneck data flow directions of the bottleneck devices; output and display the device delay tables, path delay tables, and throughput tables to a preset interface respectively. It can be understood that after determining the bottleneck points in the system by using the chip bottleneck analysis method in this embodiment, the bottleneck points can also be displayed in the form of a table. Specifically, the above bottleneck points can be converted into tables in HTML (i.e., Hyper Text Markup Language) format by using a preset table conversion method, so as to output the corresponding device delay tables, path delay tables, and throughput tables to a preset interface for display.

[0113] In this embodiment, by using a preset file parsing script to parse a pre-obtained first file, target information for performance calculation is determined based on the first file, and the target information is saved to a preset first data structure. Then, a preset performance calculation script is used to process the target information in the first data structure to obtain the device delay information, path delay information, and throughput information of the system-level chip, and a second file for bottleneck analysis is determined based on the device delay information, path delay information, and throughput information. Then, the second file is input to a preset bottleneck analysis script, so that the bottleneck analysis script uses the second file to determine the performance bottleneck points in the system-level chip. In this way, this solution uses a preset file parsing script, performance calculation script, and bottleneck analysis script to perform a series of parsing and calculations on a pre-configured first file, and quickly determines the performance bottleneck points in the chip by using three performance parameters of the chip. This solution transmits the required performance standard reference values and log file paths customized by the user according to different scenarios in the form of a configuration file, and can thus quickly analyze the performance indicators of each module of the entire SOC system during the RTL simulation process, and finally locate the performance bottleneck points, making the performance indicators of the SOC chip tend to converge before tape-out, and reducing the tape-out risk to a certain extent.

[0114] Figure 2 It is a flowchart of a specific performance calculation method provided by an embodiment of the present application. Refer to Figure 2 As shown, the method includes:

[0115] Step S21: Process the target information in the first data structure using a preset performance calculation script to obtain the device delay information, path delay information, and throughput information of the system-on-chip.

[0116] Step S22: Save the device delay information and path delay information into a preset second data structure and third data structure respectively.

[0117] In this embodiment, after determining the device delay information and path delay information, the device delay information and path delay information will be saved into a preset second data structure and third data structure respectively. In a specific implementation, the device delay information saved in the second data structure may include: device name (PCIE EP, PCIE RC), data flow direction (TX, RX), write operation or read operation. The path delay information saved in the third data structure may include: path start device name, path end device name, data flow direction (TX, RX), write operation or read operation.

[0118] Step S23: Determine the storage paths of the second data structure and the third data structure, and determine the standard reference values of device delay, path delay, and throughput according to the first file.

[0119] Step S24: Determine a second file for bottleneck analysis using the storage paths of the second data structure and the third data structure, the throughput information, the standard reference value of device delay, the standard reference value of path delay, and the standard reference value of throughput.

[0120] In this step, a second file is generated by using the storage paths of the second data structure and the third data structure, the standard reference value of device delay, the standard reference value of path delay, and the standard reference value of throughput obtained in S23. In a specific implementation, when generating the second file, the standard reference values of device delay, path delay, and throughput in the first file can be directly extracted into the second file, and then the device delay information and path delay information are determined according to the storage paths of the second data structure and the third data structure. Finally, the second file is generated by combining the throughput information generated by the second performance calculation script.

[0121] In this embodiment, after the performance calculation script obtains the target information, it processes the target information to obtain the device delay information, path delay information, and throughput information of the system-on-chip, and then saves the device delay information and path delay information to a preset second data structure and third data structure respectively, and generates a second file based on the storage paths of the second data structure and the third data structure, the standard reference values of the device delay, the standard reference values of the path delay, and the standard reference values of the throughput. In this way, by comprehensively analyzing the chip performance with these three parameters of device delay, path delay, and throughput, and combining with subsequent bottleneck analysis, the performance bottleneck points in the system-on-chip can be determined more efficiently, improving the comprehensiveness and efficiency of chip optimization.

[0122] Figure 3 It is a flowchart of a specific method for determining the device delay bottleneck point provided by an embodiment of the present application.

[0123] See Figure 3 As shown, the method includes:

[0124] Step S31: Use the preset bottleneck analysis script to determine the storage path of the second data structure from the second file, and determine the second data structure based on the storage path.

[0125] In the present application, after the second file is input into the preset bottleneck analysis script, the script internally performs bottleneck analysis on three types of data: device delay, path delay, and throughput. Specifically, the preset bottleneck analysis script internally has three independent modules, which are respectively used to analyze the performance bottlenecks of three parameters: path delay, device delay, and throughput. Further, what this embodiment proposes is the process of analyzing the path delay bottleneck point using the first module inside the preset bottleneck analysis script. In this step, what is completed is the process that the preset bottleneck analysis script determines the storage path of the second data structure from the second file, and obtains the second data structure according to the storage path of the second data structure.

[0126] Step S32: Split each row of data in the second data structure into strings, determine the first target column and the second target column in each row of data in the second data structure, and then extract the first target data corresponding to the first target column and the second target data corresponding to the second target column.

[0127] It should be noted that in this embodiment, the first target column is the device name column in each row of data in the second data structure, and the second target column is other columns in each row of data except the device name column. Therefore, the first target data in this embodiment is the name of each device.

[0128] It should be noted that in some specific embodiments, since only the delay information of a certain device needs to be analyzed when analyzing the device delay, there is only one column for the device name in the second data structure. Therefore, when the first module extracts the device name, it can directly extract the only device name column in the second data structure. When analyzing the path delay, the device names include the starting device and the ending device on the path (for example, the path from EP to EP, the path from RC to EP). Therefore, the second module needs to extract the device name columns where the starting device and the ending device in the path are located. Further, for the convenience of information extraction, the "device name" column in the second data structure can be directly placed in the first column, and then the first module can directly extract the first column in the second data structure to complete the extraction of the device name.

[0129] It should be noted that for the first module, the data flow direction and read / write operations are distinguished in the device delay, so it can be divided into four types: write send, write receive, read send, and read receive. Each type corresponds to its maximum value, minimum value, and average value. Therefore, each device corresponds to 12 delay parameters when analyzing the device delay. For the second module, the path delay does not distinguish between send and receive, and only focuses on whether it is a read operation or a write operation. The read and write operations respectively correspond to their maximum value, minimum value, and average value. Therefore, each path corresponds to 6 delay parameters when analyzing the path delay.

[0130] Step S33: Convert the data type of the second target data into a preset data type to obtain the corresponding converted data, and form a first list based on the converted data.

[0131] It should be noted that the preset data type in this embodiment can be the int type.

[0132] Step S34: Use the second file to determine the device delay information and the standard reference value of the device delay, and based on the device delay information and the standard reference value of the device delay, use a preset bottleneck calculation function to calculate each parameter in the first list to obtain replacement parameters corresponding to each parameter in the first list, and then replace each parameter with the replacement parameter to form a second list.

[0133] In the specific implementation process, the device delay information can be compared with its corresponding standard reference value. If a certain value in the current values corresponding to the device delay information is greater than its corresponding standard reference value, the delay of the current value is calculated according to a preset delay calculation formula to generate a corresponding replacement parameter. If a certain value in the current values corresponding to the device delay information is not greater than its corresponding standard reference value, the corresponding replacement parameter is directly set to 0. After determining the replacement parameter for each parameter, the various parameters are replaced based on the replacement parameter to form a second list. It should be noted that the process of numerical comparison for completing this step using the path delay information in the second module is similar to the above process.

[0134] The relevant code for generating the second list in this embodiment is as follows:

[0135] / / Process the first list through a deviation calculation function to obtain the second list

[0136] for i in range(len(list)):

[0137] list[i]=calcu_device(list[i],golden_value)

[0138] Step S35: Use a preset data partitioning method to partition the data in the second list to obtain partitioned data, and use the partitioned data and the first target data to generate a two-layer nested dictionary with the device name as the key.

[0139] It should be noted that the device name, the data flow direction corresponding to the device name, and the parameters corresponding to the data flow direction in the two-layer nested dictionary are all key-value pairs with a mapping relationship.

[0140] In this embodiment, after determining the second list, a data partitioning method is preset to partition the data in the second list to generate a double-layer nested dictionary with an internal dictionary and an external dictionary. In a specific implementation, the process of generating the internal dictionary may specifically include: dividing the elements in the second list into groups of three from left to right (i.e., the maximum value, the minimum value, and the average value) to form an independent list, and implementing the key-value pair mapping between the list and its corresponding data flow direction. That is to say, the internal dictionary is composed of the data flow direction and the maximum value, the minimum value, and the average value of the device delay corresponding to the data flow direction, and there are multiple independent lists. In the process of generating the external dictionary, it may include: forming the key-value pairs of the outer dictionary according to the device name and the data flow direction corresponding to the device name. Among them, after extracting the first target data of the first target column in step S32, the extracted device name can be directly used as the key to generate the double-layer nested dictionary. It should be noted that the process of generating the dictionary in this step using the path delay information in the second module is similar to the above process.

[0141] It should be noted that this step can also generate a multi-layer nested dictionary according to different scenarios.

[0142] Step S36: Calculate the corresponding parameters in the double-layer nested dictionary using a preset screening formula. If the preset bottleneck condition is satisfied, the key-value pairs corresponding to the corresponding device names are retained; if the preset bottleneck condition is not satisfied, the key-value pairs corresponding to the corresponding device names are deleted to generate an updated dictionary.

[0143] In this embodiment, it is possible to determine whether there is a performance bottleneck by indexing the internal dictionary layer by layer, traversing each group of values inside the dictionary, and using a preset screening formula to calculate whether each group of values satisfies the preset bottleneck condition. If the preset bottleneck condition is satisfied, it means that there is a performance bottleneck in this group of values, and the corresponding key-value pairs are retained; if the preset bottleneck condition is not satisfied, it means that there is no performance bottleneck in this device, and the corresponding key-value pairs are deleted. After performing bottleneck analysis on all the key-value pairs, an updated dictionary containing only performance bottleneck items will be generated.

[0144] The relevant code for generating the updated dictionary in this embodiment is as follows:

[0145]

[0146] Step S37: Determine the bottleneck device and the bottleneck data flow direction of the bottleneck device from the updated dictionary, and determine the device delay bottleneck point in the system-on-chip according to the bottleneck device and the bottleneck data flow direction of the bottleneck device.

[0147] In this embodiment, determining the bottleneck device and the bottleneck data flow direction of the bottleneck device from the updated dictionary, and determining the performance bottleneck point in the system-on-chip according to the bottleneck device and the bottleneck data flow direction of the bottleneck device may include: converting the updated dictionary by using a preset table conversion method to generate a visualization table; determining the bottleneck device and the bottleneck data flow direction of the bottleneck device from the visualization table, and determining the performance bottleneck point in the system-on-chip according to the bottleneck device and the bottleneck data flow direction of the bottleneck device.

[0148] In this embodiment, the preset bottleneck analysis script internally has three independent modules that respectively analyze device latency, path latency, and throughput. In a specific implementation, a dictionary-to-HTML table sub-module can be built inside each module. After generating the updated dictionary, the three modules can respectively use their own dictionary-to-HTML table sub-modules to convert the updated dictionary into corresponding visualization tables. As shown in Table 1, Table 2, and Table 3, they are the visualization tables corresponding to device latency bottleneck, path latency bottleneck, and throughput latency bottleneck respectively. Users can intuitively obtain from the tables which device the performance bottleneck point is located in and its data flow direction.

[0149] Table 1

[0150]

[0151] Table 2

[0152]

[0153] Table 3

[0154]

[0155] In a specific implementation, such as Figure 4The following shows the performance analysis process of the bottleneck analysis script for device latency. First, determine the latency data according to the second file, and then extract the device name and latency values (the maximum, minimum, and average values of the latency) from the latency data. Specifically, perform string splitting on each line of the data structure 2 (i.e., the above-mentioned second data structure), and use the content of the first column of each line as the key value of the dictionary 1 (i.e., the above-mentioned double-layer nested dictionary). Then convert the remaining characters into int-type data and store them as values in the list 1 (i.e., the above-mentioned first list). Then transfer all the elements in the list to the bottleneck calculation function, calculate the latency (i.e., the latency value) corresponding to each parameter using the preset bottleneck calculation formula, and replace the parameters in the list 1 with the latency corresponding to each parameter to generate the list 2 (i.e., the above-mentioned second list). Then perform multi-layer dictionary nesting according to the list 2 and the device name to generate a dictionary 1 containing an internal dictionary and an external dictionary. Then traverse the dictionary 1 and calculate each group of data in the dictionary 1. If the sum of the maximum, minimum, and average values of the corresponding data is 0, it means that there is no latency bottleneck in this group of data. If the sum of the maximum, minimum, and average values of the corresponding data is greater than 0, it means that there is a latency bottleneck in this group of data. Then delete the key-value pairs corresponding to the data groups without latency bottlenecks to generate a new dictionary (i.e., the above-mentioned updated dictionary), and use the dictionary-to-HTML table sub-module inside the bottleneck analysis script to output the new dictionary in tabular form.

[0156] In this application, the second module in the preset bottleneck analysis script is used to analyze the path delay bottleneck points. The process of analyzing the path delay bottleneck points is similar to the method for analyzing device delay bottleneck points proposed in this embodiment, and specifically may include: using the preset bottleneck analysis script to determine the storage path of the third data structure by means of the second file, and determining the third data structure based on the storage path; splitting each row of data in the third data structure into strings, and performing data extraction on the third data structure after string splitting to obtain the third target column and the fourth target column in each row of data, and then determining the third target data corresponding to the third target column and the fourth target data corresponding to the fourth target column; the third target column is the device name column of the starting device and the device name column of the ending device in each row of data, and the fourth target column is other columns in each row of data except the third target column; converting the data type of the fourth target data into a preset data type to obtain corresponding converted data, and forming a third list based on the converted data; using the second file to determine the path delay information and the standard reference value of the path delay, and based on the path delay information and the standard reference value of the path delay, using a preset bottleneck calculation function to calculate each parameter in the third list to obtain replacement parameters corresponding to each parameter in the third list, and then replacing each parameter with the replacement parameter to form a fourth list; using a preset data partitioning method to partition the data in the fourth list to obtain partitioned data, and generating a double-layer nested dictionary with the device name as the key by using the partitioned data and the third target data; in the double-layer nested dictionary, the data transmission direction between device names and the key-value pairs corresponding to the parameters corresponding to the data transmission directions have mapping relationships; using a preset screening formula to calculate the corresponding parameters in the double-layer nested dictionary, if the preset bottleneck condition is met, the key-value pairs corresponding to the corresponding device names are retained, if the preset bottleneck condition is not met, the key-value pairs corresponding to the corresponding device names are deleted to generate an updated dictionary; determining the bottleneck device and the bottleneck data flow direction of the bottleneck device from the updated dictionary, and determining the path delay bottleneck point in the system-on-chip according to the bottleneck device and the bottleneck data flow direction of the bottleneck device. It can be understood that the data transmission direction refers to the transmission from one device to another device. In a specific implementation manner, there may be a situation where the RC device transmits data to the EP device, and in this case, the data transmission direction is from the RC device to the EP device. It should be noted that Figure 4The overall process in FIG. 1 is also applicable to the performance analysis process of path delay. It should be noted that when extracting the device name, the content of the first column of each row in data structure 2 is no longer extracted, but the device name column of the starting device and the device name column of the end device in data structure 3 are extracted. The remaining steps except this step are the same as those in FIG. Figure 4 The process in is similar.

[0157] In this embodiment, by using a preset bottleneck analysis script, and using the storage path of the second file and the second data structure, and the preset standard reference values of each parameter, the device performance bottleneck point in the system-level chip is determined, and the device performance bottleneck point can be displayed in the form of a visual table. In this way, using the method in this embodiment can facilitate chip engineers to quickly locate the performance bottleneck point of the system or device, and implement targeted improvement plans, thereby improving the performance of the chip.

[0158] like Figure 5 This is an overall flow chart of a chip bottleneck analysis method proposed in the present application. After obtaining the top-level configuration file (i.e., the first file mentioned above), the protocol type and log path and other information in the top-level configuration file are parsed, and then the device mode, log path, protocol type and other information of the device are saved in a dictionary or hash format, and a first data structure is generated in data format 1 based on this information, and then a second data structure and a third data structure are generated based on data format 2. Then, a preset bottleneck analysis configuration file (i.e., the second file mentioned above) is used, and the second data structure and the third data structure are used to perform bottleneck analysis on the current system-level chip, wherein the deviation of device delay, the deviation of path delay and the deviation of throughput are calculated accordingly through three independent modules, and the performance bottleneck points obtained by analysis are finally displayed in the form of an HTML table through the respective HTML conversion sub-modules of the modules.

[0159] See also Figure 6 As shown, the embodiment of the present application discloses a chip bottleneck analysis device, which may specifically include:

[0160] A file parsing module 11, configured to parse a pre-acquired first file using a preset file parsing script, so as to determine target information for performance calculation based on the first file, and save the target information into a preset first data structure;

[0161] A performance calculation module 12 is used to process the target information in the first data structure using a preset performance calculation script to obtain device delay information, path delay information and throughput information of the system-level chip, and determine a second file for bottleneck analysis based on the device delay information, path delay information and throughput information;

[0162] A bottleneck analysis module 13 is configured to input the second file into a preset bottleneck analysis script, so as to determine performance bottleneck points in the system-on-chip by using the second file through the bottleneck analysis script.

[0163] In this application, a preset file parsing script is used to parse a pre-acquired first file, so as to determine target information for performance calculation based on the first file, and save the target information into a preset first data structure. Then, a preset performance calculation script is used to process the target information in the first data structure to obtain device delay information, path delay information, and throughput information of the system-on-chip, and a second file for bottleneck analysis is determined based on the device delay information, path delay information, and throughput information. Then, the second file is input into a preset bottleneck analysis script, so as to determine performance bottleneck points in the system-on-chip by using the second file through the bottleneck analysis script. In this way, this solution uses a preset file parsing script, performance calculation script, and bottleneck analysis script to perform a series of parsing and calculations on a pre-configured first file, can quickly analyze various performance indicators in the chip during RTL simulation, and finally efficiently locate the performance bottleneck points and convert them into a visual table for display. This solution can quickly analyze the performance indicators of each module of the entire SOC system during RTL simulation, and finally locate the performance bottleneck points and convert them into a visual table for display, so that the performance indicators of the SOC chip tend to converge before tape-out, and to a certain extent reduce the tape-out risk.

[0164] Furthermore, an embodiment of this application also discloses an electronic device Figure 7 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of this application.

[0165] Figure 7 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a display screen 24, an input / output interface 25, a communication interface 26, and a communication bus 27. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the chip bottleneck analysis method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0166] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 26 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is imposed here.

[0167] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, magnetic disk, optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be transient storage or permanent storage.

[0168] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the chip bottleneck analysis method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0169] Furthermore, this application also discloses a computer-readable storage medium. The computer-readable storage medium mentioned here includes random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, magnetic disks, optical disks, or any other form of storage medium well-known in the technical field. Among them, when the computer program is executed by a processor, it implements the chip bottleneck analysis method disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.

[0170] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section. Professionals can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0171] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0172] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0173] The above has introduced in detail the chip bottleneck analysis method, device, equipment, and storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for analyzing chip bottlenecks, characterized in that, Applied to a system-on-chip, including: Parsing a pre-acquired first file using a preset file parsing script to determine target information for performance calculation based on the first file, and saving the target information to a preset first data structure; Processing the target information in the first data structure using a preset performance calculation script to obtain device delay information, path delay information, and throughput information of the system-on-chip, and determining a second file for bottleneck analysis based on the device delay information, path delay information, and throughput information; Inputting the second file into a preset bottleneck analysis script to determine performance bottleneck points in the system-on-chip by the bottleneck analysis script using the second file; After processing the target information in the first data structure using the preset performance calculation script to obtain device delay information, path delay information, and throughput information of the system-on-chip, it further includes: Respectively saving the device delay information and path delay information to a preset second data structure and third data structure; Correspondingly, determining the second file for bottleneck analysis based on the device delay information, path delay information, and throughput information includes: Determining the storage paths of the second data structure and the third data structure, and determining a standard reference value for device delay, a standard reference value for path delay, and a standard reference value for throughput according to the first file; Determining the second file for bottleneck analysis using the storage paths of the second data structure and the third data structure, the throughput information, the standard reference value for device delay, the standard reference value for path delay, and the standard reference value for throughput; During the process of determining performance bottleneck points in the system-on-chip by the bottleneck analysis script using the second file, it includes: Determining device delay bottleneck points in the system-on-chip by the bottleneck analysis script using first information in the second file; the first information is determined based on the storage path of the second data structure and the standard reference value for device delay; Determining path delay bottleneck points in the system-on-chip by the bottleneck analysis script using second information in the second file; the second information is determined based on the storage path of the third data structure and the standard reference value for path delay; Determining throughput bottleneck points in the system-on-chip by the bottleneck analysis script using third information in the second file; the third information is determined based on the throughput information and the standard reference value for throughput.

2. The chip bottleneck analysis method according to claim 1, characterized in that Before parsing the pre-acquired first file using the preset file parsing script, it further includes: Obtaining a first file using a preset file acquisition interface; the first file includes the name of the project and the register transfer level version, the log file path, the device name of the target device, the working mode of the target device, the bus protocol type, the standard reference value for device delay, the standard reference value for path delay, the standard reference value for throughput, the bus bit width, the clock frequency information, and the time unit of the log. Input the first file into a preset file parsing script.

3. The chip bottleneck analysis method according to claim 2, wherein Parsing the first file obtained in advance using a preset file parsing script to determine target information for performance calculation based on the first file, including: Parsing the first file obtained in advance using a preset file parsing script to obtain the log file path in the first file; Determining a target log file according to the log file path, and invoking a preset log parsing script to parse the target log file to determine target data corresponding to the target data packet from the target log file, and using a preset character matching method to determine target information for performance calculation from the target data.

4. The chip bottleneck analysis method according to claim 3, wherein Before using the preset character matching method to determine target information for latency calculation from the target data, it further includes: Creating corresponding hash variables for storing data according to different data flow directions; the data flow directions include data sending direction and data receiving direction; Correspondingly, determining target information for latency calculation from the target log file includes: Using a preset character matching method to classify the target data according to different data flow directions, and storing the classified target data in the corresponding hash variables; Analyzing the target data based on the hash variables to determine target information for latency calculation from the target data.

5. The chip bottleneck analysis method according to claim 1, characterized in that Using the first information in the second file to determine the device latency bottleneck point in the system-on-chip, including: Determining the storage path of the second data structure using the second file through the preset bottleneck analysis script, and determining the second data structure based on the storage path; Performing string splitting on each row of data in the second data structure, and determining the first target column and the second target column in each row of data in the second data structure, and then extracting the first target data corresponding to the first target column and the second target data corresponding to the second target column; the first target column is the device name column in each row of data, and the second target column is other columns in each row of data except the device name column; Converting the data type of the second target data to a preset data type to obtain corresponding converted data, and forming a first list based on the converted data; Using the second file to determine the device latency information and the standard reference value of the device latency, and based on the device latency information and the standard reference value of the device latency, using a preset bottleneck calculation function to calculate each parameter in the first list to obtain replacement parameters corresponding to each parameter in the first list, and then replacing each parameter with the replacement parameter to form a second list; Divide the data in the second list using a preset data division method to obtain the divided data, and generate a double-layer nested dictionary with the device name as the key using the divided data and the first target data; in the double-layer nested dictionary, the device name, the data flow direction corresponding to the device name, and the parameters corresponding to the data flow direction are all key-value pairs with a mapping relationship; Calculate the corresponding parameters in the double-layer nested dictionary using a preset screening formula. If the preset bottleneck condition is satisfied, retain the key-value pair corresponding to the corresponding device name. If the preset bottleneck condition is not satisfied, delete the key-value pair corresponding to the corresponding device name to generate an updated dictionary; Determine the bottleneck device and the bottleneck data flow direction of the bottleneck device from the updated dictionary, and determine the device delay bottleneck point in the system-on-chip based on the bottleneck device and the bottleneck data flow direction of the bottleneck device.

6. The chip bottleneck analysis method according to claim 1, wherein Further includes: Use a preset table conversion method to convert the device delay bottleneck point, path delay bottleneck point, and throughput bottleneck point respectively to generate a corresponding device delay table, path delay table, and throughput table; the device delay table, path delay table, and throughput table all contain the bottleneck device of each bottleneck point and the bottleneck data flow direction of the bottleneck device; Output and display the device delay table, path delay table, and throughput table to a preset interface respectively.

7. A chip bottleneck analysis device, characterized in that Includes: A file parsing module for parsing a pre-obtained first file using a preset file parsing script to determine target information for performance calculation based on the first file, and saving the target information to a preset first data structure; A performance calculation module for processing the target information in the first data structure using a preset performance calculation script to obtain device delay information, path delay information, and throughput information of the system-on-chip, and determining a second file for bottleneck analysis based on the device delay information, path delay information, and throughput information; A bottleneck analysis module for inputting the second file into a preset bottleneck analysis script, so as to determine the performance bottleneck point in the system-on-chip using the second file through the bottleneck analysis script; After using the preset performance calculation script to process the target information in the first data structure to obtain the device delay information, path delay information, and throughput information of the system-on-chip, further includes: Save the device delay information and path delay information to a preset second data structure and third data structure respectively; Correspondingly, determining the second file for bottleneck analysis based on the device delay information, path delay information, and throughput information includes: Determine the storage paths of the second data structure and the third data structure, and determine the standard reference value of device delay, the standard reference value of path delay, and the standard reference value of throughput according to the first file; Determine a second file for bottleneck analysis by using the storage paths of the second data structure and the third data structure, the throughput information, the standard reference value of device latency, the standard reference value of path latency, and the standard reference value of throughput; In the process of determining the performance bottleneck points in the system-on-chip by using the bottleneck analysis script with the second file, it includes: Determine the device latency bottleneck points in the system-on-chip by using the first information in the second file through the bottleneck analysis script; the first information is determined based on the storage path of the second data structure and the standard reference value of device latency; Determine the path latency bottleneck points in the system-on-chip by using the second information in the second file through the bottleneck analysis script; the second information is determined based on the storage path of the third data structure and the standard reference value of path latency; Determine the throughput bottleneck points in the system-on-chip by using the third information in the second file through the bottleneck analysis script; the third information is determined based on the throughput information and the standard reference value of throughput.

8. An electronic device, characterized in that, It includes a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the chip bottleneck analysis method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, it implements the chip bottleneck analysis method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Information processing method and terminal device

    CN109997154A

  • System-on-chip performance test method, apparatus and device, and readable storage medium

    CN112540902A