Evaluation method for SoC system optimization and storage medium

By building a delay statistics module in the SoC system, counting the delay data of IP and subsystems, and evaluating and optimizing the SoC system architecture, the problems of improper resource allocation and unclear optimization guidance in the existing technology are solved, and the processing speed and efficiency of the SoC system are improved.

CN119940247APending Publication Date: 2025-05-06BEIJING TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411863099.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing SoC system optimization techniques are difficult to determine whether each IP or subsystem has obtained the optimal resource configuration, and lack clear optimization guidance when bandwidth allocation is unreasonable, resulting in the optimization process being time-consuming and inefficient.

Method used

By building a delay statistics module, using hardware description and verification language, it is mounted into the SoC system, and running the SoC system allows each IP to work under extreme load, counting the average delay of read/write data of the IP and subsystem, and evaluating and optimizing the SoC system architecture based on this data.

Benefits of technology

This realizes an intuitive evaluation of the performance of each IP and subsystem in the SoC system, provides clear optimization guidance, avoids local optimal situations, and improves optimization efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940247A_ABST
    Figure CN119940247A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of SoC systems, and particularly discloses an evaluation method for SoC system optimization and a storage medium. According to the invention, a delay statistics module is constructed by hardware description and verification language, the delay statistics module is mounted on the IP and the subsystem in the SoC system, and then all the IPs and all the subsystems in the SOC system are sequentially judged and optimized according to the counted IP read / write data average delay and subsystem read / write data average delay. The method provided by the invention not only can provide quantifiable reference basis for optimization of the SoC system, but also can sequentially judge and optimize the IP and the subsystem, and can avoid the situation that the local IP or the local subsystem is optimal in the SOC system. In addition, the method provided by the invention is simple and easy to use, and can be directly mounted on the IP in an instantiation mode without increasing extra workload and project time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of SoC systems, and in particular relates to an evaluation method and a storage medium for SoC system optimization. Background Art

[0002] SoC system design is the core link of chip design. When designing a SoC system, it is necessary to combine the application scenarios of chip products, comprehensively consider the role and resources required by each IP or subsystem, and make each IP or subsystem more reasonably distributed. When the bus and public storage of the SoC system are determined, the total bandwidth of the system is fixed. If the bandwidth allocation is unreasonable, some IPs in the SoC system will not be able to perform at their maximum performance, and resources may also be wasted due to redundant allocation. Therefore, whether the bandwidth of the SoC system is reasonably allocated directly determines the processing speed and efficiency of the chip.

[0003] However, existing technical solutions for optimizing SoC systems often have limitations. On the one hand, these solutions can only evaluate whether the current resource allocation meets the system requirements, but cannot determine whether each IP or subsystem has obtained the most optimized resource configuration; on the other hand, when unreasonable bandwidth allocation is detected, these solutions usually cannot provide clear optimization guidance, but rely on the experience of engineers to make multiple iterative attempts, which is not only time-consuming but also inefficient. Summary of the invention

[0004] In order to solve at least one of the problems mentioned in the above background technology, the present disclosure proposes an evaluation method for SoC system optimization.

[0005] An evaluation method for SoC system optimization includes the following steps:

[0006] S1, construct a delay statistics module using a hardware description and verification language, and mount the constructed delay statistics module into the SoC system.

[0007] S2, running the SoC system to make each IP in the SoC system work under extreme load.

[0008] S3, performing an IP granularity operation performance evaluation on the SOC system, and optimizing the SoC system architecture when the IP granularity operation performance evaluation result does not meet its operation performance requirements. The specific process is as follows:

[0009] S301, collecting statistics on the average delay of reading data and the average delay of writing data of the IP under extreme load.

[0010] The calculation formula for the average delay of IP data reading under extreme load is:

[0011]

[0012] In the formula, j is the number of statistical times, M is the total number of statistical times, and t u is the start time of data reading of IP number i under extreme load, t v Read is the data reading cutoff time of the corresponding IP number i under extreme load. i It is the average delay of reading data corresponding to IP number i under extreme load.

[0013] Specifically, t u and t v The recording method comprises the steps of:

[0014] First, under extreme load, the IP corresponding to number i reads data from the memory.

[0015] Then, the ARVALID and ARREADY signals on the corresponding IP number i are detected. When it is detected that the ARVALID and ARREADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i reading data under the extreme load. u , correspondingly, record the ID of this data read operation, recorded as ID First .

[0016] Next, detect the RLAST and RVALID signals on the IP number i. When it is detected that the RLAST and RVALID signals on the IP number i are both 1, record the ID of this data read operation, which is recorded as ID Second .

[0017] Finally, determine the ID First and ID Second If they are equal, the time when the RLAST and RVALID signals on IP number i are both 1 is taken as t v .

[0018] The calculation formula for the average delay of IP writing data under extreme load is:

[0019]

[0020] In the formula, j is the number of statistical times, M is the total number of statistical times, and t u is the start time of data reading of IP number i under extreme load, t v Write is the data read cutoff time for the corresponding IP number i under extreme load. i It is the average delay of writing data corresponding to IP number i under extreme load.

[0021] Specifically, t m and t n The recording method comprises the steps of:

[0022] First, under extreme load, IP number i writes data to the memory.

[0023] Then, the AWVALID and AWEADY signals on the IP number i are detected. When it is detected that the AWVALID and AWEADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i writing data under the extreme load. m , correspondingly, record the ID of this write data operation, recorded as ID First .

[0024] Next, detect the BVALID and BREADY signals on the IP number i. When it is detected that the BVALID and BREADY signals on the IP number i are both 1, record the ID of this write data operation, which is recorded as ID Second .

[0025] Finally, determine the ID First and ID Second If they are equal, the time when the BVALID and BREADY signals on the corresponding IP number i are both 1 is taken as t n .

[0026] After the above steps, the performance of each IP in the SoC system when working in different application scenarios can be intuitively reflected based on the calculated average delay in reading data and the average delay in writing data. This can also be used to evaluate whether the current SoC system can meet the bandwidth requirements of each IP, and whether the current SoC system meets the delay requirements of each IP in each application scenario.

[0027] S302: Determine whether the SoC system meets the operating performance requirement based on the statistically obtained average delay for reading data and the average delay for writing data.

[0028] If the operating performance requirement is not met, the SoC system architecture is optimized and the process returns to S301 to S302 .

[0029] S4, performing subsystem granularity operation performance evaluation on the SOC system, and optimizing the SoC system architecture when the subsystem operation performance evaluation result does not meet its operation performance requirements. The specific process is as follows:

[0030] S401 , counting the average latency of reading data and the average latency of writing data of the subsystem.

[0031] Among them, the calculation formula for the average delay of subsystem reading data under extreme load is:

[0032]

[0033]

[0034]

[0035] In the formula, j is the number of statistical times, M is the total number of statistical times, i is the number of IP in the subsystem, N is the number of IP in the subsystem, t u is the start time of data reading of IP number i under extreme load, t v Read is the data reading cutoff time of the corresponding IP number i under extreme load. i is the average delay of reading data corresponding to IP number i under extreme load, All Read Ave is the total delay of reading data of all IPs in the subsystem under extreme load. Read It is the average delay of reading data of all IPs in the subsystem under extreme load.

[0036] Specifically, t u and t v The recording method comprises the steps of:

[0037] First, under extreme load, the IP corresponding to number i reads data from the memory.

[0038] Then, the ARVALID and ARREADY signals on the corresponding IP number i are detected. When it is detected that the ARVALID and ARREADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i reading data under the extreme load. u , correspondingly, record the ID of this data read operation, recorded as ID First .

[0039] Next, detect the RLAST and RVALID signals on the IP number i. When it is detected that the RLAST and RVALID signals on the IP number i are both 1, record the ID of this data read operation, which is recorded as ID Second .

[0040] Finally, determine the ID First and ID Second If they are equal, the time when the RLAST and RVALID signals on IP number i are both 1 is taken as t v .

[0041] The calculation formula for the average delay of subsystem writing data under extreme load is:

[0042]

[0043]

[0044]

[0045] In the formula, j is the number of statistical times, M is the total number of statistical times, i is the number of IP in the subsystem, N is the number of IP in the subsystem, t m is the start time of writing data to IP number i under extreme load, t n Write is the data writing deadline for the corresponding IP number i under extreme load. i is the average delay of writing data corresponding to IP number i under extreme load, All Write Ave is the total delay of writing data for all IPs in the subsystem under extreme load. Write It is the average delay of writing data for all IPs in the subsystem under extreme load.

[0046] Specifically, t m and t n The recording method comprises the steps of:

[0047] First, under extreme load, IP number i writes data to the memory.

[0048] Then, the AWVALID and AWEADY signals on the IP number i are detected. When it is detected that the AWVALID and AWEADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i writing data under the extreme load. m , correspondingly, record the ID of this write data operation, recorded as ID First .

[0049] Next, detect the BVALID and BREADY signals on the IP number i. When it is detected that the BVALID and BREADY signals on the IP number i are both 1, record the ID of this write data operation, which is recorded as ID Second .

[0050] Finally, determine the ID First and ID Second If they are equal, the time when the BVALID and BREADY signals on the corresponding IP number i are both 1 is taken as t n .

[0051] After the above steps, the performance of each subsystem in the SoC system when working in different application scenarios can be intuitively reflected based on the calculated average delay in reading data and the average delay in writing data, and this can be used to evaluate whether the current SoC system can meet the bandwidth requirements of each subsystem, and whether the current SoC system meets the delay requirements of each subsystem in each application scenario.

[0052] S402, judging whether the SoC system meets the operation performance requirement according to the statistically obtained average delay of reading data and the average delay of writing data;

[0053] If the operating performance requirement is not met, the SoC system architecture is optimized and the process returns to S401 to S402 .

[0054] A computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the evaluation method for SoC system optimization described in any of the above embodiments.

[0055] The present disclosure proposes to construct a delay statistics module by hardware description and verification language, and the delay statistics module is mounted on the IP and subsystem in the SoC system, and then all IPs and all subsystems in the SOC system are judged and optimized in turn according to the statistical average delay of IP read / write data and the average delay of subsystem read / write data. Among them, the average delay of single IP read / write data can intuitively represent the operating performance of each IP in the SoC system; the average delay of single subsystem read / write data can intuitively reflect the bandwidth pressure of each subsystem. The method proposed in the present disclosure can not only provide a quantifiable reference basis for the optimization of the SoC system, but also judge and optimize the IP and subsystem method in turn, which can avoid the situation where the SOC system has a local IP optimum or a local subsystem optimum. In addition, the method proposed in the present disclosure is simple and easy to use, and can be directly mounted on the IP in an instantiated manner, without adding additional workload and project time. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a system flow chart of the present disclosure;

[0057] Figure 2 It is a schematic diagram of the architecture of the SoC system in the embodiment of the present disclosure. DETAILED DESCRIPTION

[0058] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments.

[0059] The terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances. This is just a way of distinguishing objects with the same attributes when describing the embodiments of the present disclosure.

[0060] Embodiment 1:

[0061] An evaluation method for SoC system optimization, the implementation steps are as follows:

[0062] Step S1, constructing a delay statistics module by using a hardware description and verification language, and mounting the constructed delay statistics module into the SoC system.

[0063] Optionally, the SOC system selects AXI bus as the on-chip interconnection protocol; the hardware description and verification language is SystemVerilog, and the module is mounted on the IP in the SoC system through an instantiation method.

[0064] The purpose of this step is to construct a mixed scenario case for the SoC system.

[0065] S2, running the SoC system to make each IP in the SoC system work under extreme load.

[0066] S3, performing an IP granularity operation performance evaluation on the SOC system, and optimizing the SoC system architecture when the IP granularity operation performance evaluation result does not meet its operation performance requirements. The specific process is as follows:

[0067] S301, collecting statistics on the average delay of reading data and the average delay of writing data of the IP under extreme load.

[0068] In the specific implementation process, if the average delay of reading data of all IPs in the statistical subsystem is considered, the extreme load refers to the state where all IPs in the SOC system reach the maximum bandwidth usage.

[0069] The calculation formula for the average delay of IP data reading under extreme load is:

[0070]

[0071] In the formula, j is the number of statistical times, M is the total number of statistical times, and t u is the start time of data reading of IP number i under extreme load, t v Read is the data reading cutoff time of the corresponding IP number i under extreme load. i It is the average delay of reading data corresponding to IP number i under extreme load.

[0072] Specifically, t u and t v The recording method comprises the steps of:

[0073] First, under extreme load, the IP corresponding to number i reads data from the memory.

[0074] Then, the ARVALID and ARREADY signals on the corresponding IP number i are detected. When it is detected that the ARVALID and ARREADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i reading data under the extreme load. u , correspondingly, record the ID of this data read operation, recorded as ID First .

[0075] Next, detect the RLAST and RVALID signals on the IP number i. When it is detected that the RLAST and RVALID signals on the IP number i are both 1, record the ID of this data read operation, which is recorded as ID Second .

[0076] Finally, determine the ID First and ID Second If they are equal, the time when the RLAST and RVALID signals on IP number i are both 1 is taken as t v .

[0077] The calculation formula for the average delay of IP writing data under extreme load is:

[0078]

[0079] In the formula, j is the number of statistical times, M is the total number of statistical times, and t u is the start time of data reading of IP number i under extreme load, t v Write is the data read cutoff time for the corresponding IP number i under extreme load. i It is the average delay of writing data corresponding to IP number i under extreme load.

[0080] Specifically, t m and t n The recording method comprises the steps of:

[0081] First, under extreme load, IP number i writes data to the memory.

[0082] Then, the AWVALID and AWEADY signals on the IP number i are detected. When it is detected that the AWVALID and AWEADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i writing data under the extreme load. m , correspondingly, record the ID of this write data operation, recorded as ID First .

[0083] Next, detect the BVALID and BREADY signals on the IP number i. When it is detected that the BVALID and BREADY signals on the IP number i are both 1, record the ID of this write data operation, which is recorded as IDSecond .

[0084] Finally, determine the ID First and ID Second If they are equal, the time when the BVALID and BREADY signals on the corresponding IP number i are both 1 is taken as t n .

[0085] Optionally, an IP is made to work at extreme load in different application scenarios to obtain the read data delay and write data delay of the IP in different application scenarios, and then the average read data delay and write data delay of the IP are statistically obtained.

[0086] After the above steps, the performance of each IP in the SoC system when working in different application scenarios can be intuitively reflected based on the calculated average delay in reading data and the average delay in writing data. This can also be used to evaluate whether the current SoC system can meet the bandwidth requirements of each IP, and whether the current SoC system meets the delay requirements of each IP in each application scenario.

[0087] S302: Determine whether the SoC system meets the operating performance requirement based on the statistically obtained average delay for reading data and the average delay for writing data.

[0088] In an optional implementation, each IP is operated under extreme load in turn, and the operating performance data of each IP under extreme load is counted; the above step S3 is executed based on the operating performance data of each IP under extreme load; and the operating performance data of each subsystem is obtained based on the motion performance data of the IP of each subsystem under extreme load, and the above step S4 is executed based on the operating performance data of each subsystem.

[0089] In another optional implementation, each IP is operated under extreme load in turn; based on a single IP, the motion performance of the SOC system is evaluated and optimized at the IP granularity; after completing the motion performance evaluation and optimization at the IP granularity, the motion performance of the SOC system is evaluated and optimized at the IP granularity based on a single subsystem until the motion performance evaluation and optimization at the IP granularity is completed.

[0090] As an example and not limitation, the IP granularity operating performance includes the average read latency of at least one IP and / or the average write latency of at least one IP. It should be noted that other indicators can also be selected as the IP operating performance according to actual application scenarios or application requirements, and this disclosure does not limit this.

[0091] In an optional implementation, taking the operating performance of IP granularity including the average delay of reading data of multiple IPs and the average delay of writing data of multiple IPs as an example, it is determined whether the current SOC system architecture (such as the subsystem division method of multiple IPs) is adapted to the operating bandwidth demand distribution of multiple IPs reflected by the average delay of reading data of multiple IPs and the average delay of writing data. If it is adapted, it is satisfied, otherwise, it is not satisfied and needs to be optimized. The adaptation standard is determined according to the actual situation or actual needs, and the present disclosure does not limit this. In an optional implementation, taking the operating performance of IP granularity including the average delay of reading data of a single IP and the average delay of writing data of a single IP as an example, the average delay of reading data and the average delay of writing data of a single IP all meet the operating performance requirements of the IP, then the operating performance based on the IP granularity of the IP meets the operating performance requirements, otherwise it is not met.

[0092] Among them, the operating performance requirements may include but are not limited to at least one of the following: the operating bandwidth requirement distribution of multiple IPs reflected by the read data delay and / or write data delay of multiple IPs, the operating bandwidth requirement of the IP, the read data delay requirement of the IP, the write data delay requirement of the IP, the read data delay requirement of the subsystem where the IP is located, and the write data delay requirement of the subsystem where the IP is located.

[0093] There are many ways to determine whether the operating performance of the IP granularity meets its operating bandwidth requirements. For example, the mapping relationship between the operating performance of the IP (such as the read data delay and write data delay of the IP) and the operating bandwidth requirements can be determined in advance, and whether the operating performance of the IP meets its operating bandwidth requirements can be determined based on the mapping relationship.

[0094] There are many ways to determine whether the operating performance of an IP meets its read data delay requirement (or write data delay requirement, or the read data delay requirement of the subsystem in which it is located, or the write data delay requirement of the subsystem in which it is located). For example, it can be predetermined whether the read data delay (or write data delay) of the IP exceeds the preset read data delay threshold (or write data delay threshold). If it exceeds, it is not satisfied. The present disclosure does not limit the specific setting method of the read data delay threshold and the write data delay threshold. Taking the read data delay requirement of the IP as an example, the preset read data delay threshold can be the read data delay requirement of the IP, or the product of the read data delay requirement of the IP and a preset weight (for example, 0.9). Taking the read data delay requirement of the subsystem in which the IP is located as an example, the preset read data delay threshold can be the read data delay requirement of the subsystem in which the IP is located, or the product of the read data delay requirement of the subsystem in which the IP is located and a preset weight (for example, 0.9).

[0095] If the operating performance requirement is not met, the SoC system architecture is optimized and the process returns to S301 to S302 .

[0096] Optionally, different SoC system architecture optimization strategies may be used for different operation performance requirements. For example, if the operation performance does not meet the operation bandwidth distribution requirements of multiple IPs, then the optimization is performed according to the SoC system architecture optimization strategy corresponding to the operation bandwidth distribution requirements; if the operation performance does not meet the operation bandwidth requirements, then the optimization is performed according to the SoC system architecture optimization strategy corresponding to the operation bandwidth requirements; if the operation performance does not meet the corresponding delay requirements, then the optimization is performed according to the SoC system architecture optimization strategy corresponding to the delay requirements.

[0097] Among them, different delay requirements can have the same optimization strategy, or they can correspond to different optimization strategies respectively, which are determined according to actual needs and actual conditions, and the present disclosure does not limit this.

[0098] S4, performing subsystem granularity operation performance evaluation on the SOC system, and optimizing the SoC system architecture when the subsystem operation performance evaluation result does not meet its operation performance requirements. The specific process is as follows:

[0099] S401 , counting the average latency of reading data and the average latency of writing data of the subsystem.

[0100] In the specific implementation process, if the average delay of reading data of all IPs in the statistical subsystem is considered, the extreme load refers to the state where all IPs in the SOC system reach the maximum bandwidth usage.

[0101] Among them, the calculation formula for the average delay of subsystem reading data under extreme load is:

[0102]

[0103]

[0104]

[0105] In the formula, j is the number of statistical times, M is the total number of statistical times, i is the number of IP in the subsystem, N is the number of IP in the subsystem, t u is the start time of data reading of IP number i under extreme load, t v Read is the data reading cutoff time of the corresponding IP number i under extreme load. i is the average delay of reading data corresponding to IP number i under extreme load, All Read Ave is the total delay of reading data of all IPs in the subsystem under extreme load. Read It is the average delay of reading data of all IPs in the subsystem under extreme load.

[0106] Specifically, t u and t v The recording method comprises the steps of:

[0107] First, under extreme load, the IP corresponding to number i reads data from the memory.

[0108] Then, the ARVALID and ARREADY signals on the corresponding IP number i are detected. When it is detected that the ARVALID and ARREADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i reading data under the extreme load. u , correspondingly, record the ID of this data read operation, recorded as ID First .

[0109] Next, detect the RLAST and RVALID signals on the IP number i. When it is detected that the RLAST and RVALID signals on the IP number i are both 1, record the ID of this data read operation, which is recorded as ID Second .

[0110] Finally, determine the ID First and ID Second If they are equal, the time when the RLAST and RVALID signals on IP number i are both 1 is taken as t v .

[0111] The calculation formula for the average delay of subsystem writing data under extreme load is:

[0112]

[0113]

[0114]

[0115] In the formula, j is the number of statistical times, M is the total number of statistical times, i is the number of IP in the subsystem, N is the number of IP in the subsystem, t m is the start time of writing data to IP number i under extreme load, t n Write is the data writing deadline for the corresponding IP number i under extreme load. i is the average delay of writing data corresponding to IP number i under extreme load, All Write Ave is the total delay of writing data for all IPs in the subsystem under extreme load. Write It is the average delay of writing data for all IPs in the subsystem under extreme load.

[0116] Specifically, t m and t n The recording method comprises the steps of:

[0117] First, under extreme load, IP number i writes data to the memory.

[0118] Then, the AWVALID and AWEADY signals on the IP number i are detected. When it is detected that the AWVALID and AWEADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i writing data under the extreme load. m , correspondingly, record the ID of this write data operation, recorded as ID First .

[0119] Next, detect the BVALID and BREADY signals on the IP number i. When it is detected that the BVALID and BREADY signals on the IP number i are both 1, record the ID of this write data operation, which is recorded as ID Second .

[0120] Finally, determine the ID First and ID Second If they are equal, the time when the BVALID and BREADY signals on the corresponding IP number i are both 1 is taken as t n .

[0121] Optionally, a subsystem is made to work at a limit load in different application scenarios to obtain the read data delay and write data delay of the subsystem in different application scenarios, and then the average read data delay and the average write data delay of the subsystem are statistically obtained.

[0122] After the above steps, the performance of each subsystem in the SoC system when working in different application scenarios can be intuitively reflected based on the calculated average delay in reading data and the average delay in writing data, and this can be used to evaluate whether the current SoC system can meet the bandwidth requirements of each subsystem, and whether the current SoC system meets the delay requirements of each subsystem in each application scenario.

[0123] S402, judging whether the SoC system meets the operation performance requirement according to the statistically obtained average delay of reading data and the average delay of writing data;

[0124] As an example and not limitation, the subsystem granularity operating performance includes the average read latency of at least one subsystem and / or the average write latency of at least one subsystem. It should be noted that other indicators can also be selected as the subsystem operating performance according to actual application scenarios or application requirements, and this disclosure does not limit this.

[0125] In an optional implementation, taking the subsystem granularity operation performance including the average delay of reading data of multiple subsystems and the average delay of writing data of multiple subsystems as an example, it is determined whether the current delay of reading data and / or delay of writing data of multiple subsystems meets the operation bandwidth distribution requirements of multiple subsystems. If so, it is achieved; otherwise, it is not achieved and needs to be optimized. The satisfaction standard is determined according to the actual situation or actual needs, and the present disclosure does not limit this.

[0126] In an optional implementation, taking the subsystem-level operating performance including the average delay in reading data of a single subsystem and the average delay in writing data of a single subsystem as an example, if the average delay in reading data and the average delay in writing data of a single subsystem all meet the operating performance requirements of the subsystem, then the operating performance at the subsystem-level based on the subsystem meets the operating performance requirements, otherwise it does not meet the requirements.

[0127] As an example and not limitation, the operating performance of the subsystem includes at least one of the average read latency of at least one subsystem and / or the average write latency of at least one subsystem. It should be noted that other indicators can also be selected as the operating performance of the subsystem according to actual application scenarios or application requirements, and this disclosure does not limit this.

[0128] Among them, the operating performance requirements of the subsystem may include but are not limited to at least one of the following: the read data delay and / or write data delay of multiple subsystems meets the operating bandwidth distribution requirements of multiple subsystems, the read data delay requirements of the subsystem, and the write data delay requirements of the subsystem.

[0129] There are many ways to determine whether the operating performance of the subsystem granularity meets the read data delay requirement (or write data delay requirement). For example, it can be predetermined whether the read data delay (or write data delay) of the subsystem exceeds the preset read data delay threshold (or write data delay threshold). If it exceeds, it is not satisfied. The present disclosure does not limit the specific setting method of the read data delay threshold and the write data delay threshold. Taking the read data delay requirement of the subsystem as an example, the preset read data delay threshold can be the read data delay requirement of the subsystem, or the product of the read data delay requirement of the subsystem and a preset weight (for example, 0.9).

[0130] If the operating performance requirement is not met, the SoC system architecture is optimized and the process returns to S401 to S402 .

[0131] Optionally, different SoC system architecture optimization strategies can be used for different operating performance requirements. For example, if the operating performance of the subsystem granularity does not meet its operating bandwidth distribution requirements, then the optimization is performed according to the SoC system architecture optimization strategy corresponding to the operating bandwidth distribution requirements; if the operating performance of the subsystem does not meet the corresponding delay requirements, then the optimization is performed according to the SoC system architecture optimization strategy corresponding to the delay requirements.

[0132] Among them, different delay requirements can have the same optimization strategy, or they can correspond to different optimization strategies respectively, which are determined according to actual needs and actual conditions, and the present disclosure does not limit this.

[0133] Embodiment 2:

[0134] like Figure 2 As shown, in the SoC system of this embodiment, IP1, IP2 and IP3 are in the subsystem SUBSYS-A, and IP4, IP5, IP6 and IP7 are in the subsystem SUBSYS-B.

[0135] First, control each IP in the SoC system to work under extreme load, and calculate the average delay of reading and writing data of IP under extreme load, as well as the average delay of reading and writing data of each subsystem.

[0136] According to the obtained average delay of reading data of multiple IPs and the average delay of writing data of multiple IPs, it is judged whether the current SOC architecture is compatible with the bandwidth demand distribution of multiple IPs. In an optional implementation, it is assumed that the delay results of each IP are analyzed and it is found that the bandwidth demand of IP1 and IP2 is large, while the bandwidth demand of IP4, IP5, and IP6 is small. At this time, the SoC system in this embodiment can be considered to be optimized as follows: Optimization 1, adjust the distribution of IPs in each subsystem until the bandwidth demand of IP1 and IP2 is met; Optimization 2, if the NOC bus is used, consider adjusting the priority of the corresponding subsystem SUBSYS-A port on the NOC bus. After the optimization is completed, run again and count the average delay of reading data and the average delay of writing data until the SoC system optimization of IP granularity is successful.

[0137] Then, based on the obtained average delays for reading data of multiple subsystems and the average delays for writing data of multiple IPs, determine whether the average delays for reading data and writing data of multiple subsystems meet the bandwidth distribution requirements of multiple subsystems. Assume that the average delay of subsystem SUBSYS-A is 1000, and the average delay of subsystem SUBSYS-B is 10000. In an optional implementation, if the design requirement of the SoC system in this embodiment is to evenly distribute the bandwidth of the two subsystems, it is obvious that the SoC system in the embodiment does not meet the optimization requirements, then some IPs in subsystem SUBSYS-B can be divided into subsystem SUBSYS-A. After the optimization is completed, run again and count the average delay for reading data and the average delay for writing data until the SoC system optimization at the subsystem granularity is successful.

[0138] Example 3

[0139] This embodiment provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the evaluation method for SoC system optimization described in any of the above embodiments.

[0140] It should be noted that the serial numbers of the embodiments of the present disclosure are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0141] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present disclosure.

[0142] The above are only preferred embodiments of the present disclosure, and are not intended to limit the patent scope of the present disclosure. Any equivalent structure or equivalent process transformation made using the contents of the present disclosure and the drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present disclosure.

Claims

1. An evaluation method for SoC system optimization, characterized in that: Includes steps: S1, construct a delay statistics module using a hardware description and verification language, and mount the constructed delay statistics module into the SoC system; S2, running the SoC system so that each IP in the SoC system works under extreme load; S3, performing an IP-granularity operation performance evaluation on the SOC system, and optimizing the SoC system architecture when the IP-granularity operation performance evaluation result does not meet its operation performance requirements; S4, performing a subsystem-level operational performance evaluation on the SOC system, and optimizing the SoC system architecture when the subsystem operational performance evaluation result does not meet its operational performance requirements.

2. The evaluation method for SoC system optimization according to claim 1, characterized in that: S3 includes the following steps: S301, collecting statistics on the average latency of reading and writing data of IP under extreme load; S302, judging whether the SoC system meets the operation performance requirement according to the statistically obtained average delay of reading data and the average delay of writing data; If the operating performance requirement is not met, the SoC system architecture is optimized and the process returns to S301 to S302 .

3. The evaluation method for SoC system optimization according to claim 1, characterized in that: S4 includes the steps: S401, counting the average delay of reading data and the average delay of writing data of the subsystem; S402, judging whether the SoC system meets the operation performance requirement according to the statistically obtained average delay of reading data and the average delay of writing data; If the operating performance requirement is not met, the SoC system architecture is optimized and the process returns to S401 to S402 .

4. The evaluation method for SoC system optimization according to claim 2, characterized in that: The calculation formula for the average delay of IP reading data under the extreme load is: In the formula, j is the number of statistical times, M is the total number of statistical times, and t u is the start time of data reading of IP number i under extreme load, t v Read is the data reading cutoff time of the corresponding IP number i under extreme load. i is the average delay of reading data corresponding to IP number i under extreme load; The calculation formula for the average delay of IP writing data under the extreme load is: In the formula, j is the number of statistical times, M is the total number of statistical times, and t u is the start time of data reading of IP number i under extreme load, t v Write is the data read cutoff time for the corresponding IP number i under extreme load. i It is the average delay of writing data corresponding to IP number i under extreme load.

5. The evaluation method for SoC system optimization according to claim 3, characterized in that: The calculation formula for the average delay of subsystem reading data under the extreme load is: In the formula, j is the number of statistical times, M is the total number of statistical times, i is the number of IP in the subsystem, N is the number of IP in the subsystem, t u is the start time of data reading of IP number i under extreme load, t v Read is the data reading cutoff time of the corresponding IP number i under extreme load. i is the average delay of reading data corresponding to IP number i under extreme load, All Read Ave is the total delay of reading data of all IPs in the subsystem under extreme load. Read It is the average delay of reading data of all IPs in the subsystem under extreme load; The calculation formula for the average delay of subsystem writing data under the extreme load is: In the formula, j is the number of statistical times, M is the total number of statistical times, i is the number of IP in the subsystem, N is the number of IP in the subsystem, t m is the start time of writing data to IP number i under extreme load, t n Write is the data writing deadline for the corresponding IP number i under extreme load. i is the average delay of writing data corresponding to IP number i under extreme load, All Write Ave is the total delay of writing data for all IPs in the subsystem under extreme load. Write It is the average delay of writing data for all IPs in the subsystem under extreme load.

6. The evaluation method for SoC system optimization according to claim 4 or 5, characterized in that: The u and t v The recording method comprises the steps of: A1, under extreme load, the corresponding IP number i reads data from the memory; A2, detect the ARVALID and ARREADY signals on the corresponding IP number i. When it is detected that the ARVALID and ARREADY signals on the IP number i are both 1, the time is used as the start time t of the IP number i reading data under the extreme load. u , correspondingly, record the ID of this data read operation, recorded as ID First ; A3, detect the RLAST and RVALID signals on the IP number i. When it is detected that the RLAST and RVALID signals on the IP number i are both 1, record the ID of this data read operation, which is recorded as ID Second ; A4, determine ID First and ID Second If they are equal, the time when the RLAST and RVALID signals on IP number i are both 1 is taken as t v .

7. The evaluation method for SoC system optimization according to claim 4 or 5, characterized in that: The m and t n The recording method comprises the steps of: B1, under extreme load, data is written to the memory by IP number i; B2, detect the AWVALID and AWEADY signals on the IP number i. When it is detected that the AWVALID and AWEADY signals on the IP number i are both 1, the time is taken as the start time t of the IP number i writing data under the extreme load. m , correspondingly, record the ID of this write data operation, recorded as ID First ; B3, detect the BVALID and BREADY signals on the IP number i. When it is detected that the BVALID and BREADY signals on the IP number i are both 1, record the ID of this data write operation, which is recorded as ID Second ; B4, determine ID First and ID Second If they are equal, the time when the BVALID and BREADY signals on the corresponding IP number i are both 1 is taken as t n .

8. The evaluation method for SoC system optimization according to claim 1, characterized in that: The IP granularity operation performance includes an average read delay of at least one IP and an average write delay of at least one IP.

9. The evaluation method for SoC system optimization according to claim 1, characterized in that: The subsystem-granularity operating performance includes an average read latency of at least one subsystem and an average write latency of at least one subsystem.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which is loaded and executed by a processor to implement the evaluation method for SoC system optimization described in any one of claims 1-11.