Network-on-chip test system and method
By using a pseudo-computing core for NoC testing in the SoC system, the NoC data throughput pressure problem is solved, and the time overlap between NoC testing and computing core design is achieved, thereby improving SoC development efficiency.
Patent Information
- Application Number
- CN202511296850.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-11
AI Technical Summary
In SoC systems, the increased data throughput pressure on NoC makes it difficult to optimize and adjust NoC during the design phase, affecting the design cycle and efficiency of the SoC.
A pseudo-computing core is used to replace the computing core for memory access function testing. The data transmission throughput is detected through the parameter configuration module and the test data acquisition module. The results are analyzed using the data analysis module and optimized through the data dump module.
Shorten NoC testing time, improve testing efficiency, and achieve overlap between NoC testing and computing core design time, thereby shortening the SoC development cycle.
Smart Images

Figure CN120785799B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of chip testing, and in particular to a network-on-chip testing system and method. BACKGROUND
[0002] A NoC (Network on Chip) is a communication network structure used in a SoC (System on Chip). With the development of SoC systems, the number of integrated chips (or dies) is increasing, and the number of various computing cores in the chips is also increasing, thereby bringing more and more pressure to the data throughput of the NoC.
[0003] In order to ensure the stability of communication when a large number of computing cores access a memory and the balance of the amount of data transmitted between the computing cores when accessing the memory, the NoC needs to be optimized and adjusted in the design stage. Therefore, how to complete the design and optimization of the NoC as early as possible in the design stage to help shorten the design cycle of the SoC has become a problem to be solved. SUMMARY
[0004] Therefore, the present disclosure provides a network-on-chip testing system to help improve the testing efficiency of the NoC, shorten the testing time of the NoC, and help complete the testing of the NoC before the design of the computing cores in the SoC system is completed, so that the NoC testing time can overlap with the design time of the computing cores, thereby helping to shorten the development cycle of the SoC and helping to improve the development efficiency of the SoC.
[0005] The technical solution of the present disclosure is implemented as follows:
[0006] According to an aspect of an embodiment of the present disclosure, a network-on-chip testing system is provided, comprising:
[0007] a parameter configuration module configured to provide a memory access parameter for testing of a network-on-chip under test;
[0008] at least one pseudo computing core coupled to the parameter configuration module and coupled to a memory through the network-on-chip under test, and configured to access the memory through the network-on-chip under test based on the memory access parameter, wherein the pseudo computing core has the same memory access function as a computing core; and
[0009] a test data acquisition module coupled to the network-on-chip under test, and configured to obtain data transmission throughput detection data of the network-on-chip under test by detecting the data transmission throughput of the network-on-chip under test during the access of the at least one pseudo computing core to the memory.
[0010] In a possible implementation manner, the network-on-chip testing system further comprises:
[0011] at least one compute core coupled to the parameter configuration module and to the memory through the on-chip network under test for performing computation and accessing the memory through the on-chip network under test based on the memory access parameters;
[0012] wherein the at least one pseudo compute core has the same external transmission bandwidth as the at least one compute core.
[0013] In a possible implementation, the on-chip network test system further comprises:
[0014] a data analysis module coupled to the test data acquisition module for obtaining a data transmission throughput analysis result of the on-chip network under test according to the data transmission throughput detection data of the on-chip network under test and the preset throughput theoretical data.
[0015] In a possible implementation, the on-chip network test system further comprises:
[0016] a data dump module for dumping the waveform data of the on-chip network under test during the memory access if the data transmission throughput analysis result indicates that the data transmission throughput detection data of the on-chip network under test does not meet the requirement of the throughput theoretical data.
[0017] In a possible implementation, the access of the at least one pseudo compute core to the memory is writing data to the memory.
[0018] The memory access parameters comprise:
[0019] a start parameter, a write address parameter, a write request burst transmission length parameter, a write request length parameter, a write request pen number parameter, a write request step length parameter, and a maximum number of uncompleted write request parameter.
[0020] In a possible implementation, the access of the at least one pseudo compute core to the memory is reading data from the memory.
[0021] The memory access parameters comprise:
[0022] a start parameter, a read address parameter, a read request length parameter, a read request pen number parameter, and a read request step length parameter.
[0023] In a possible implementation, the data transmission throughput detection data comprises:
[0024] at least one of a total bandwidth of memory data access detection of the on-chip network under test and a memory access detection bandwidth of each of the at least one pseudo compute core.
[0025] In a possible implementation, the data transfer throughput detection data comprises at least one of a memory data access detection total bandwidth of the measured network-on-chip and a memory access detection bandwidth of each of the at least one pseudo compute core;
[0026] The throughput theory data comprises at least one of a memory data access theory total bandwidth of the measured network-on-chip and a memory access theory bandwidth of each of the at least one pseudo compute core;
[0027] The data transfer throughput analysis result of the measured network-on-chip comprises at least one of a difference between the memory data access detection total bandwidth and the memory data access theory total bandwidth and a difference between the memory access detection bandwidth and the memory access theory bandwidth of any one of the pseudo compute cores.
[0028] According to another aspect of the embodiments of the present disclosure, a network-on-chip testing method is provided, comprising:
[0029] configuring a memory access parameter to at least one pseudo compute core coupled to a measured network-on-chip;
[0030] The at least one pseudo compute core accesses a memory through the measured network-on-chip based on the configured memory access parameter, wherein the memory is coupled to the measured network-on-chip;
[0031] During the access of the memory by the at least one pseudo compute core, data transfer throughput detection data of the measured network-on-chip is obtained through detection of data transfer throughput of the measured network-on-chip.
[0032] In a possible implementation, the configuring a memory access parameter to at least one pseudo compute core coupled to a measured network-on-chip comprises:
[0033] configuring a write parameter to the at least one pseudo compute core or configuring a read parameter to the at least one pseudo compute core.
[0034] In a possible implementation, the at least one pseudo compute core accesses a memory through the measured network-on-chip based on the configured memory access parameter comprises:
[0035] In a case where the memory access parameter is a write parameter, the at least one pseudo compute core writes data into the memory through the measured network-on-chip based on the configured write parameter;
[0036] In a case where the memory access parameter is a read parameter, the at least one pseudo compute core reads data from the memory through the measured network-on-chip based on the configured read parameter.
[0037] In a possible implementation, the obtaining the data transmission throughput detection data of the measured NoC comprises:
[0038] In a possible implementation, the obtaining the data transmission throughput detection data of the measured NoC comprises:
[0039] In a possible implementation, the obtaining the data transmission throughput detection data of the measured NoC comprises:
[0040] In a possible implementation, the on-chip network testing method further comprises:
[0041] In a possible implementation, the obtaining the data transmission throughput analysis result of the measured NoC comprises:
[0042] In a possible implementation, the data transmission throughput detection data of the measured NoC comprises at least one of a memory write data detection total bandwidth of the measured NoC, a memory write data detection bandwidth of any one of the at least one pseudo computing core, a memory read data detection total bandwidth of the measured NoC, and a memory read data detection bandwidth of any one of the at least one pseudo computing core.
[0043] In a possible implementation, the throughput theoretical data comprises at least one of a memory write data theoretical total bandwidth of the measured NoC, a memory write data theoretical bandwidth of any one of the at least one pseudo computing core, a memory read data theoretical total bandwidth of the measured NoC, and a memory read data theoretical bandwidth of any one of the at least one pseudo computing core.
[0044] The data transmission throughput analysis result of the measured network-on-chip comprises at least one of a difference between the memory write data detection total bandwidth and the memory write data theoretical total bandwidth, a difference between the memory write data detection bandwidth of the arbitrary one of the pseudo computing cores and the memory write data theoretical bandwidth of the arbitrary one of the pseudo computing cores, a difference between the memory read data detection total bandwidth and the memory read data theoretical total bandwidth, and a difference between the memory read data detection bandwidth of the arbitrary one of the pseudo computing cores and the memory read data theoretical bandwidth of the arbitrary one of the pseudo computing cores.
[0045] In a case where the data transmission throughput detection data of the measured network-on-chip comprises the memory write data detection total bandwidth of the measured network-on-chip, the data transmission throughput analysis result of the measured network-on-chip is obtained according to the data transmission throughput detection data of the measured network-on-chip and the preset throughput theoretical data, and comprises:
[0046] A difference between the memory write data detection total bandwidth and the memory write data theoretical total bandwidth of the measured network-on-chip is obtained according to the memory write data detection total bandwidth of the measured network-on-chip and the preset memory write data theoretical total bandwidth of the measured network-on-chip.
[0047] In a case where the data transmission throughput detection data of the measured network-on-chip comprises the memory write data detection bandwidth of the arbitrary one of the pseudo computing cores, the data transmission throughput analysis result of the measured network-on-chip is obtained according to the data transmission throughput detection data of the measured network-on-chip and the preset throughput theoretical data, and comprises:
[0048] A difference between the memory write data detection bandwidth and the memory write data theoretical bandwidth of the arbitrary one of the pseudo computing cores is obtained according to the memory write data detection bandwidth of the arbitrary one of the pseudo computing cores and the preset memory write data theoretical bandwidth of the arbitrary one of the pseudo computing cores.
[0049] In a case where the data transmission throughput detection data of the measured network-on-chip comprises the memory read data detection total bandwidth of the measured network-on-chip, the data transmission throughput analysis result of the measured network-on-chip is obtained according to the data transmission throughput detection data of the measured network-on-chip and the preset throughput theoretical data, and comprises:
[0050] A difference between the memory read data detection total bandwidth and the memory read data theoretical total bandwidth of the measured network-on-chip is obtained according to the memory read data detection total bandwidth of the measured network-on-chip and the preset memory read data theoretical total bandwidth of the measured network-on-chip.
[0051] In a case that the data transmission throughput detection data of the measured NoC includes memory read data detection bandwidth of any one of the at least one pseudo computing core, the data transmission throughput analysis result of the measured NoC is obtained according to the data transmission throughput detection data of the measured NoC and the preset throughput theoretical data, including:
[0052] According to the memory read data detection bandwidth of the any one pseudo computing core and the preset memory read data theoretical bandwidth of the any one pseudo computing core, the difference between the memory read data detection bandwidth and the memory read data theoretical bandwidth of the any one pseudo computing core is obtained.
[0053] In a possible implementation, after the data transmission throughput analysis result of the measured NoC is obtained, the NoC testing method further includes:
[0054] In a case that the data transmission throughput analysis result indicates that the data transmission throughput detection data of the measured NoC does not meet the requirement of the throughput theoretical data, the waveform data of the measured NoC during the access of the memory is dumped.
[0055] As can be seen from the above solutions, the NoC testing system and method of the present disclosure uses a pseudo computing core to replace a computing core to realize the memory access function of the computing core, so that on the one hand, the pseudo computing core can be used to enter the testing stage of the NoC in advance when the design of the computing core is not completed, and the testing time of the NoC can be overlapped with the design time of the computing core, thereby helping to shorten the development cycle of the SoC and improve the development efficiency of the SoC. On the other hand, because the pseudo computing core does not have the computing function of the computing core, but only has the same memory access function as the computing core, the use of the pseudo computing core also helps to save the time required for the computing of the computing core, thereby helping to shorten the testing time of the measured NoC and improve the testing efficiency of the measured NoC. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 is a connection structure diagram of a computing core, a NoC and a memory in a chip in the related art;
[0057] Figure 2 is an embodiment structure diagram of a NoC testing system according to an illustrative embodiment;
[0058] Figure 3 is another embodiment structure diagram of a NoC testing system according to an illustrative embodiment;
[0059] Figure 4 is a memory space structure diagram corresponding to a memory write data parameter according to an illustrative embodiment;
[0060] Figure 5 is a memory space structure diagram corresponding to a memory read data parameter according to an illustrative embodiment;
[0061] Figure 6 is a flowchart of a network-on-chip test method according to an illustrative embodiment;
[0062] Figure 7 is a step diagram of one specific application scenario of the network-on-chip test method according to an illustrative embodiment.
[0063] In the drawings, the components represented by the numbers are as follows:
[0064] 101, a chiplet,
[0065] 1011, a computing core,
[0066] 102, a NoC,
[0067] 103, a memory,
[0068] 1031, a level 2 cache,
[0069] 1032, a memory,
[0070] 201, a parameter configuration module,
[0071] 202, a pseudo computing core,
[0072] 203, a test data acquisition module,
[0073] 204, a network-on-chip under test,
[0074] 205, a data analysis module,
[0075] 206, a data dump module. DETAILED DESCRIPTION
[0076] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the present disclosure is further described in detail below with reference to the drawings and examples.
[0077] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0078] As used in the specification and claims of this disclosure, “coupled” or “connected” can mean either a direct or indirect connection, such that a first device is coupled (or connected) to a second device, which means that the first device can be directly connected to the second device or that some other device or connection is intervening.
[0079] Figure 1 is a schematic diagram of a connection structure of a computing core, a NoC and a memory in a chip in the related art. As shown in Figure 1 , a chip, in particular a SoC chip, can include a plurality of small chips 101, each of which includes a plurality of computing cores 1011, and the computing cores 1011 are coupled to a memory 103 through a NoC 102. In order to improve the read and write speed of the memory 103, the memory 103 further includes a second-level cache 1031 and a memory 1032. Because the second-level cache 1031 has a higher access speed, the frequently used data in the memory 1032 and the data just accessed from the memory 103 are usually backed up in the second-level cache 1031 for fast access by each small chip 101.
[0080] In the process of chip development, each component of the chip needs to be tested to ensure that the chip meets the design requirements. The complexity of the chip structure causes the development progress of each component in the chip to be different. If the chip is designed in its entirety and then the chip as a whole or each part of the chip is tested, and then modified or passed according to the test results, the design and testing will be alternated in time, and the development time and testing time cannot be executed synchronously, which makes it difficult to shorten the chip development cycle. Therefore, testing the completed design part of the chip during the design process will help shorten the chip development cycle.
[0081] Referring to Figure 1 , in the chip development stage, if the NoC 102 needs to be tested, the computing cores 1011 coupled to the NoC 102 and the memory 103 need to be involved. In one possible case, the NoC 102 and the memory 103 have completed the design, but the computing cores 1011 have not completed the entire design, so it is difficult to implement the test of the NoC 102; in another possible case, the NoC 102, the memory 103 and all the computing cores 1011 have completed the design, so the NoC 102 can be tested, but in the testing process, the computing cores 1011 need to occupy part of the time in the testing process due to the implementation of their own related computing functions, so the testing of the NoC 102 will be delayed due to the waiting for the completion of the computing task of the computing cores 1011.
[0082] Therefore, the embodiments of the present disclosure provide a network-on-chip test system and method to help improve the test efficiency of the NoC, shorten the test time of the NoC, and help complete the test of the NoC before the design of the computing core in the SoC is completed, so that the test time of the NoC can overlap with the design time of the computing core, thereby helping to shorten the development cycle of the SoC and improve the development efficiency of the SoC.
[0083] Figure 2 is an embodiment structure schematic diagram of a network-on-chip test system according to an illustrative embodiment, as shown in Figure 2 The network-on-chip test system mainly includes a parameter configuration module 201, a pseudo computing core 202, and a test data acquisition module 203. The parameter configuration module 201 is configured to provide memory access parameters for a test of a network-on-chip under test (also referred to as a test NoC) 204. The number of the pseudo computing core 202 is at least one. The at least one pseudo computing core 202 is coupled to the parameter configuration module 201 and coupled to the memory 103 through the test NoC 204. The pseudo computing core 202 is configured to access the memory 103 through the test NoC 204 based on the memory access parameters, wherein the pseudo computing core 202 has the same memory access function as the computing core 1011 but does not have the computing function of the computing core 1011. The test data acquisition module 203 is coupled to the test NoC 204 and configured to obtain data transmission throughput detection data of the test NoC 204 by detecting the data transmission throughput of the test NoC 204 during the access of the memory 103 by the at least one pseudo computing core 202. In an illustrative embodiment, the test data acquisition module 203 can be implemented by a bus monitor.
[0084] In the network-on-chip test system of the embodiments of the present disclosure, the pseudo computing core is used to replace the computing core to implement the memory access function of the computing core, so that on the one hand, the pseudo computing core can be used to enter the test phase of the NoC in advance when the design of the computing core is not completed, and on the other hand, because the pseudo computing core does not have the computing function of the computing core but only has the same memory access function as the computing core, the use of the pseudo computing core also helps to save the time required for the computing of the computing core, thereby helping to shorten the test time of the test NoC and improve the test efficiency of the test NoC.
[0085] In an illustrative embodiment, in the case that part of the computing cores are completed in design and test, the network-on-chip test system can further include the computing cores completed in design and test. Figure 3 is another embodiment structure schematic diagram of a network-on-chip test system according to an illustrative embodiment, as shown in Figure 3 The network-on-chip test system includes the parameter configuration module 201, the pseudo computing core 202, and the test data acquisition module 203. Figure 2In addition to the components of the illustrated embodiments, the computing cores 1011 are included, the number of the computing cores 1011 is at least one, the at least one computing core 1011 is coupled to the parameter configuration module 201 and coupled to the memory 103 through the measured on-chip network 204, and is configured to perform computation and access the memory 103 through the measured on-chip network 204 based on the memory access parameter. In the illustrative embodiment, the external transmission bandwidth of the at least one pseudo computing core 202 is the same as the external transmission bandwidth of the at least one computing core 1011.
[0086] In the illustrative embodiment, the computing cores 1011 can be at least one of a vector core or a tensor core. For example, in the case of one computing core 1011, the computing core 1011 can be a vector core or a tensor core, in the case of more than one computing core 1011, all of the computing cores 1011 can be vector cores, or all of the computing cores 1011 can be tensor cores, or a part of the computing cores 1011 are vector cores and the other part of the computing cores 1011 are tensor cores.
[0087] Figure 3 The computing cores 1011 are introduced in the on-chip network test system of the illustrated embodiments, so that the test process of the NoC 102 can be closer to the real running scenario of the chip, and the data transmission throughput detection data of the measured on-chip network 204 can be more reliable.
[0088] For Figure 3 In order to shorten the test time of the measured on-chip network 204 as much as possible and improve the test efficiency of the measured on-chip network 204, in the on-chip network test system of the illustrated embodiments, the number of the pseudo computing cores 202 is greater than the number of the computing cores 1011.
[0089] Because the number of the pseudo computing cores 202 is greater than the number of the computing cores 1011, it helps to reduce the proportion of the time occupied by the computing cores 1011 in performing the computing task in the test process in the whole test process, and thus helps to achieve a better balance between the reliability of the data transmission throughput detection data of the measured on-chip network 204 and the test efficiency of the measured on-chip network 204.
[0090] As Figure 2 , Figure 3 In order to obtain the test conclusion of the measured on-chip network 204, in the illustrative embodiment, the on-chip network test system of the embodiments of the present disclosure further includes a data analysis module 205. The data analysis module 205 is coupled to the test data acquisition module 203, and is configured to obtain the data transmission throughput analysis result of the measured on-chip network 204 according to the data transmission throughput detection data of the measured on-chip network 204 and the preset throughput theoretical data.
[0091] AsFigure 2 、 Figure 3 As shown in FIG. 2, in an illustrative embodiment, to facilitate modifying and optimizing the measured NoC 204 according to the analysis result, in an illustrative embodiment, the NoC test system of the embodiments of the present disclosure further comprises a data dump module 206. The data dump module 206 is configured to dump the waveform data of the measured NoC 204 during the access of the memory 103 in a case where the data transfer throughput analysis result indicates that the data transfer throughput detection data of the measured NoC 204 does not meet the throughput theoretical data requirement.
[0092] In an illustrative embodiment, the data transfer throughput detection data comprises at least one of a memory data access detection total bandwidth of the measured NoC 204 and a memory access detection bandwidth of each of the at least one pseudo compute core 202.
[0093] In an illustrative embodiment, the throughput theoretical data comprises at least one of a memory data access theoretical total bandwidth of the measured NoC 204 and a memory access theoretical bandwidth of each of the at least one pseudo compute core 202.
[0094] In an illustrative embodiment, the data transfer throughput analysis result of the measured NoC 204 comprises at least one of a difference between the memory data access detection total bandwidth and the memory data access theoretical total bandwidth and a difference between the memory access detection bandwidth and the memory access theoretical bandwidth of any one of the at least one pseudo compute core 202. In a case where at least one of the difference between the memory data access detection total bandwidth and the memory data access theoretical total bandwidth is less than a preset total bandwidth difference threshold and the difference between the memory access detection bandwidth and the memory access theoretical bandwidth of any one of the at least one pseudo compute core 202 is less than a preset compute core bandwidth difference threshold, the data transfer throughput detection data of the measured NoC 204 does not meet the throughput theoretical data requirement. In a case where the NoC test system of the embodiments of the present disclosure further comprises the compute core 1011, in a case where at least one of the difference between the memory data access detection total bandwidth and the memory data access theoretical total bandwidth is less than a preset total bandwidth difference threshold, the difference between the memory access detection bandwidth and the memory access theoretical bandwidth of any one of the at least one pseudo compute core 202 is less than a preset compute core bandwidth difference threshold, and the difference between the memory access detection bandwidth and the memory access theoretical bandwidth of any one of the compute cores 1011 is less than a preset compute core bandwidth difference threshold, the data transfer throughput detection data of the measured NoC 204 does not meet the throughput theoretical data requirement.
[0095] As shown in FIG. 2, in an illustrative embodiment, to facilitate modifying and optimizing the measured NoC 204 according to the analysis result, in an illustrative embodiment, the NoC test system of the embodiments of the present disclosure further comprises a data dump module 206. The data dump module 206 is configured to dump the waveform data of the measured NoC 204 during the access of the memory 103 in a case where the data transfer throughput analysis result indicates that the data transfer throughput detection data of the measured NoC 204 does not meet the throughput theoretical data requirement. Figure 2 、 Figure 3As shown, in the illustrative embodiment, the memory 103 includes a level 2 cache 1031 and a main memory 1032. The level 2 cache 1031 is coupled to the on-chip network under test 204, and the main memory 1032 is coupled to the level 2 cache 1031. In the illustrative embodiment, during the on-chip network test, the memory 103 performs relevant data write and read operations based on its own design rules to cooperate with the data access of the various pseudo-compute cores 202 and compute cores 1011.
[0096] In the illustrative embodiment, the access to the memory 103 can include writing data to the memory 103 and reading data from the memory 103. The on-chip network test system of the embodiments of the present disclosure is described below with respect to writing data and reading data, respectively.
[0097] In the illustrative embodiment, the access to the memory 103 by the at least one pseudo-compute core 202 is writing data to the memory 103, and in the case that the on-chip network test system further includes at least one compute core 1011, the access to the memory 103 by the at least one compute core 1011 is also writing data to the memory 103. In this case, during the process in which all the pseudo-compute cores 202 and compute cores 1011 simultaneously write data to the memory 103, the memory write data detection total bandwidth of the on-chip network under test 204 can be detected by the test data acquisition module 203, and the memory write data detection bandwidth of each pseudo-compute core 202 and the memory write data detection bandwidth of each compute core 1011 can also be obtained.
[0098] In the illustrative embodiment, in the case that the access to the memory 103 is writing data to the memory 103, the memory access parameters can be referred to as memory write data parameters, or can be simply referred to as write parameters. In the illustrative embodiment, the memory write data parameters include a start parameter, a stop parameter, a write address parameter, a write request burst length parameter, a write request length parameter, a write request pen number parameter, a write request step parameter, and a maximum outstanding write request number parameter.
[0099] In the hardware system of the chip, the control of the pseudo-computing core 202 and the computing core 1011 is usually implemented by registers, and therefore, the registers configured by the memory write data parameters include a start register (cfg_start), an end register (cfg_end), a write address register (write_address), a write request burst transmission length register (write_burst_length), a write request length register (write_length), a write request number register (write_num), a write request stride register (write_stride), and a maximum number of outstanding write requests register (write_outstanding). In the illustrative embodiment, the start register is configured to 1 to indicate that the write request and write data are started to be issued and the write return signal is received; the end register is configured to 1 to indicate that the write request is immediately stopped to be issued, and therefore, based on the end register, the tester can stop the pseudo-computing core 202 from issuing the write request and data in the middle of the process; the write address register is configured to the start memory address of the issued write request; the write request burst transmission length register is configured to the burst transmission length of each write request, i.e., the data length corresponding to each write request, in the illustrative embodiment, burst_length=0 represents that the data length corresponding to each write request is 128 bytes, burst_length=1 represents that the data length corresponding to each write request is 256 bytes, and burst_length=3 represents that the data length corresponding to each write request is 512 bytes; the write request length register is configured to the continuous memory address length of each write request; the write request number register is configured to the total number of write requests; the write request stride register is configured to the interval stride between the continuous memory addresses of each two write requests; and the maximum number of outstanding write requests register is configured to the maximum number of outstanding write requests, i.e., when the number of write return signals not received is equal to the maximum number of outstanding write requests, the issuance of the write request is paused to limit the write data flow on the bus.
[0100] In the related art, the SoC writes data to the memory 103 for the computing core 1011, sets the start register, the abort register, the write address register, the write request burst transfer length register, the write request length register, the write request pen number register, the write request step register, and the maximum number of unfinished write request register of each computing core 1011, and these registers are arranged around the designed computing core 1011 for the computing core 1011 to read. Based on this, in the illustrative embodiment, the pseudo computing core 202 can reuse the start register, the abort register, the write address register, the write request burst transfer length register, the write request length register, the write request pen number register, the write request step register, and the maximum number of unfinished write request register required when the computing core 1011 writes data to the memory 103. Therefore, in the on-chip network test system of the embodiment of the disclosure, there is no need to separately design the corresponding registers required when writing data to the memory 103 for each pseudo computing core 202, which can save the design time of the related registers.
[0101] Since the pseudo computing core 202 reuses the related registers required when the computing core 1011 writes data to the memory 103, the pseudo computing core 202 can simulate the real computing core 1011 to issue a write request and write data to the memory 103 and receive a write return signal, so that the tester of each stage can simulate the real data access behavior of the computing core 1011 by configuring these registers without software programming, and then test the write performance of the measured on-chip network 204. Since the pseudo computing core 202 is not a real computing core 1011, the pseudo computing core 202 does not have related computing functions, and therefore the content of the write data can be set according to the test needs. In the illustrative embodiment, the write request ID (Identification, identity) can be used as the content of the write data, so that the tester can directly locate the write request in the measured on-chip network 204 according to the content of the write data, which can make the test process more efficient. For example, in the write test process, the write request IDs of the pseudo computing cores 202 are different from each other, and the write request ID can be one-to-one bound to each pseudo computing core 202 to indicate which write request is issued from which pseudo computing core 202. Therefore, in the test process, the corresponding pseudo computing core 202 can be located according to the content of the write data detected on the measured on-chip network 204, so that the memory write data detection bandwidth of any one pseudo computing core 202 can be detected.
[0102] Figure 4is a memory space structure diagram corresponding to a memory write data parameter according to an illustrative embodiment. Wherein, burst represents a burst transmission length of each write request represented by a write request burst transmission length parameter, length represents a continuous memory address length of each write request represented by a write request length parameter, and it can be seen that the continuous memory address length of each write request contains the burst transmission length of multiple write requests, stride represents an interval step length between continuous memory addresses of each two write requests represented by a write request step length parameter, and address represents a starting memory address of a write request represented by a write address parameter, which is located at a starting position of the first burst, Figure 4 In the diagram, each burst corresponds to a write request respectively, and the total number of bursts is the total number of write requests represented by a write request number parameter.
[0103] In the illustrative embodiment, in the case of writing data to the memory 103, the data transmission throughput detection data includes at least one of a memory write data detection total bandwidth of the measured on-chip network 204 and a memory write data detection bandwidth of each of the at least one pseudo computing core 202.
[0104] In the illustrative embodiment, in the case of writing data to the memory 103, the throughput theoretical data includes at least one of a memory write data theoretical total bandwidth of the measured on-chip network 204 and a memory write data theoretical bandwidth of each of the at least one pseudo computing core 202.
[0105] In the illustrative embodiment, in the case that the access to the memory 103 is writing data to the memory 103, the data transmission throughput analysis result of the measured network-on-chip 204 includes at least one of a difference between a memory write data detection total bandwidth and a memory write data theoretical total bandwidth, and a difference between a memory write data detection bandwidth of any one of the pseudo computing cores 202 and a memory write data theoretical bandwidth. In the case that at least one of the difference between the memory write data detection total bandwidth and the memory write data theoretical total bandwidth is less than a preset total write data bandwidth difference threshold, and the difference between the memory write data detection bandwidth of any one of the pseudo computing cores 202 and the memory write data theoretical bandwidth is less than a preset computing core write data bandwidth difference threshold, the data transmission throughput detection data of the write data of the measured network-on-chip 204 does not meet the throughput theoretical data requirement of the write data. In the case that the on-chip network test system of the embodiment of the present disclosure further includes the computing cores 1011, in the case that at least one of the difference between the memory write data detection total bandwidth and the memory write data theoretical total bandwidth is less than a preset total write data bandwidth difference threshold, the difference between the memory write data detection bandwidth of any one of the pseudo computing cores 202 and the memory write data theoretical bandwidth is less than a preset computing core write data bandwidth difference threshold, and the difference between the memory write data detection bandwidth of any one of the computing cores 1011 and the memory write data theoretical bandwidth is less than a preset computing core write data bandwidth difference threshold, the data transmission throughput detection data of the write data of the measured network-on-chip 204 does not meet the throughput theoretical data requirement.
[0106] In the illustrative embodiment, the access to the memory 103 by the at least one pseudo computing core 202 is reading data from the memory 103, and in the case that the on-chip network test system further includes the at least one computing core 1011, the access to the memory 103 by the at least one computing core 1011 is also reading data from the memory 103. At this time, in the process that all the pseudo computing cores 202 and the computing cores 1011 simultaneously read data from the memory 103, the memory read data detection total bandwidth of the measured network-on-chip 204 can be detected by the test data acquisition module 203, and the memory read data detection bandwidth of each pseudo computing core 202 and the memory read data detection bandwidth of each computing core 1011 can also be obtained.
[0107] In the illustrative embodiment, in the case that the access to the memory 103 is reading data from the memory 103, the memory access parameter can be referred to as a memory read data parameter, or can be simply referred to as a read parameter. In the illustrative embodiment, the memory read data parameter includes a start parameter, a read address parameter, a read request length parameter, a read request pen number parameter, and a read request step length parameter.
[0108] In the illustrative embodiment, corresponding to the memory read data parameters, the configured registers include a start register (cfg_start), a read address register (read_address), a read request length register (read_length), a read request pen number register (read_num), and a read request stride register (read_stride). In the illustrative embodiment, the start register is set to 1 to indicate that a read request is to be issued and the returned read data is to be received; the read address register is configured as the starting memory address of the issued read request; the read request length register is configured as the length of the continuous memory address of each pen of the issued read request; the read request pen number register is configured as the total number of pens of the issued read request; and the read request stride register is configured as the interval stride between the continuous memory addresses of each two pens of the issued read request. The start register involved in the read request and the start register involved in the write request can be multiplexed. In the write access, because in addition to the relevant information in the start register, there are also information in the write address register, the write request burst transmission length register, the write request length register, the write request pen number register, the write request stride register, and the maximum number of unfinished write requests register, which are related to the write request, so whether it is a pseudo-computing core or a computing core, the write access is implemented based on the information in the associated registers related to the write request. Similarly, in the read access, because in addition to the relevant information in the start register, there are also information in the read address register, the read request length register, the read request pen number register, and the read request stride register, which are related to the read request, so whether it is a pseudo-computing core or a computing core, the read access is implemented based on the information in the associated registers related to the read request. Moreover, for the same pseudo-computing core or computing core, the write access and the read access cannot be performed simultaneously. Therefore, the read request and the write request can multiplex the same start register.
[0109] In the related art, the SoC reads data from the memory 103 for the compute core 1011, sets the start-up register, the read address register, the read request length register, the read request pen number register, and the read request stride register of each compute core 1011, and these registers are arranged around the designed compute core 1011 for the compute core 1011 to read. Based on this, in the illustrative embodiment, the pseudo compute core 202 can reuse the start-up register, the read address register, the read request length register, the read request pen number register, and the read request stride register required when the compute core 1011 reads data from the memory 103, so that in the on-chip network test system of the embodiment of the disclosure, there is no need to separately design the registers required for reading data from the memory 103 for each pseudo compute core 202, and the design time of the related registers can be saved. Since the pseudo compute core 202 reuses the related registers required when the compute core 1011 reads data from the memory 103, the pseudo compute core 202 can simulate the real compute core 1011 issuing a read request to the memory 103 and receiving the read data returned by the memory 103, so that the tester of each stage can simulate the real behavior of the compute core 1011 reading data by configuring these registers without software programming, and then test the read performance of the measured on-chip network 204.
[0110] Figure 5 is a memory space structure diagram corresponding to the memory read data parameters according to an illustrative embodiment. Wherein, length represents the continuous memory address length of each read request represented by the read request length parameter, stride represents the interval stride between the continuous memory addresses of each two read requests represented by the read request stride parameter, and address represents the starting memory address of the read request represented by the read address parameter, which is located at the starting position of the first length, Figure 5 In the diagram, the total number of length is the total number of read requests represented by the read request pen number parameter.
[0111] In the illustrative embodiment, in the case of accessing the memory 103 to read data from the memory 103, the data transmission throughput detection data includes at least one of the memory read data detection total bandwidth of the measured on-chip network 204 and the memory read data detection bandwidth of each of the at least one pseudo compute core 202.
[0112] In the illustrative embodiment, in the case of accessing the memory 103 to read data from the memory 103, the throughput theoretical data includes at least one of the memory read data theoretical total bandwidth of the measured on-chip network 204 and the memory read data theoretical bandwidth of each of the at least one pseudo compute core 202.
[0113] In the illustrative embodiment, in the case of accessing the memory 103 to read data from the memory 103, the data transmission throughput analysis result of the measured on-chip network 204 includes at least one of a difference between a memory read data detection total bandwidth and a memory read data theoretical total bandwidth, and a difference between a memory read data detection bandwidth of any one of the pseudo computing cores 202 and a memory read data theoretical bandwidth of the any one of the pseudo computing cores 202. If at least one of the difference between the memory read data detection total bandwidth and the memory read data theoretical total bandwidth is less than a preset total read data bandwidth difference threshold, and the difference between the memory read data detection bandwidth of any one of the pseudo computing cores 202 and the memory read data theoretical bandwidth of the any one of the pseudo computing cores 202 is less than a preset computing core read data bandwidth difference threshold, then the data transmission throughput detection data of the read data of the measured on-chip network 204 does not meet the throughput theoretical data requirement of the read data. In the case that the on-chip network test system of the embodiment of the present disclosure further includes the computing core 1011, if at least one of the difference between the memory read data detection total bandwidth and the memory read data theoretical total bandwidth is less than a preset total read data bandwidth difference threshold, the difference between the memory read data detection bandwidth of any one of the pseudo computing cores 202 and the memory read data theoretical bandwidth of the any one of the pseudo computing cores 202 is less than a preset computing core read data bandwidth difference threshold, and the difference between the memory read data detection bandwidth of any one of the computing cores 1011 and the memory read data theoretical bandwidth of the any one of the computing cores 1011 is less than a preset computing core read data bandwidth difference threshold, then the data transmission throughput detection data of the read data of the measured on-chip network 204 does not meet the throughput theoretical data requirement.
[0114] In the illustrative embodiment, according to different designs, the implementation of at least one of the parameter configuration module 201, the pseudo computing core 202, the computing core 1011, the measured on-chip network 204, the memory 103, the parameter configuration module 201, the test data acquisition module 203, the data analysis module 205, and the data dump module 206 can be a combination of multiple of hardware, firmware, and software (i.e., programs).
[0115] In hardware form, at least one of the parameter configuration module 201, the pseudo computing core 202, the computing core 1011, the measured on-chip network 204, the memory 103, the parameter configuration module 201, the test data acquisition module 203, the data analysis module 205, and the data dump module 206 can be implemented as logic circuitry on an integrated circuit. For example, the functions of at least one of the parameter configuration module 201, the pseudo computing core 202, the computing core 1011, the measured on-chip network 204, the memory 103, the parameter configuration module 201, the test data acquisition module 203, the data analysis module 205, and the data dump module 206 can be implemented in various logic blocks, modules, and circuits of one or more hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), central processing units (CPUs), or other processing units. The functions of at least one of the parameter configuration module 201, the pseudo computing core 202, the computing core 1011, the measured on-chip network 204, the memory 103, the parameter configuration module 201, the test data acquisition module 203, the data analysis module 205, and the data dump module 206 can be implemented as hardware circuits, such as various logic blocks, modules, and circuits in an integrated circuit, using hardware description languages (such as Verilog HDL or VHDL) or other suitable programming languages.
[0116] In software or firmware, the functions of at least one of the parameter configuration module 201, the pseudo-compute core 202, the compute core 1011, the measured NoC 204, the memory 103, the parameter configuration module 201, the test data acquisition module 203, the data analysis module 205, and the data dump module 206 can be implemented as programming codes. For example, at least one of the parameter configuration module 201, the pseudo-compute core 202, the compute core 1011, the measured NoC 204, the memory 103, the parameter configuration module 201, the test data acquisition module 203, the data analysis module 205, and the data dump module 206 can be implemented by using a general programming language (such as C, C++, or assembly language) or other suitable programming language. The programming codes can be recorded and stored in a non-transitory machine-readable storage medium. In some embodiments, the non-transitory machine-readable storage medium includes, for example, a semiconductor memory and / or a storage device. An electronic device (such as a CPU, a hardware controller, a microcontroller, a hardware processor, or a microprocessor) can read and execute the programming codes from the non-transitory machine-readable storage medium, thereby implementing the functions of at least one of the parameter configuration module 201, the pseudo-compute core 202, the compute core 1011, the measured NoC 204, the memory 103, the parameter configuration module 201, the test data acquisition module 203, the data analysis module 205, and the data dump module 206.
[0117] In illustrative embodiments, the NoC testing apparatus of the present disclosure is applicable to a SoC chip, where the SoC chip can be any one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), a NPU (Neural network Processing Unit), a DPU (Deep learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit).
[0118] The embodiment of the present disclosure further discloses a network-on-chip testing method. The network-on-chip testing method can be used to further understand the network-on-chip testing system of the embodiment of the present disclosure.
[0119] Figure 6 Fig. 1 is a flowchart showing a network-on-chip testing method according to an illustrative embodiment. As shown in Fig. 1, the network-on-chip testing method mainly includes the following steps 601-603. Figure 6
[0120] Step 601: configuring a memory access parameter for at least one pseudo computing core coupled to a network-on-chip under test.
[0121] Step 602: accessing a memory by the at least one pseudo computing core based on the configured memory access parameter through the network-on-chip under test, wherein the memory is coupled to the network-on-chip under test.
[0122] Step 603: obtaining data transmission throughput detection data of the network-on-chip under test by detecting a data transmission throughput of the network-on-chip under test during the access of the memory by the at least one pseudo computing core.
[0123] In the illustrative embodiment, the network-on-chip testing method of the embodiment of the present disclosure can be applied to the network-on-chip testing system of the above-mentioned embodiments of the present disclosure.
[0124] Based on the access to the memory including both write access and read access, in the illustrative embodiment, step 601 can specifically include: configuring a write parameter for the at least one pseudo computing core or configuring a read parameter for the at least one pseudo computing core.
[0125] Based on the access to the memory including both write access and read access, in the illustrative embodiment, step 602 can specifically include: in the case that the memory access parameter is a write parameter, the at least one pseudo computing core writes data into the memory through the network-on-chip under test based on the configured write parameter; in the case that the memory access parameter is a read parameter, the at least one pseudo computing core reads data from the memory through the network-on-chip under test based on the configured read parameter. Wherein the write parameter includes a start parameter, a stop parameter, a write address parameter, a write request burst transmission length parameter, a write request length parameter, a write request pen number parameter, a write request step parameter, and a maximum number of unfinished write request parameter; the read parameter includes a start parameter, a read address parameter, a read request length parameter, a read request pen number parameter, and a read request step parameter.
[0126] In the illustrative embodiment, in the write access process, when the arbitrary pseudo-computing core reads that the start register is 1, it starts to send write requests to the memory through the measured on-chip network from the start memory address recorded in the write address register, each write request corresponds to write data with a length specified by the write request length register, the interval step length between the continuous memory addresses of two write requests is specified by the write request step length register, the total number of write requests sent is specified by the write request number register, and after receiving the write return signal returned by the measured on-chip network for the number of write requests specified by the write request number register, the write access process of the arbitrary pseudo-computing core is stopped. When the arbitrary pseudo-computing core reads that the stop register is 1, it immediately stops sending write requests, and the write data corresponding to the write requests that have been sent continues to be sent until the number of write requests corresponding to the write requests that have been sent is reached, and after receiving the return signals of all the sent write requests, all the behaviors of the arbitrary pseudo-computing core are stopped. The write data corresponding to each write request is aligned with the write request, and the write data will not be sent before the corresponding write request, so that the write data sent when the stop occurs will not be more than the write request.
[0127] In the illustrative embodiment, the pseudo-computing core includes an unfinished write request counter, which is used to count the number of unfinished write requests. When the value of the unfinished write request counter is greater than or equal to the configuration value of the maximum unfinished write request number register, the pseudo-computing core stops sending write requests, and only when the value of the unfinished write request counter is less than the configuration value of the maximum unfinished write request number register, the pseudo-computing core can continuously send write requests, so that the traffic control on the measured on-chip network when writing data to the memory can be realized. In the illustrative embodiment, the sending of write requests and write data by the pseudo-computing core and the receiving of write returns are all based on the communication protocol designed inside the SoC chip, so that when the write return signal is received, the bus will not be back-pressured.
[0128] As can be seen, in the on-chip network testing system and method of this disclosure, a pseudo-computing core is used to simulate the writing of data from a computing core to memory. This only requires reusing the relevant registers of the real computing core to simulate the write requests, data transmission, and write return signal recovery of multiple cores within the chip. Therefore, it can be directly used to test the write performance of NoC and inter-chip buses. Furthermore, because the pseudo-computing core only has memory access functions and lacks other functions such as computation, it helps reduce the complexity of related software programming. In addition, in this disclosure, the write data flow of the pseudo-computing core on the NoC can be directly controlled through registers, making flow control on the NoC more flexible and convenient, and simulating flow closer to real-world scenarios. Moreover, the content of the write data of the pseudo-computing core is a write request ID, thereby enabling customization of the write data content and facilitating write data tracking on the NoC bus. The corresponding write request and the pseudo-computing core that issued the write request can be directly located based on the content of the write data, which helps improve the efficiency of solving related problems.
[0129] In an illustrative embodiment, during the read access process, when any pseudo-computing core reads that the start register is 1, it sends a read request to the memory through the network under test (BIT) starting from the starting memory address recorded in the read address register to read data from the starting memory address recorded in the read address register. The length of the read data corresponding to each read request is the length specified by the read request length register, the interval step size between consecutive memory addresses of every two read requests is the interval step size specified by the read request step size register, and the total number of read requests sent is the number of read requests specified by the read request count register. After receiving the number of read data specified by the read request count register returned by the BIT, the read access process of that pseudo-computing core is stopped.
[0130] In the illustrative embodiment, the pseudo-computing core sends read requests and receives read data based on the communication protocol designed inside the SoC chip, so that the bus is not pressured when receiving read data.
[0131] As can be seen, the on-chip network testing system and method of this disclosure uses a pseudo-computing core that mimics a computing core reading data from memory. This only requires reusing the relevant registers of the real computing core to simulate the issuance of read requests from multiple cores within the chip and the reclamation of read data. Therefore, it can be directly used to test the read performance of NoC and inter-chip buses. Furthermore, because the pseudo-computing core only has memory access functions and lacks other functions such as computation, it helps reduce the complexity of related software programming. Since the relevant registers of the real computing core are reused, and these registers are built into the chip, no new hardware configuration consumption is added to the pseudo-computing core. Moreover, in this disclosure, the read data flow of the pseudo-computing core on the on-chip network under test can be directly controlled through registers. Therefore, the flow control on the on-chip network under test is more flexible and convenient, and it can also simulate flow that is closer to real-world scenarios. The on-chip network testing system and method of this disclosure are applicable to different testing stages before and after chip tape-out.
[0132] Since memory access includes both write access and read access, in the illustrative embodiment, step 603 may specifically include:
[0133] 1) Data transfer throughput detection for write access: During the writing of data to memory by at least one pseudo-computing core, the total memory write data detection bandwidth of the on-chip network under test and at least one of the memory write data detection bandwidth of each pseudo-computing core in at least one pseudo-computing core are obtained by detecting the data transfer throughput of the on-chip network under test.
[0134] 2) Data transfer throughput detection for read access: During the reading of data from memory by at least one pseudo-computing core, the total memory read data detection bandwidth of the on-chip network under test and at least one of the memory read data detection bandwidth of each pseudo-computing core in at least one pseudo-computing core are obtained by detecting the data transfer throughput of the on-chip network under test.
[0135] In a specific application scenario, a large network topology is deployed on a hardware emulator, including 64 pseudo-computing cores connected to the L2 cache and memory via the on-chip network under test (BTC). The bandwidth of the BTC is defined as the amount of data transmitted by the NoC per unit time. The external transmission bandwidth of the pseudo-computing cores is consistent with that of the real computing cores. Assume that the maximum write and read data transmission bandwidth of both the pseudo-computing cores and the real computing cores is 80 GB / s, and the maximum designed write and read data bandwidth of the BTC is 3000 GB / s.
[0136] In this specific application scenario, 1) the data transmission throughput detection for write access can specifically include: configuring the maximum continuous memory address length of each write request for each of the 64 pseudo-computing cores through their respective write request length registers, so that each pseudo-computing core sends write requests to the on-chip network under test at full bandwidth; measuring the amount of write data transmitted by the on-chip network under test per unit time, the actual write bandwidth of the on-chip network under test (i.e., the total bandwidth of memory write data detection of the on-chip network under test) can be obtained; the 64 pseudo-computing cores start simultaneously and send write requests with the same amount of tasks, measuring the completion time of the corresponding write request of each pseudo-computing core, the total amount of write data of the write request of each pseudo-computing core is a known configuration value, and the write bandwidth of each pseudo-computing core (i.e., the memory write data detection bandwidth of each pseudo-computing core in at least one pseudo-computing core) can be obtained by dividing the total amount of write data of each pseudo-computing core by the completion time of each pseudo-computing core.
[0137] In this specific application scenario, 2) the data transfer throughput detection for read access can specifically include: configuring the continuous memory address length of each read request for each of the 64 pseudo-computing cores through their respective read request length registers, so that each pseudo-computing core sends read requests to the on-chip network under test at full bandwidth; measuring the amount of read data transmitted by the on-chip network under test per unit time, the actual read bandwidth of the on-chip network under test (i.e., the total bandwidth of memory read data detection of the on-chip network under test) can be obtained; the 64 pseudo-computing cores start simultaneously and send read requests with the same amount of tasks, measuring the completion time of the corresponding read request of each pseudo-computing core, the total amount of read data of each pseudo-computing core's read request is a known configuration value, and the read bandwidth of each pseudo-computing core (i.e., the memory read data detection bandwidth of each pseudo-computing core in at least one pseudo-computing core) can be obtained by dividing the total amount of read data of each pseudo-computing core by the completion time of each pseudo-computing core.
[0138] The purpose of the test is not only to obtain data transmission throughput test data of the network under test (NIC), but also to draw relevant conclusions to determine whether it meets the corresponding design requirements. Based on this, in the illustrative embodiment, such as Figure 6 As shown, after obtaining the data transmission throughput test data of the on-chip network under test in step 603, the on-chip network testing method further includes:
[0139] Step 604: Based on the data transmission throughput test data of the network under test and the preset theoretical throughput data, obtain the data transmission throughput analysis results of the network under test.
[0140] In the illustrative embodiment, since the data transmission throughput detection of the network under test includes 1) data transmission throughput detection for write access and 2) data transmission throughput detection for read access, and 1) the data transmission throughput detection for write access also includes the result of at least one of the total memory write data detection bandwidth of the network under test and the memory write data detection bandwidth of each pseudo-computing core in at least one pseudo-computing core, and 2) the data transmission throughput detection for read access also includes the result of at least one of the total memory read data detection bandwidth of the network under test and the memory read data detection bandwidth of each pseudo-computing core in at least one pseudo-computing core, step 604 will be explained from the two aspects of data transmission throughput detection for write access and data transmission throughput detection for read access, and their respective results.
[0141] In an illustrative embodiment, the data transmission throughput detection data of the network under test includes at least one of the following: the total memory write data detection bandwidth of the network under test, the memory write data detection bandwidth of any one of the at least one pseudo-computing cores, the total memory read data detection bandwidth of the network under test, and the memory read data detection bandwidth of any one of the at least one pseudo-computing cores.
[0142] In an illustrative embodiment, the theoretical throughput data includes at least one of the following: a preset theoretical total bandwidth of memory write data of the network under test, a preset theoretical bandwidth of memory write data of any pseudo-computing core, a preset theoretical total bandwidth of memory read data of the network under test, and a preset theoretical bandwidth of memory read data of any pseudo-computing core.
[0143] In an illustrative embodiment, the data transmission throughput analysis results of the network under test include at least one of the following: the difference between the total bandwidth detected by memory write data and the theoretical total bandwidth detected by memory write data, the difference between the bandwidth detected by memory write data of any pseudo-computing core and the theoretical bandwidth detected by memory write data of that pseudo-computing core, the difference between the total bandwidth detected by memory read data and the theoretical total bandwidth detected by memory read data, and the difference between the bandwidth detected by memory read data of any pseudo-computing core and the theoretical bandwidth detected by memory read data of that pseudo-computing core.
[0144] In an illustrative embodiment, when the data transmission throughput detection data of the network-on-chip under test includes the total bandwidth of the memory write data detection of the network-on-chip under test, step 604 includes:
[0145] The difference between the total bandwidth for memory write data detection and the theoretical total bandwidth for memory write data is obtained based on the total bandwidth for memory write data detection of the network under test and the preset theoretical total bandwidth for memory write data of the network under test.
[0146] In the specific application scenario example above, assume that the maximum transmission bandwidth for writing and reading data for both the pseudo-computing core and the computing core is 80 GB / s, and the maximum designed bandwidth for writing and reading data for the on-chip network under test is 3000 GB / s. Assuming that the currently measured total bandwidth for memory write data detection in the on-chip network under test is 2000 GB / s, then the difference between the total bandwidth for memory write data detection and the theoretical total bandwidth for memory write data is 2000 GB / s - 3000 GB / s = -1000 GB / s. This indicates a difference of 1000 GB / s, and the total bandwidth for memory write data detection does not meet the design requirement of the theoretical total bandwidth for memory write data.
[0147] In an illustrative embodiment, when the data transmission throughput detection data of the network-on-chip under test includes the memory write data detection bandwidth of any one of the pseudo-computing cores, step 604 includes:
[0148] Based on the memory write data detection bandwidth of any pseudo-computing core and the preset theoretical memory write data bandwidth of any pseudo-computing core, the difference between the memory write data detection bandwidth and the theoretical memory write data bandwidth of any pseudo-computing core is obtained.
[0149] In the specific application scenario example above, it is assumed that the maximum transmission bandwidth for writing and reading data for both the pseudo computing core and the computing core is 80 GB / s, and the maximum designed bandwidth for writing and reading data for the on-chip network under test is 3000 GB / s. Assuming that the actual write bandwidth (i.e., memory write data detection bandwidth) of one of the 64 pseudo-computing cores (the first, third, and fourth pseudo-computing cores) is 80 GB / s, and the actual write bandwidth of one of the 64 pseudo-computing cores (the second pseudo-computing core) is 50 GB / s, then the difference between the memory write data detection bandwidth and the theoretical memory write data bandwidth for the first, third, and fourth pseudo-computing cores is 80 GB / s - 80 GB / s = 0 GB / s. However, the difference between the memory write data detection bandwidth and the theoretical memory write data bandwidth for the second pseudo-computing core is 50 GB / s - 80 GB / s = -30 GB / s. This indicates that there is no difference between the memory write data detection bandwidth and the theoretical memory write data bandwidth for the first, third, and fourth pseudo-computing cores, while the difference for the second pseudo-computing core is 30 GB / s. GB / s indicates that the load on the written data of the network under test for each pseudo-computing core is unbalanced, and the load balancing of the written data of the network under test for each pseudo-computing core does not meet the corresponding design requirements.
[0150] In an illustrative embodiment, when the data transmission throughput detection data of the network-on-chip under test includes the total bandwidth of the memory read data detection of the network-on-chip under test, step 604 includes:
[0151] The difference between the total bandwidth detected by memory read data of the network under test and the preset theoretical total bandwidth of memory read data of the network under test is obtained.
[0152] In the specific application scenario example above, assume that the maximum transmission bandwidth for both the pseudo-computing core and the computing core for writing and reading data is 80 GB / s, and the maximum designed bandwidth for writing and reading data of the on-chip network under test is 3000 GB / s. Assuming that the currently measured total bandwidth for memory read data detection of the on-chip network under test is 2500 GB / s, then the difference between the total bandwidth for memory read data detection and the theoretical total bandwidth for memory write data is 2500 GB / s - 3000 GB / s = -500 GB / s. This indicates a difference of 500 GB / s, and the total bandwidth for memory read data detection does not meet the design requirement of the theoretical total bandwidth for memory read data.
[0153] In an illustrative embodiment, when the data transmission throughput detection data of the network-on-a-chip under test includes the memory read data detection bandwidth of any one of the pseudo-computing cores, step 604 includes:
[0154] Based on the memory read data detection bandwidth of any pseudo-computing core and the preset theoretical memory read data bandwidth of any pseudo-computing core, the difference between the memory read data detection bandwidth and the theoretical memory read data bandwidth of any pseudo-computing core is obtained.
[0155] In the specific application scenario example above, it is assumed that the maximum transmission bandwidth for writing and reading data for both the pseudo computing core and the computing core is 80 GB / s, and the maximum designed bandwidth for writing and reading data for the on-chip network under test is 3000 GB / s. Assuming that the actual read bandwidth (i.e., memory read data detection bandwidth) of one of the 64 pseudo-computing cores (the first, third, and fourth pseudo-computing cores) is 80 GB / s, and the actual read bandwidth of one of the 64 pseudo-computing cores (the second pseudo-computing core) is 50 GB / s, then the difference between the memory read data detection bandwidth and the theoretical memory read data bandwidth for the first, third, and fourth pseudo-computing cores is 80 GB / s - 80 GB / s = 0 GB / s. However, the difference between the memory read data detection bandwidth and the theoretical memory read data bandwidth for the second pseudo-computing core is 50 GB / s - 80 GB / s = -30 GB / s. This indicates that there is no difference between the memory read data detection bandwidth and the theoretical memory read data bandwidth for the first, third, and fourth pseudo-computing cores, while the difference for the second pseudo-computing core is 30 GB / s. GB / s indicates that the load of read data for each pseudo-computing core in the network under test is unbalanced, and the load balance of read data for each pseudo-computing core in the network under test does not meet the corresponding design requirements.
[0156] After drawing relevant conclusions and determining whether the corresponding design requirements are met, if the obtained data transmission throughput analysis results of the tested on-chip network indicate that the tested on-chip network fails to meet the corresponding design requirements, then relevant test data needs to be fed back for relevant designers to refer to and improve. Based on this, in the illustrative embodiment, after step 604, the on-chip network testing method of this disclosure embodiment may further include:
[0157] When the data transmission throughput analysis results show that the data transmission throughput test data of the network under test (NAT) does not meet the theoretical throughput requirements, the waveform data of the NAT during the access period of the dump memory is used.
[0158] Figure 7 This is a schematic diagram illustrating the steps of a specific application scenario of an on-chip network testing method according to an illustrative embodiment, such as... Figure 7 As shown, this application scenario mainly includes the following steps 701 to 711.
[0159] Step 701: Send memory access parameters to each pseudo-computing core in the on-chip network test system, and then execute step 702.
[0160] If the on-chip network testing system also includes computing cores, then in step 701, memory access parameters are also sent to each computing core.
[0161] Step 702: Each pseudo-computing core starts based on memory access parameters, and then step 703 is executed.
[0162] If the on-chip network testing system also includes computing cores, then step 702 further includes each computing core starting up based on memory access parameters.
[0163] Step 703: Each pseudo-computing core performs a memory access, during which steps 704 and 707 are executed.
[0164] Among them, memory access refers to writing data to or reading data from memory.
[0165] Step 704: Read the network traffic information on the chip under test, and then proceed to step 705.
[0166] Step 705: Calculate the actual access bandwidth of the network on the tested chip based on the traffic information on the network on the tested chip, and then proceed to step 706.
[0167] The actual access bandwidth of the network on the tested chip is either the total bandwidth for memory write data detection or the total bandwidth for memory read data detection.
[0168] Step 706: Determine whether the actual access bandwidth of the network under test is less than the preset theoretical access bandwidth of the network under test. If yes, proceed to step 710; otherwise, proceed to step 711.
[0169] Wherein, corresponding to memory access is writing data to memory or reading data from memory, the theoretical access bandwidth of the network under test is the theoretical total bandwidth of the network under test for writing data to memory or the theoretical total bandwidth of the network under test for reading data to memory.
[0170] Step 707: Read the traffic information of each pseudo-computing core, and then proceed to step 708.
[0171] If the on-chip network testing system also includes computing cores, then step 707 further includes reading the traffic information of each computing core.
[0172] Step 708: Calculate the actual access bandwidth of each pseudo-computing core based on the traffic information of each pseudo-computing core, and then proceed to step 709.
[0173] The actual access bandwidth of each pseudo-computing core is either the memory write data detection bandwidth or the memory read data detection bandwidth of each pseudo-computing core. If the on-chip network test system also includes computing cores, then step 708 further includes calculating the actual access bandwidth of each computing core based on the traffic information of each computing core, wherein the actual access bandwidth of each computing core is either the memory write data detection bandwidth or the memory read data detection bandwidth of each computing core.
[0174] Step 709: Determine whether the load distribution of memory access for each pseudo-computing core of the on-chip network under test is balanced. If so, proceed to step 710; otherwise, proceed to step 711.
[0175] The determination of whether the load distribution of memory access for each pseudo-computing core in the network under test is balanced can be achieved by judging the difference between the actual access bandwidth and the theoretical access bandwidth of each pseudo-computing core. Balanced load distribution means that the actual access bandwidth of each pseudo-computing core is equal or nearly equal. Therefore, the determination of whether the load distribution of memory access for each pseudo-computing core in the network under test is balanced can be achieved by directly judging whether the difference between the actual access bandwidth of each pseudo-computing core is within a preset range, thus achieving the judgment in step 709. For example, in step 709, the maximum and minimum values of the actual access bandwidth of each pseudo-computing core can be obtained, and it can be determined whether the difference between the maximum and minimum values is less than a preset difference range. This determines whether the load distribution of memory access for each pseudo-computing core by the network under test is balanced. If the difference between the maximum and minimum values is less than the preset difference range, the load distribution of memory access for each pseudo-computing core by the network under test is balanced (then proceed to step 710); otherwise, the load distribution of memory access for each pseudo-computing core by the network under test is unbalanced (then proceed to step 711).
[0176] If the on-chip network testing system also includes computing cores, then step 709 includes determining whether the load distribution of memory accesses by the on-chip network under test for each pseudo-computing core and the computing core is balanced.
[0177] Step 710: The test is passed and the test is completed.
[0178] Step 711: Dump the waveform data of the on-chip network under test during memory access and complete the test.
[0179] The on-chip network testing system and method of this disclosure replace the designed computing core with a pseudo-computing core. By using the pseudo-computing core to enter the on-chip network testing phase in advance, the on-chip network testing time can overlap with the design time of the computing core, thereby helping to shorten the SoC development cycle and improve SoC development efficiency. Simultaneously, because the pseudo-computing core does not have the computing power of the computing core, but only has the same memory access capabilities, it also helps to save the time required for the computing core to perform calculations, further shortening the on-chip network testing time and improving the testing efficiency.
[0180] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A network-on-chip test system, comprising: The network-on-chip test system comprises: a parameter configuration module configured to provide memory access parameters for a network-on-chip under test; at least one dummy compute core coupled to the parameter configuration module and coupled to a memory through the network-on-chip under test, and configured to access the memory through the network-on-chip under test based on the memory access parameters, wherein the dummy compute core has the same memory access function as a compute core but does not have the computing function of the compute core, and wherein the compute core is configured to perform computation and access the memory through the network-on-chip under test based on the memory access parameters; and a test data acquisition module coupled to the network-on-chip under test, and configured to obtain data transmission throughput detection data of the network-on-chip under test by detecting data transmission throughput of the network-on-chip under test during access of the memory by the at least one dummy compute core.
2. The network-on-chip test system of claim 1, wherein, The network-on-chip test system further comprises: at least one compute core coupled to the parameter configuration module and coupled to the memory through the network-on-chip under test; wherein the external transmission bandwidth of the at least one dummy compute core is the same as the external transmission bandwidth of the at least one compute core.
3. The network-on-chip test system of claim 1, wherein, The network-on-chip test system further comprises: a data analysis module coupled to the test data acquisition module, and configured to obtain data transmission throughput analysis results of the network-on-chip under test based on the data transmission throughput detection data of the network-on-chip under test and preset throughput theoretical data.
4. The network-on-chip test system of claim 3, wherein, The network-on-chip test system further comprises: a data dump module configured to dump waveform data of the network-on-chip under test during access of the memory in a case where the data transmission throughput analysis results indicate that the data transmission throughput detection data of the network-on-chip under test does not meet the throughput theoretical data requirements.
5. The network-on-chip test system of claim 1, wherein: the access of the memory by the at least one dummy compute core is writing data to the memory; the memory access parameters comprise: start parameters, abort parameters, write address parameters, write request burst transmission length parameters, write request length parameters, write request pen number parameters, write request step length parameters, and maximum number of uncompleted write request parameters.
6. The network-on-chip test system of claim 1, wherein: the access of the memory by the at least one dummy compute core is reading data from the memory; the memory access parameters comprise: start parameters, read address parameters, read request length parameters, read request pen number parameters, and read request step length parameters.
7. The network-on-chip test system of claim 1, wherein, The data transmission throughput detection data comprise at least one of a total memory data access detection bandwidth of the network-on-chip under test and a memory access detection bandwidth of each of the at least one dummy compute core.
8. The network-on-chip test system of claim 3, wherein: the data transmission throughput detection data comprise at least one of a total memory data access detection bandwidth of the network-on-chip under test and a memory access detection bandwidth of each of the at least one dummy compute core. The throughput theoretical data includes at least one of a memory data access detection total bandwidth of the measured NoC and a memory access theoretical bandwidth of each of the at least one pseudo-compute core. The data transmission throughput analysis result of the measured NoC includes at least one of a difference between the memory data access detection total bandwidth and the memory data access theoretical total bandwidth and a difference between a memory access detection bandwidth of any one pseudo-compute core and a memory access theoretical bandwidth of the any one pseudo-compute core. 9.A NoC testing method, comprising: configuring memory access parameters to at least one pseudo-compute core coupled to a measured NoC; accessing, by the at least one pseudo-compute core, a memory through the measured NoC based on the configured memory access parameters, wherein the memory is coupled to the measured NoC; obtaining data transmission throughput detection data of the measured NoC by detecting a data transmission throughput of the measured NoC during the accessing of the memory by the at least one pseudo-compute core. The pseudo-compute core has a same memory access function as a compute core and does not have a computing function of the compute core, wherein the compute core is configured to perform a computation and access the memory through the measured NoC based on the memory access parameters.
10. The network-on-chip testing method of claim 9, wherein, The configuring memory access parameters to at least one pseudo-compute core coupled to a measured NoC, comprises: configuring write parameters to the at least one pseudo-compute core or configuring read parameters to the at least one pseudo-compute core.
11. The network-on-chip testing method of claim 9, wherein, The accessing, by the at least one pseudo-compute core, a memory through the measured NoC based on the configured memory access parameters, comprises: in a case that the memory access parameters are write parameters, writing, by the at least one pseudo-compute core, data into the memory through the measured NoC based on the configured write parameters; in a case that the memory access parameters are read parameters, reading, by the at least one pseudo-compute core, data from the memory through the measured NoC based on the configured read parameters.
12. The network-on-chip testing method of claim 9, wherein, The obtaining data transmission throughput detection data of the measured NoC by detecting a data transmission throughput of the measured NoC during the accessing of the memory by the at least one pseudo-compute core, comprises: in a case that the at least one pseudo-compute core writes data into the memory, obtaining at least one of a memory write data detection total bandwidth of the measured NoC and a memory write data detection bandwidth of each of the at least one pseudo-compute core by detecting a data transmission throughput of the measured NoC during the writing of data into the memory by the at least one pseudo-compute core; in a case that the at least one pseudo-compute core reads data from the memory, obtaining at least one of a memory read data detection total bandwidth of the measured NoC and a memory read data detection bandwidth of each of the at least one pseudo-compute core by detecting a data transmission throughput of the measured NoC during the reading of data from the memory by the at least one pseudo-compute core.
13. The network-on-chip testing method of claim 9, wherein, After the obtaining data transmission throughput detection data of the measured NoC, the NoC testing method further comprises: According to the data transmission throughput detection data of the measured network-on-chip and preset throughput theoretical data, data transmission throughput analysis results of the measured network-on-chip are obtained.
14. The network-on-chip testing method of claim 13, wherein: the data transmission throughput detection data of the measured network-on-chip comprises at least one of memory write data detection total bandwidth of the measured network-on-chip, memory write data detection bandwidth of any one of the at least one pseudo computing core, memory read data detection total bandwidth of the measured network-on-chip, and memory read data detection bandwidth of any one of the at least one pseudo computing core; the throughput theoretical data comprises at least one of preset memory write data theoretical total bandwidth of the measured network-on-chip, memory write data theoretical bandwidth of any one of the at least one pseudo computing core, memory read data theoretical total bandwidth of the measured network-on-chip, and memory read data theoretical bandwidth of any one of the at least one pseudo computing core; the data transmission throughput analysis results of the measured network-on-chip comprise at least one of difference between the memory write data detection total bandwidth and the memory write data theoretical total bandwidth, difference between the memory write data detection bandwidth of any one of the at least one pseudo computing core and the memory write data theoretical bandwidth of any one of the at least one pseudo computing core, difference between the memory read data detection total bandwidth and the memory read data theoretical total bandwidth, and difference between the memory read data detection bandwidth of any one of the at least one pseudo computing core and the memory read data theoretical bandwidth of any one of the at least one pseudo computing core; in a case where the data transmission throughput detection data of the measured network-on-chip comprises the memory write data detection total bandwidth of the measured network-on-chip, the obtaining of the data transmission throughput analysis results of the measured network-on-chip according to the data transmission throughput detection data of the measured network-on-chip and the preset throughput theoretical data comprises: obtaining, according to the memory write data detection total bandwidth of the measured network-on-chip and preset memory write data theoretical total bandwidth of the measured network-on-chip, difference between the memory write data detection total bandwidth and the memory write data theoretical total bandwidth; in a case where the data transmission throughput detection data of the measured network-on-chip comprises the memory write data detection bandwidth of any one of the at least one pseudo computing core, the obtaining of the data transmission throughput analysis results of the measured network-on-chip according to the data transmission throughput detection data of the measured network-on-chip and the preset throughput theoretical data comprises: obtaining, according to the memory write data detection bandwidth of any one of the at least one pseudo computing core and preset memory write data theoretical bandwidth of any one of the at least one pseudo computing core, difference between the memory write data detection bandwidth of any one of the at least one pseudo computing core and the memory write data theoretical bandwidth of any one of the at least one pseudo computing core; in a case where the data transmission throughput detection data of the measured network-on-chip comprises the memory read data detection total bandwidth of the measured network-on-chip, the obtaining of the data transmission throughput analysis results of the measured network-on-chip according to the data transmission throughput detection data of the measured network-on-chip and the preset throughput theoretical data comprises: According to the memory read data detection total bandwidth of the measured NoC and the preset memory read data theoretical total bandwidth of the measured NoC, a difference between the memory read data detection total bandwidth and the memory read data theoretical total bandwidth is obtained; In a case where the data transmission throughput detection data of the measured NoC includes the memory read data detection bandwidth of any one of the at least one pseudo computing core, the data transmission throughput analysis result of the measured NoC is obtained according to the data transmission throughput detection data of the measured NoC and the preset throughput theoretical data, including: According to the memory read data detection bandwidth of the any one pseudo computing core and the preset memory read data theoretical bandwidth of the any one pseudo computing core, a difference between the memory read data detection bandwidth and the memory read data theoretical bandwidth of the any one pseudo computing core is obtained.
15. The network-on-chip testing method of claim 13, wherein, After the data transmission throughput analysis result of the measured NoC is obtained, the NoC testing method further includes: In a case where the data transmission throughput analysis result indicates that the data transmission throughput detection data of the measured NoC does not meet the requirement of the throughput theoretical data, the waveform data of the measured NoC during the access of the memory is dumped.
Citation Information
Patent Citations
Multi-core system based on network-on-chip and data transmission method
CN116610630A
Mobile device throughput testing
US20120300649A1