Data access delay evaluation system and method
By using pseudo computing cores to simulate the memory access function of computing cores in SoC systems, early testing and optimization of memory and on-chip networks are achieved, solving the problem of synchronization between testing and design in the SoC development cycle and improving development efficiency.
Patent Information
- Application Number
- CN202511296745.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-11
AI Technical Summary
In SoC systems, as the number of chiplets and computing cores increases, the data throughput pressure of on-chip networks and memories increases, making it difficult to synchronize the design and optimization of memories and on-chip networks during the SoC development cycle, affecting development efficiency.
A pseudo computing core is used to simulate the memory access function of the computing core, and the data access delay evaluation system is used for testing. The parameter configuration module, data access unit, delay data register group and delay evaluation module are used to evaluate the delay performance of the memory and on-chip network, realizing early testing and optimization.
It shortens the test time of memory and on-chip network, improves test efficiency, reduces the waiting time of computing core design, shortens the SoC development cycle, and improves development efficiency.
Smart Images

Figure CN120763115A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of chip testing technology, and in particular to a data access delay evaluation system and method. Background Art
[0002] A Network on Chip (NoC) is a communication network structure used within a System on Chip (SoC). In an SoC, integrated chiplets (or bare chips) read and write data with memory, such as High Bandwidth Memory (HBM), through the NoC. To improve memory read and write efficiency, the memory can be divided into two parts: main memory and secondary cache. Compared with main memory, secondary cache has a higher read and write speed. When a chiplet reads memory, it first checks whether the data to be read is stored in the secondary cache. If the data to be read is stored in the secondary cache (this is called a hit), the data is read directly from the secondary cache. If the data to be read is not stored in the secondary cache (this is called a miss), the data is read from the main memory.
[0003] With the development of SoC systems, the number of integrated chiplets (or bare dies) continues to increase, and the number of various computing cores in the chiplets is also increasing, which puts increasing pressure on the data throughput of on-chip networks and memories.
[0004] To ensure that the NOC and memory meet the design requirements for access to the compute core, they must be optimized and adjusted during the design phase. Therefore, how to complete the design and optimization of memory and NOC as early as possible during the design phase to help shorten the SoC design cycle has become a pressing issue. Summary of the Invention
[0005] In view of this, the present disclosure provides a data access delay evaluation system and method to help improve the test efficiency of memory and on-chip network, shorten the test time of memory and on-chip network, and help complete the test of memory and on-chip network before completing the design of the computing core in the SoC system, so that the test time of memory and on-chip network can overlap with the design time of the computing core, thereby helping to shorten the development cycle of SoC and help improve the development efficiency of SoC.
[0006] The technical solution of the present disclosure is achieved as follows: According to one aspect of an embodiment of the present disclosure, a data access delay evaluation system is provided, comprising: A parameter configuration module, used for providing memory access parameters during data access delay testing; at least two data access units coupled to the parameter configuration module and to a memory through a network-on-chip, configured to read and write a two-dimensional data matrix in the memory through the network-on-chip based on memory access parameters, wherein at least one of the at least two data access units is a dummy compute core, the dummy compute core having the same memory access function as a compute core, and the memory comprises a level two cache and a main memory; a delay data register set coupled to the at least two data access units, configured to obtain delay data of the at least two data access units during the access to the memory by the at least two data access units; and a delay evaluation module coupled to the delay data register set, configured to obtain a delay evaluation result according to the delay data of the at least two data access units and a preset delay design condition.
[0007] In a possible implementation, at least one of the at least two data access units is a compute core.
[0008] In a possible implementation, the data access delay evaluation system further comprises: a data dump module, configured to dump waveform data during the reading and writing of the two-dimensional data matrix in the memory in a case where the delay evaluation result indicates that the delay data does not meet the delay design condition requirement.
[0009] In a possible implementation, the memory access parameters comprise at least one of a start parameter, a stop parameter, a process identification parameter, a global process identification parameter, a read-write request arbitration weight parameter, a maximum number of uncompleted write requests parameter, a maximum number of uncompleted read requests parameter, a write request length parameter, a write request height parameter, a write request burst transmission length parameter, a write request step parameter, a write address parameter, a read request length parameter, a read request height parameter, a read request burst transmission length parameter, a read request step parameter, and a read address parameter.
[0010] In a possible implementation, the delay data comprises at least one of a number of read requests for statistical read delay, a total read request delay, a number of completed read requests, a number of completed write requests, and a total busy duration of the data access unit. The delay data register set comprises at least one of a read delay request number register, a read delay register, a completed read request number register, a completed write request number register, and a busy period register. The read delay request number register is configured to record the number of read requests for counting read delay, the read delay register is configured to record the total read delay of the read request, the completed read request number register is configured to record the number of completed read requests, the completed write request number register is configured to record the number of completed write requests, and the busy period register is configured to record the total busy time of the data access unit.
[0011] In a possible implementation, the delay evaluation module comprises: a calculation sub-module configured to calculate, according to delay data of the at least two data access units, a first delay of any one of the at least two data access units in a case of all cache miss of secondary cache data read and a second delay of the any one of the at least two data access units in a case of cache hit of secondary cache data read; a judgment sub-module configured to obtain the delay evaluation result according to the first delay, the second delay and the delay design condition.
[0012] In a possible implementation, the delay design condition is: the first delay of any one of the at least two data access units in a case of all cache miss of secondary cache data read is greater than the second delay of the any one of the at least two data access units in a case of cache hit of secondary cache data read, and a hit delay ratio obtained according to the first delay and the second delay is within a preset hit delay ratio interval.
[0013] According to another aspect of the embodiments of the present disclosure, a data access delay evaluation method is provided, comprising: configuring memory access parameters for at least two data access units coupled to a memory through a network on chip, wherein at least one of the at least two data access units is a pseudo computing core, the pseudo computing core has the same memory access function as a computing core, and the memory comprises a secondary cache and a main memory; reading and writing a two-dimensional data matrix in the memory by the at least two data access units based on the configured memory access parameters through the network on chip; obtaining delay data of the at least two data access units during the access of the at least two data access units to the memory; obtaining a delay evaluation result according to the delay data of the at least two data access units and a preset delay design condition.
[0014] In a possible implementation, the memory access parameter configures no overlap between the two-dimensional data matrices respectively read and written by different data access units, or the memory access parameter configures overlap between the two-dimensional data matrices respectively read and written by different data access units.
[0015] In a possible implementation, the delay evaluation result is obtained according to delay data of the at least two data access units and a preset delay design condition. The first delay of any one of the at least two data access units in a case of all misses in the secondary cache data read and the second delay of the any one of the at least two data access units in a case of a hit in the secondary cache data read are calculated according to the delay data of the at least two data access units. The delay evaluation result is obtained according to the first delay, the second delay, and the delay design condition.
[0016] As can be seen from the above scheme, in the data access delay evaluation system and method disclosed in the present invention, a pseudo computing core is used to implement the memory access function of the computing core. On the one hand, it is possible to enter the memory and on-chip network test phase in advance when the design of the computing core is not completed, so that the test time of the memory and on-chip network can overlap with the design time of the computing core, thereby helping to shorten the development cycle of the SoC and help improve the development efficiency of the SoC. On the other hand, because the pseudo computing core does not have the computing function of the computing core, but only has the same memory access function as the computing core, the use of the pseudo computing core also helps to save the time required for the computing core to perform calculations, thereby helping to shorten the test time of the memory and on-chip network and improve the test efficiency of the memory and on-chip network. At the same time, because in the related art, the SoC will set relevant registers for the memory access of the computing core, and then in the data access delay evaluation system and method disclosed in the present invention, the pseudo computing core can reuse these relevant registers required for the computing core to access the memory. Therefore, there is no need to design corresponding registers for accessing the memory for each pseudo computing core separately, which can save the design time of the relevant registers. Moreover, since the pseudo-computing core reuses these registers for accessing the memory, the pseudo-computing core can simulate the real computing core accessing the memory, so that the testers at each stage can simulate the real behavior of the computing core accessing the memory by configuring these registers without software programming, and then test the performance of the memory and the on-chip network. On this basis, the present disclosure can also allow the process control of multiple cores in the chip to be realized by configuring these registers, and realize the multi-core and multi-process operation in the chip, so that the real chip behavior can be simulated during the test. Therefore, it can be directly used to test the read and write performance of the on-chip bus and inter-chip bus between the multiple processes of the chip, reducing the complexity of software programming. In addition, the present disclosure can also allow the maximum outstanding request for reading and writing and the related read and write request arbitration weight to be controlled by registers to more accurately control the read and write traffic on the bus, which is more flexible and convenient, and closer to the actual scene. The present disclosure also allows the memory to be read and written by the two-dimensional data matrix directly through the configuration of the register. In the process of reading performance test for the memory, the overlap between different two-dimensional data matrices and the influence of the secondary cache are also taken into account, making the test scenario closer to the actual operation scenario of the artificial intelligence chip and making the test results more reliable. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a schematic diagram of the connection structure of a computing core, NoC and memory in a chip in the related art; Figure 2 is a structural diagram of a data access delay evaluation system according to an exemplary embodiment; Figure 3A2 is a schematic structural diagram of an embodiment in which all data access units in a data access delay evaluation system adopt pseudo computing cores; Figure 3B This is a schematic structural diagram of an embodiment in which a part of the data access units in the data access delay evaluation system adopt pseudo computing cores and another part of the data access units adopt computing cores; Figure 4A is a schematic diagram of a logical arrangement structure of a two-dimensional data matrix according to an exemplary embodiment; Figure 4B is a schematic diagram showing the distribution of a two-dimensional data matrix in a memory according to an exemplary embodiment; Figure 5 is a schematic structural diagram of a delay evaluation module according to an exemplary embodiment; Figure 6 is a flow chart illustrating a method for evaluating data access delay according to an exemplary embodiment; Figure 7A is a schematic diagram showing the relationship between a two-dimensional data matrix accessed by multiple data access units according to an exemplary embodiment; Figure 7B is a schematic diagram showing the relationship between another two-dimensional data matrix accessed by multiple data access units according to an exemplary embodiment; Figure 8 is a schematic diagram showing steps for obtaining a delay evaluation result according to an exemplary embodiment; Figure 9 This is a flowchart of an application scenario using the data access delay evaluation system and method of the embodiment of the present disclosure.
[0018] In the accompanying drawings, the names of the components represented by the reference numbers are as follows: 101. Small chips, 1011, computing core, 102. Network on Chip, 103. Memory, 1031, Second Level Cache, 1032. Main memory, 201. Parameter configuration module, 202, data access unit, 2021, pseudo computing core, 203, delay data register group, 204. Delay evaluation module, 2041, calculation submodule, 2042, judgment submodule, 205. Data dump module, 701. First data matrix, 702. Second data matrix, 703, the third data matrix, 704, the fourth data matrix, 710. Large two-dimensional data matrix, 711. First data access unit, 712. Second data access unit, 713. A third data access unit, 714. Fourth data access unit. DETAILED DESCRIPTION
[0019] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and examples.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0021] As used in the specification and claims of this disclosure, “coupled (or connected)” may refer to any direct or indirect connection means. For example, if a first device is coupled (or connected) to a second device, it should be interpreted that the first device can be directly connected to the second device, or the first device can be indirectly connected to the second device via other devices or some connection means.
[0022] Figure 1 This is a schematic diagram of the connection structure of the computing core, on-chip network and memory in a chip in the related technology. Figure 1 As shown, a chip, particularly a SoC chip, may include multiple chiplets 101, each of which includes multiple computing cores 1011. These computing cores 1011 are coupled to a memory 103 via an on-chip network 102. To improve the read and write speed of the memory 103, the memory 103 also includes a secondary cache 1031 and a main memory 1032. Because the secondary cache 1031 has a higher access speed, frequently used data in the main memory 1032 and data that has just been accessed from the memory 103 are usually backed up in the secondary cache 1031 for quick access by each chiplet 101.
[0023] During chip development, each component of the chip needs to be tested to ensure it meets design requirements. The complexity of chip structures leads to varying development schedules for each component. If the chip design is completed before testing the entire chip or individual components, and then rework or testing based on the test results, the design and testing phases will overlap in time. Consequently, the chip development cycle will be difficult to shorten due to the inability to synchronize development and testing. Therefore, interspersing testing of completed chip components during the design process, allowing testing and design to be performed simultaneously, can help shorten the chip development cycle.
[0024] refer to Figure 1 For the chip shown, during the chip development phase, if the memory 103 and the on-chip network 102 need to be tested, the participation of all computing cores 1011 coupled to the on-chip network 102 is required. In one possible scenario, the on-chip network 102 and the memory 103 have been designed, but the computing cores 1011 have not yet completed their design, making it difficult to test the on-chip network 102 and the memory 103. In another possible scenario, the on-chip network 102, the memory 103, and all the computing cores 1011 have been designed, and the on-chip network 102 and the memory 103 can be tested. However, during the test, the computing core 1011 needs to take up part of the test time because of the implementation of its own related computing functions. Therefore, the test of the on-chip network 102 and the memory 103 will be prolonged because it needs to wait for the completion of the computing tasks of the computing core 1011.
[0025] In view of this, the embodiments of the present disclosure provide a data access delay evaluation system and method to help improve the test efficiency of memory and on-chip network, shorten the test time of memory and on-chip network, and help complete the test of memory and on-chip network before completing the design of the computing core in the SoC system, so that the test time of memory and on-chip network can overlap with the design time of the computing core, thereby helping to shorten the development cycle of SoC and help improve the development efficiency of SoC.
[0026] Figure 2 This is a structural diagram of a data access delay evaluation system according to an exemplary embodiment. The data access delay evaluation system can be used to test data access to memory and on-chip network to obtain test data such as data transmission bandwidth of on-chip network and data read and write bandwidth of memory. Figure 2As shown, the data access delay evaluation system mainly includes a parameter configuration module 201, at least two data access units 202, a delay data register group 203, and a delay evaluation module 204. The parameter configuration module 201 is used to provide memory access parameters during data access delay testing. The at least two data access units 202 are coupled to the parameter configuration module 201 and coupled to the memory 103 via the on-chip network 102. They are used to read and write a two-dimensional data matrix in the memory 103 via the on-chip network 102 based on the memory access parameters. At least one of the at least two data access units 202 is a pseudo-computing core having the same memory access function as a computing core. The memory 103 includes a secondary cache 1031 and a main memory 1032. The delay data register group 203 is coupled to the at least two data access units 202 and is used to obtain delay data of the at least two data access units 202 during the period when the at least two data access units 202 access the memory 103. Delay evaluation module 204 is coupled to delay data register bank 203 and is configured to generate a delay evaluation result based on the delay data of at least two data access units 202 and predetermined delay design conditions. In a more specific embodiment, the pseudo-computing core does not have the computing functionality of a computing core, but only has the same memory access functionality as a computing core.
[0027] Among them, the delay design condition is a condition used to evaluate whether a system or device meets the design requirements. If the delay data meets the delay design condition, the system or device meets the design requirements. If the delay data does not meet the delay design condition, the system or device does not meet the design requirements. For example, in the embodiment of the present disclosure, if the on-chip network and memory are referred to as a storage device as a whole, the delay design condition can be a condition used to evaluate whether the storage device (including the on-chip network and memory) meets the design requirements. If the delay data meets the delay design condition, the storage device meets the design requirements. If the delay data does not meet the delay design condition, the storage device does not meet the design requirements. For example, in the embodiment of the present disclosure, assuming that the on-chip network has been previously tested or evaluated to show that it meets the design requirements, the delay design condition can be a condition used to evaluate whether the memory meets the design requirements. If the delay data meets the delay design condition, the memory meets the design requirements. If the delay data does not meet the delay design condition, the memory does not meet the design requirements.
[0028] In the data access delay evaluation system of the embodiment of the present disclosure, at least one of the at least two data access units 202 is a pseudo computing core, thereby realizing the use of the pseudo computing core to replace the computing core to realize the memory access function of the computing core. Therefore, on the one hand, when the design of the computing core is not completed, the pseudo computing core can be used to enter the test phase of the on-chip network 102 and the memory 103 in advance. On the other hand, because the pseudo computing core does not have the computing function of the computing core and only has the same memory access function as the computing core, the pseudo computing core can save the relevant circuit space of the computing function in the computing core and help save the time required for the computing core to perform calculations, thereby helping to save the software and hardware resources of the data access delay evaluation system and helping to shorten the test time of the on-chip network and memory, thereby improving the test efficiency of the on-chip network and memory.
[0029] In the data access delay evaluation system of the embodiment of the present disclosure, at least two data access units 202 read and write the two-dimensional data matrix in the memory 103 through the on-chip network 102 based on the memory access parameters, thereby helping to adapt to scenario requirements such as matrix operations (such as the operations of artificial intelligence chips). For example, in a matrix operation scenario, a large two-dimensional data matrix is stored in the memory 103, and each data access unit 202 imitates the corresponding computing core to obtain small two-dimensional data matrices in different areas of the large two-dimensional data matrix. These small two-dimensional data matrices may or may not overlap. Among them, one case where the small two-dimensional data matrices overlap is, for example, between two small two-dimensional data matrices, at least the last row of data of one small two-dimensional data matrix and at least the first row of data of the other small two-dimensional data matrix are data stored at the same address location in the memory 103, and the same address location needs to be accessed by at least two computing cores. When there is no overlap between small two-dimensional data matrices, when each data access unit 202 reads its own two-dimensional data matrix from the memory 103: if it is the first time to read from the memory 103, then because there is no corresponding backup data in the secondary cache 1031, each data access unit 202 reads from the main memory 1032. At this time, the delay of each data access unit 202 reading the main memory 1032 is obtained, and then whether the delay in this case meets the design requirements can be evaluated; if it is not the first time to read from the memory 103, then because there may be corresponding backup data in the secondary cache 1031, each data access unit 202 can read its own two-dimensional data matrix from the secondary cache 1031. At this time, the delay of each data access unit 202 reading the secondary cache 1031 is obtained, and then whether the delay in this case meets the design requirements can be evaluated.
[0030] In the case where there is overlap between small two-dimensional data matrices, when each data access unit 202 reads its own two-dimensional data matrix from the memory 103: if it is the first time to read from the memory 103, then because there is no corresponding backup data in the secondary cache 1031, each data access unit 202 reads from the main memory 1032, wherein, because the two-dimensional data matrix is read, and the reading of the two-dimensional data matrix is read row by row according to the rows of the matrix, at the same time, in order to improve the access speed of the memory 103, the read data will be backed up in the secondary cache 1031 for subsequent fast reading. In this case, when two data access units 202 read two overlapping small two-dimensional data matrices at the same time, the two data access units 202 will simultaneously read from the first row of their respective corresponding small two-dimensional data matrices. Data reading begins. Because at least the last row of data in one small two-dimensional data matrix (assuming it's matrix A) and at least the first row of data in another small two-dimensional data matrix (assuming it's matrix B) are stored at the same address in memory 103 (assuming this is called overlapping data), the data access unit 202 reading the B matrix will read the overlapping data before the data access unit 202 reading the A matrix. Furthermore, the data read from the B matrix will also be backed up in the secondary cache 1031. When the data access unit 202 reading the A matrix reads the overlapping data, it can directly read the overlapping data from the secondary cache 1031 without having to read it from the main memory 1032. This improves the data reading speed of the data access unit 202 of the A matrix and reduces data reading latency. At this point, the latency of each data access unit 202 reading the two-dimensional data matrices from memory 103 when there is overlap can be determined, allowing evaluation of whether the latency in this case meets the design requirements.
[0031] The above are examples of many test purposes that can be achieved when at least two data access units 202 read and write the two-dimensional data matrix in the memory 103 through the on-chip network 102. Other test purposes, such as delay testing when some data access units 202 perform data reading and other data access units 202 perform data writing, can also be tested by reading and writing the two-dimensional data matrix. At least two data access units 202 use the method of reading and writing the two-dimensional data matrix to achieve more test purposes.
[0032] Figure 3A FIG. 1 is a schematic diagram of an embodiment structure in which all data access units in the data access delay evaluation system adopt pseudo computing cores. Figure 3AAs shown, in the illustrative embodiment, the data access delay evaluation system mainly comprises a parameter configuration module 201, at least two pseudo computing cores 2021, a delay data register group 203, and a delay evaluation module 204. The parameter configuration module 201 is configured to provide memory access parameters during data access delay testing. The at least two pseudo computing cores 2021 are coupled to the parameter configuration module 201 and coupled to the memory 103 through the network on chip 102, and are configured to read and write a two-dimensional data matrix in the memory 103 through the network on chip 102 based on the memory access parameters, wherein the pseudo computing cores 2021 have the same memory access function as the computing cores, and the memory 103 comprises a secondary cache 1031 and a main memory 1032. The delay data register group 203 is coupled to the at least two pseudo computing cores 2021, and is configured to obtain delay data of the at least two pseudo computing cores 2021 during the at least two pseudo computing cores 2021 accessing the memory 103. The delay evaluation module 204 is coupled to the delay data register group 203, and is configured to obtain a delay evaluation result according to the delay data of the at least two pseudo computing cores 2021 and a preset delay design condition.
[0033] In the illustrative embodiment, in the case that part of the computing cores are completed in design and testing, the data access delay evaluation system of the embodiments of the present disclosure can further comprise the computing cores completed in design and testing. Figure 3B is an embodiment structure schematic diagram in the case that part of data access units in the data access delay evaluation system adopt pseudo computing cores and another part of data access units adopt computing cores, as shown in Figure 3BAs shown, the data access delay evaluation system mainly comprises a parameter configuration module 201, at least one pseudo-compute core 2021, at least one compute core 1011, a delay data register group 203, and a delay evaluation module 204. The parameter configuration module 201 is configured to provide memory access parameters during data access delay testing. The at least one pseudo-compute core 2021 and the at least one compute core 1011 are coupled to the parameter configuration module 201 and coupled to the memory 103 through the network on chip 102, and are configured to read and write a two-dimensional data matrix in the memory 103 through the network on chip 102 based on the memory access parameters, wherein the pseudo-compute core 2021 has the same memory access function as the compute core 1011, and the memory 103 comprises a secondary cache 1031 and a main memory 1032. The delay data register group 203 is coupled to the at least one pseudo-compute core 2021 and the at least one compute core 1011, and is configured to obtain delay data of the at least one pseudo-compute core 2021 and the at least one compute core 1011 during access to the memory 103 by the at least one pseudo-compute core 2021 and the at least one compute core 1011. The delay evaluation module 204 is coupled to the delay data register group 203, and is configured to obtain a delay evaluation result according to the delay data of the at least two pseudo-compute cores 2021 and a preset delay design condition. In an illustrative embodiment, the compute core 1011 is configured to perform computation in addition to reading and writing the two-dimensional data matrix in the memory 103 through the network on chip 102 based on the memory access parameters. In an illustrative embodiment, the transmission bandwidth of the at least one pseudo-compute core 2021 is the same as that of the at least one compute core 1011.
[0034] In an illustrative embodiment, as shown in Figure 2 、 Figure 3A all the data access units 202 can be pseudo-compute cores 2021; in an illustrative embodiment, as shown in Figure 2 、 Figure 3B some of the data access units 202 can be pseudo-compute cores 2021 and the others can be compute cores 1011; in an illustrative embodiment, as shown in Figure 2 、 Figure 3B in the case where some of the data access units 202 are pseudo-compute cores 2021 and the others are compute cores 1011, the number of pseudo-compute cores 2021 is greater than that of compute cores 1011, for example, as shown in Figure 3B the data access delay evaluation system can comprise at least two pseudo-compute cores 2021 and at least one compute core 1011, and the number of pseudo-compute cores 2021 is greater than that of compute cores 1011.
[0035] In an exemplary embodiment, a computing core may be at least one of a vector core and a tensor core. For example, when there is only one computing core, the computing core may be a vector core or a tensor core. When there are more than one computing core, all computing cores may be vector cores, all computing cores may be tensor cores, or some computing cores may be vector cores and others may be tensor cores.
[0036] Figure 3B The computing core 1011 is introduced into the data access delay evaluation system of the illustrated embodiment, so that the testing process of the on-chip network 102 and the memory 103 can be closer to the actual operation scenario of the chip, and the obtained delay data may be more reliable.
[0037] against Figure 3B In the data access delay evaluation system of the illustrated embodiment, to minimize data access delay testing time and improve data access delay testing efficiency, in the exemplary embodiment, the number of pseudo computing cores 2021 is greater than the number of computing cores 1011. Because the number of pseudo computing cores 2021 is greater than the number of computing cores 1011, this helps reduce the proportion of time spent by computing cores 1011 executing computing tasks during the test, thereby helping to achieve a better balance between the reliability of delay data and improved testing efficiency.
[0038] like Figure 2 、 Figure 3A 、 Figure 3B As shown, in the exemplary embodiment, in order to facilitate modification and optimization of the on-chip network 102 and the memory 103 based on the delay evaluation results, in the exemplary embodiment, the data access delay evaluation system of the present disclosure further includes a data dump module 205. The data dump module 205 is configured to dump waveform data during the reading and writing of the two-dimensional data matrix in the memory 103 when the delay evaluation results indicate that the delay data does not meet the delay design condition requirements.
[0039] In an illustrative embodiment, the memory access parameters include at least one of a start parameter, an abort parameter, a maximum number of outstanding write requests parameter, a maximum number of outstanding read requests parameter, a write request length parameter, a write request height parameter, a write request burst transfer length parameter, a write request step parameter, a write address parameter, a read request length parameter, a read request height parameter, a read request burst transfer length parameter, a read request step parameter, and a read address parameter.
[0040] In the hardware system of the chip, the control of pseudo-computing cores and computing cores is usually implemented using registers. Therefore, corresponding to the memory access parameters, the configured registers include at least one of the start register (cfg_start), the end register (cfg_end), the maximum number of outstanding write requests register (write_outstanding), the maximum number of outstanding read requests register (read_outstanding), the write request length register (write_length), the write request height register (write_height), the write request burst transfer length register (write_burst_length), the write request step register (write_stride), the write address register (write_address), the read request length register (read_length), the read request height register (read_height), the read request burst transfer length register (read_burst_length), the read request step register (read_stride), and the read address register (read_address).
[0041] In the illustrative embodiment, the start register is configured to 1 to indicate to start issuing write requests and write data and to receive returned write return signals; the stop register is configured to 1 to indicate to immediately stop issuing write requests, and thus, based on the stop register, a tester can be allowed to stop the pseudo-compute core from issuing write requests and data in the middle of the process; the maximum outstanding write request number register is configured to allow a maximum number of outstanding write requests, i.e., when the number of write requests that have not received write return signals is equal to the maximum outstanding write request number, issuing of write requests is suspended to limit write data traffic on the bus; the maximum outstanding read request number register is configured to allow a maximum number of outstanding read requests, i.e., when the number of read requests that have not received returned data is equal to the maximum outstanding read request number, issuing of read requests is suspended to limit read data traffic on the bus; the write request length register is configured to the length of the main memory addresses accessed continuously by a write request, i.e., the length of each row in a two-dimensional data matrix and the total number of burst requests for each row; the write request height register is configured to the height of the two-dimensional data matrix of write requests, i.e., the number of rows in the two-dimensional data matrix; the write request burst length register is configured to the burst length of each write request issued, i.e., the length of the data corresponding to each write request, in the illustrative embodiment, burst_length = 0 represents that the length of the data corresponding to each write request is 128 bytes, burst_length = 1 represents that the length of the data corresponding to each write request is 256 bytes, and burst_length = 3 represents that the length of the data corresponding to each write request is 512 bytes; the write request step register is configured to the interval step length between the continuous main memory addresses of each two write requests issued, i.e., the interval step length between each two rows of data in the two-dimensional data matrix in the main memory; the write address register is configured to the starting main memory address of the write request issued; the read request length register is configured to the length of the main memory addresses accessed continuously by a read request, i.e., the length of each row in a two-dimensional data matrix and the total number of burst requests for each row; the read request height register is configured to the height of the two-dimensional data matrix of read requests, i.e., the number of rows in the two-dimensional data matrix; the read request burst length register is configured to the burst length of each read request issued, i.e., the length of the data corresponding to each read request, in the illustrative embodiment, burst_length = 0 represents that the length of the data corresponding to each read request is 128 bytes, burst_length = 1 represents that the length of the data corresponding to each write request is 256 bytes, and burst_length = 3 represents that the length of the data corresponding to each write request is 512 bytes; the read request step register is configured to the interval step length between the continuous main memory addresses of each two read requests issued, i.e., the interval step length between each two rows of data in the two-dimensional data matrix in the main memory; and the read address register is configured to the starting main memory address of the read request issued.
[0042] In an illustrative embodiment, in order to facilitate the management of the process and the control of the read-write weight, the memory access parameter can further include at least one of a process identification parameter, a global process identification parameter, and a read-write request arbitration weight parameter. Corresponding to the process identification parameter, the global process identification parameter, and the read-write request arbitration weight parameter, the configured register can further include at least one of a process identification register (context_id), a global process identification register (global_context_id), and a read-write request arbitration weight register (arb_weight).
[0043] The process identification register is configured to run the process of the computing core simulated by the pseudo computing core; the global process identification register is configured to run the process of the whole chip corresponding to the process run by the computing core simulated by the current pseudo computing core, and the process identification parameter and the global process identification parameter, i.e., the information configured by the process identification register and the global process identification register, represent the mapping relationship between the in-core process and the chip process.
[0044] The read-write request arbitration weight register is configured to be the arbitration weight when the read-write request is simultaneously issued. In an illustrative embodiment, the read-write request arbitration weight register is configured to 1, indicating that the weight of the write request is 1:1 compared with the read request, the read-write request arbitration weight register is configured to 2, indicating that the weight of the write request is 1:2 compared with the read request, and the read-write request arbitration weight register is configured to 3, indicating that the weight of the write request is 1:3 compared with the read request. Generally, in the SoC, the read request and the write request of the computing core to the bus (for example, UCIe (Universal Chiplet Interconnect Express)) cannot be parallel, but the data is transmitted through other channels other than the read request and the write request, so the transmission of the data can be parallel with the access request. Generally, the access on the bus is more read than write, so in an illustrative embodiment, the proportion of the read-write request of the pseudo computing core can be configured through the read-write request arbitration weight register, so that it is more consistent with the proportion relationship of the read-write request on the actual bus, and the data access of the pseudo computing core is closer to the actual application scenario.
[0045] Through the control of each register on each data access unit 202, the access to the memory 103 according to the demand can be realized.
[0046] Figure 4A is a schematic diagram of the logical arrangement structure of a two-dimensional data matrix according to an illustrative embodiment, Figure 4B is a schematic diagram of the distribution of a two-dimensional data matrix in a memory according to an illustrative embodiment. As Figure 4AAs shown, the logical arrangement of the two-dimensional data matrix is in the form of rows and columns, combined with the access of any data access unit 202 to a two-dimensional data matrix in the present disclosure, Figure 4A and Figure 4B In Chinese: If it is a write request, then each burst represents the burst transfer length of each write request represented by the write request burst transfer length parameter, height represents the height of the two-dimensional data matrix represented by the write request height parameter, that is, the number of rows in the two-dimensional data matrix, length represents the main memory address length of the write request continuously accessed by the write request represented by the write request length parameter, that is, the length of each row in the two-dimensional data matrix, and the sum of the burst transfer lengths of the write requests in each row, address represents the starting main memory address of the write request represented by the write address parameter, stride represents the interval step between each two consecutive main memory addresses of the write request represented by the write request step parameter, that is, the interval step between each two rows of data in the two-dimensional data matrix in the main memory ; If it is a read request, then each burst represents the burst transfer length of each read request represented by the read request burst transfer length parameter, height represents the height of the two-dimensional data matrix represented by the read request height parameter, that is, the number of rows of the two-dimensional data matrix, length represents the main memory address length of continuous access by the read request represented by the read request length parameter, that is, the length of each row in the two-dimensional data matrix, and the sum of the burst transfer lengths of the read requests in each row, address represents the starting main memory address of the read request represented by the read address parameter, stride represents the interval step between each two consecutive main memory addresses of the read request represented by the read request step parameter, that is, the interval step between each two rows of data in the two-dimensional data matrix in the main memory.
[0047] like Figure 4A 、 Figure 4B As shown in the figure, in terms of the logical arrangement structure, the two-dimensional data matrix is arranged in the form of rows and columns, but in terms of the actual storage location in the memory, the two-dimensional data matrix is stored continuously in the main memory space, and the data of each row in the two-dimensional data matrix are separated in the main memory space by the continuous main memory address space represented by the write request step parameter represented by stride or the read request step parameter.
[0048] like Figure 4A 、 Figure 4BAs shown, in the illustrative embodiment, during the write access process, when the start register of any one data access unit 202 is read as 1, the write request is issued through the network on a chip 102 to the memory 103 from the start memory 1032 address recorded in the write address register to start writing a two-dimensional data matrix into the memory 103, each piece of write request (including burst transmission of multiple write requests) continuously accesses the memory address length (i.e. the length of each row in the two-dimensional data matrix, the total amount of burst transmission request for each row write) specified by the write request length register, the data length corresponding to each write request is specified by the write request burst transmission length register, the interval step length between the continuous memory addresses of each two pieces of write request is specified by the write request step length register, and the total number of rows of write requests issued is specified by the write request height register. After receiving the write return signal returned by the memory 103, the operation is stopped.
[0049] During the write access process, when the stop register of any one data access unit 202 is read as 1, the write request is immediately stopped, the write data continues to be issued until the number of write requests is stopped, and after all the return signals of the issued write requests are received, the behavior of the arbitrary data access unit 202 is stopped. The write data corresponding to each write request is aligned with the write request, and the write data is not issued before the corresponding write request, so that the write data issued when stopped is not more than the write request. In the illustrative embodiment, an unfinished write request number counter is arranged in the pseudo-computing core to count the number of unfinished write requests. When the value of the unfinished write request number counter is greater than or equal to the configuration value of the maximum unfinished write request number register, the write request is stopped, and only when the value of the unfinished write request number counter is less than the configuration value of the maximum unfinished write request number register, the write request can be continuously issued, so as to realize the flow control when writing data into the memory 103. In the illustrative embodiment, the issuance of the write request and the write data by the data access unit 202 and the reception of the write return are completed based on the communication protocol designed in the SoC chip, so that the write return signal is not back-pressured to the bus.
[0050] As Figure 4A , Figure 4BAs shown, in the illustrative embodiment, during the read access process, when the start register of any one data access unit 202 is read to be 1, the read request is issued through the network-on-chip 102 to the memory 103 from the start main memory address recorded in the read address register to start reading a two-dimensional data matrix from the memory 103, each piece of read request (including burst transmission of multiple read requests) continuously accesses the main memory address length (i.e. the length of each row in the two-dimensional data matrix, the total amount of burst transmission request for each row read) specified by the read request length register, the data length corresponding to each read request is the length specified by the read request burst transmission length register, the interval step length between the continuous main memory addresses of each two pieces of read request is the interval step length specified by the read request step length register, and the total number of rows of read requests issued is the number of rows specified by the read request height register, and after receiving the read return data returned by the memory 103, the operation is stopped.
[0051] During the read access process, when the abort register of any one data access unit 202 is read to be 1, the read request is immediately stopped, and after all read return data is collected, all behaviors of the arbitrary data access unit 202 are stopped. In the illustrative embodiment, an unfinished read request number counter is arranged in the pseudo-computing core to count the number of unfinished read requests, and when the value of the unfinished read request number counter is greater than or equal to the configuration value of the maximum unfinished read request number register, the read request is stopped, and only when the value of the unfinished read request number counter is less than the configuration value of the maximum unfinished read request number register, the read request can be continuously issued, thereby realizing the flow control when reading data from the memory 103. In the illustrative embodiment, the data access unit 202 issues the read request and receives the read data based on the communication protocol designed in the SoC chip, and thus the bus is not back-pressured when receiving the read data.
[0052] In the illustrative embodiment, in order to facilitate the intuitiveness of the delay evaluation, the delay data can include the read request number for the statistical read delay, the read request total delay, the completed read request number, the completed write request number, the total busy time length of the data access unit 202, and the delay data register set 203 can include at least one of the read delay request number register (pfc_read_latecy_req_cnt), the read delay register (pfc_read_latecy_cnt), the completed read request number register (read_req_done_cnt), the completed write request number register (write_req_done_cnt), and the busy cycle register (busy_cycle_cnt) with respect to each delay data. The read delay request number register is used to record the read request number for the statistical read delay, the read delay register is used to record the read request total delay, the completed read request number register is used to record the completed read request number, the completed write request number register is used to record the completed write request number, and the busy cycle register is used to record the total busy time length of the data access unit. The average delay of the read request completed by the current data access unit can be intuitively seen through the read delay register and the read delay request number register. The read / write request number completed by the current data access unit can be intuitively seen through the completed read request number register and the completed write request number register, and the running time of the current data access unit can be intuitively seen through the busy cycle register, which facilitates the tester to understand the working time and state of the data access unit (pseudo-computing core, computing core) corresponding to the busy cycle register.
[0053] The access of the memory 103 mainly includes the read operation of the data stored in the memory 103 and the write operation of the data into the memory 103. For the read operation, because of the existence of the secondary cache 1031, the data can be obtained from the secondary cache 1031, which can help to improve the data read speed, and for the write operation, the data is usually written into the main memory 1032, which may not need the participation of the secondary cache 1031, so the process of the read operation is more complex than that of the write operation. Therefore, the evaluation condition of the read operation of the memory 103 can be different or more complex than that of the write operation, because the read operation of the memory 103 includes whether the data read of the secondary cache 1031 hits, and generally, the data read of the secondary cache 1031 hit means the improvement of the data read speed and the shortening of the delay. The following mainly combines the read data to explain the delay evaluation.
[0054] Figure 5 is a structural schematic diagram of a delay evaluation module according to an illustrative embodiment, as Figure 5In the illustrative embodiment, the delay evaluation module 204 includes a calculation submodule 2041 and a judgment submodule 2042. The calculation submodule 2041 is configured to calculate, according to the delay data of the at least two data access units 202, a first delay of any one of the at least two data access units 202 in a case of all misses in the data read of the second level cache 1031 and a second delay of the any one of the at least two data access units 202 in a case of a hit in the data read of the second level cache 1031. The judgment submodule 2042 is configured to obtain a delay evaluation result according to the first delay, the second delay and a delay design condition.
[0055] In the illustrative embodiment, for the write data, since there is no case of a hit or miss in the second level cache, the calculation submodule 2041 is configured to calculate, according to the delay data of the at least two data access units 202, a write data delay of any one of the at least two data access units 202. The judgment submodule 2042 is configured to obtain a delay evaluation result for the write data according to a comparison result of the write data delay and a preset write data delay threshold.
[0056] In the illustrative embodiment, the delay design condition is that the first delay of any one of the at least two data access units 202 in a case of all misses in the data read of the second level cache is greater than the second delay of the any one of the at least two data access units 202 in a case of a hit in the data read of the second level cache, and a hit delay ratio obtained according to the first delay and the second delay is within a preset hit delay ratio range. The delay design condition is expressed in a formula as follows: T miss > T hit (1) (T miss - T hit ) / T miss ∈ [LR Hit ratio range] (2) wherein T miss is the first delay, T hit is the second delay, and [LR Hit ratio range] is the hit delay ratio range, wherein [LR Hit ratio range] is associated with a hit ratio, and the hit ratio refers to a hit rate of a cache block in the second level cache. The relationship between [LR Hit ratio range] and the hit ratio can be obtained through relevant tests. The delay design condition is satisfied when both the formula (1) and the formula (2) are satisfied, and the delay design condition is not satisfied when either of the formula (1) and the formula (2) is not satisfied.
[0057] In the illustrative embodiment, according to different designs, at least one of the data access unit 202, the pseudo computing core 2021, the computing core 1011, the on-chip network 102, the memory 103, the parameter configuration module 201, the delay data register group 203, the delay evaluation module 204, the computing sub-module 2041, the judgment sub-module 2042, and the data dump module 205 can be implemented in a combination of hardware, firmware, and software (i.e., program).
[0058] In terms of hardware, at least one of the data access unit 202, the pseudo computing core 2021, the computing core 1011, the on-chip network 102, the memory 103, the parameter configuration module 201, the delay data register group 203, the delay evaluation module 204, the computing submodule 2041, the judgment submodule 2042, and the data dump module 205 can be implemented as a logic circuit on an integrated circuit. For example, the relevant functions of at least one of the data access unit 202, the pseudo computing core 2021, the computing core 1011, the on-chip network 102, the memory 103, the parameter configuration module 201, the delay data register group 203, the delay evaluation module 204, the computing submodule 2041, the judgment submodule 2042, and the data dump module 205 can be implemented in one or more hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), or other similar devices. Various logic blocks, modules, and circuits in a processor (DSP), a field programmable gate array (FPGA), a central processing unit (CPU), or other processing units. The functions associated with at least one of the data access unit 202, the pseudo computing core 2021, the computing core 1011, the on-chip network 102, the memory 103, the parameter configuration module 201, the delay data register group 203, the delay evaluation module 204, the computing submodule 2041, the judgment submodule 2042, and the data dump module 205 can be implemented as hardware circuits using hardware description languages (such as Verilog HDL or VHDL) or other suitable programming languages, such as various logic blocks, modules, and circuits in an integrated circuit.
[0059] In software or firmware form, the relevant functions of at least one of the data access unit 202, the dummy computing core 2021, the computing core 1011, the on-chip network 102, the memory 103, the parameter configuration module 201, the delay data register group 203, the delay evaluation module 204, the computing submodule 2041, the judgment submodule 2042, and the data dump module 205 can be implemented as programming codes. For example, at least one of the data access unit 202, the dummy computing core 2021, the computing core 1011, the on-chip network 102, the memory 103, the parameter configuration module 201, the delay data register group 203, the delay evaluation module 204, the computing submodule 2041, the judgment submodule 2042, and the data dump module 205 can be implemented using a general programming language (e.g., C, C++, or assembly language) or other suitable programming language. The programming code can be recorded and stored in a "non-transitory machine-readable storage medium." In some embodiments, the non-transitory machine-readable storage medium includes, for example, a semiconductor memory and / or a storage device. An electronic device (e.g., a CPU, hardware controller, microcontroller, hardware processor, or microprocessor) can read and execute the programming code from the non-transitory machine-readable storage medium, thereby implementing the relevant functions of at least one of the data access unit 202, pseudo computing core 2021, computing core 1011, on-chip network 102, memory 103, parameter configuration module 201, delay data register group 203, delay evaluation module 204, computing submodule 2041, judgment submodule 2042, and data dump module 205.
[0060] In an illustrative embodiment, the data access delay evaluation system of the embodiment of the present disclosure is applicable to SoC chips, etc., wherein the SoC chip can be any one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit).
[0061] Figure 6 is a flow chart of a data access delay evaluation method according to an exemplary embodiment. The data access delay evaluation method can be applied to the data access delay evaluation system of any of the above embodiments, such as Figure 6 As shown, in the exemplary embodiment, the data access delay evaluation method mainly includes the following steps 601 to 604 .
[0062] Step 601: Configure memory access parameters for at least two data access units coupled to the memory through the on-chip network, wherein at least one of the at least two data access units is a pseudo computing core, and the pseudo computing core has the same memory access function as the computing core, and the memory includes a secondary cache and a main memory.
[0063] Step 602: At least two data access units read and write a two-dimensional data matrix in a memory through an on-chip network based on configured memory access parameters.
[0064] Step 603: Acquire delay data of the at least two data access units during the period when the at least two data access units access the memory.
[0065] Step 604: Obtain a delay evaluation result according to the delay data of at least two data access units and a preset delay design condition.
[0066] In the exemplary embodiment, the two-dimensional data matrices read and written by data access units with different memory access parameter configurations do not overlap. This approach allows each data access unit to read and write a specific memory region without interfering with each other. This allows for data access latency testing without interfering with each other's access addresses, eliminating delays caused by access address conflicts between data access units. This can be used to test whether the on-chip network and / or memory's bandwidth allocation to each data access unit, as well as latency balance, meets design requirements.
[0067] Figure 7A FIG. 1 is a schematic diagram showing the relationship between a two-dimensional data matrix accessed by multiple data access units according to an exemplary embodiment. Figure 7A As shown, this embodiment includes four two-dimensional data matrices accessed by four data access units, namely a first data matrix 701, a second data matrix 702, a third data matrix 703, and a fourth data matrix 704. The first data matrix 701 is accessed by a first data access unit 711, the second data matrix 702 is accessed by a second data access unit 712, the third data matrix 703 is accessed by a third data access unit 713, and the fourth data matrix 704 is accessed by a fourth data access unit 714. In the exemplary embodiment, the first data access unit 711 may be, for example, a real computing core, and the second data access unit 712, the third data access unit 713, and the fourth data access unit 714 may be, for example, pseudo computing cores. Figure 7A As shown, the four two-dimensional data matrices may be part of a larger two-dimensional data matrix 710 . Figure 7A In the illustrated embodiment, there is no overlap between the first data matrix 701 , the second data matrix 702 , the third data matrix 703 , and the fourth data matrix 704 .
[0068] In an exemplary embodiment, the two-dimensional data matrices read and written by the data access units configured with memory access parameters overlap. This overlap causes the memory read and write operations of each data access unit to cause memory management operations for access resources and process coordination scheduling, resulting in corresponding access delays. Furthermore, the presence of the L2 cache during read operations also reduces read latency to a certain extent, allowing for testing whether the memory meets design requirements.
[0069] Figure 7B FIG. 1 is a schematic diagram showing the relationship between a two-dimensional data matrix accessed by multiple data access units according to an exemplary embodiment. Figure 7BAs shown, this embodiment includes four two-dimensional data matrices accessed by four data access units, namely a first data matrix 701, a second data matrix 702, a third data matrix 703, and a fourth data matrix 704. The first data matrix 701 is accessed by a first data access unit 711, the second data matrix 702 is accessed by a second data access unit 712, the third data matrix 703 is accessed by a third data access unit 713, and the fourth data matrix 704 is accessed by a fourth data access unit 714. In the exemplary embodiment, the first data access unit 711 may be, for example, a real computing core, and the second data access unit 712, the third data access unit 713, and the fourth data access unit 714 may be, for example, pseudo computing cores. Figure 7B As shown, the four two-dimensional data matrices can be part of a large two-dimensional data matrix 710. Figure 7A The embodiment shown is different in that Figure 7BIn the shown embodiment, there is overlap between the first data matrix 701, the second data matrix 702, the third data matrix 703, and the fourth data matrix 704. For example, the last row of data of the first data matrix 701 is the first row of data of the second data matrix 702, the last row of data of the second data matrix 702 is the first row of data of the third data matrix 703, and the last row of data of the third data matrix 703 is the first row of data of the fourth data matrix 704. Thus, the data accessed by the first data access unit 711, the second data access unit 712, the third data access unit 713, and the fourth data access unit 714 are correlated, which in turn affects the respective access delays.For example, when the first data access unit 711, the second data access unit 712, the third data access unit 713, and the fourth data access unit 714 simultaneously read the first data matrix 701, the second data matrix 702, the third data matrix 703, and the fourth data matrix 704 for the first time, the four data access units each read their corresponding two-dimensional data matrix simultaneously, that is, at the same moment, the first data access unit 711 reads the first row of data of the first data matrix 701, and the second data access unit 712 reads the first row of data of the second data matrix 702, and the third data access unit 713 reads the first row of data of the third data matrix 703, and the fourth data access unit 714 reads the fourth data matrix The first row of data of matrix 704 is stored in the main memory because it is the first time to read, and there is no corresponding copy backup in the secondary cache. After the first read, the first row of data of the first data matrix 701, the first row of data of the second data matrix 702, the first row of data of the third data matrix 703, and the first row of data of the fourth data matrix 704 are all backed up in the secondary cache. After reading the first row of data of their respective two-dimensional data matrices, the four data access units each simultaneously read the second row of data of their respective corresponding two-dimensional data matrices. Similarly, when the four access units each simultaneously read the last row of data of their respective two-dimensional data matrices, because the last row of data of the first data matrix 701 is the first row of data of the fourth data matrix The first row of data of the second data matrix 702 has been backed up in the secondary cache, so when the first data access unit 711 reads the last row of data of the first data matrix 701, it only needs to read from the secondary cache without having to read from the main memory. Similarly, when the second data access unit 712 reads the last row of data of the second data matrix 702, it only needs to read from the secondary cache without having to read from the main memory. When the third data access unit 713 reads the last row of data of the third data matrix 703, it only needs to read from the secondary cache without having to read from the main memory. However, when the fourth data access unit 714 reads the last row of data of the fourth data matrix 704, it still needs to read from the main memory. It can be seen that the first data matrix 701 and the second data matrix 702 have the same data structure as the first data matrix 701. 02. This overlapping manner between the third data matrix 703 and the fourth data matrix 704 can, to a certain extent, improve the speed at which the first data access unit 711, the second data access unit 712, and the third data access unit 713 initially read their respective two-dimensional data matrices. Since each row of data of the fourth data access unit 714 needs to be read from the main memory, the first data access unit 711, the second data access unit 712, and the third data access unit 713 initially read their respective two-dimensional data matrices faster than the fourth data access unit 714 initially reads the fourth data matrix 704, and the data reading latency is lower. However, because not all data is retrieved from the secondary cache, the latency is not minimized.In this case, the delay design condition can be set to perform delay evaluation.
[0070] Figure 8 Fig. 6 is a schematic diagram of steps of obtaining a delay evaluation result according to an illustrative embodiment, as shown in Fig. 6, for the above case, in an illustrative embodiment, step 604 can include steps 801-802. Figure 8
[0071] Step 801: according to the delay data of the at least two data access units, a first delay of any one of the at least two data access units in the case of all cache miss in the second level cache data read and a second delay of the any one of the at least two data access units in the case of cache hit in the second level cache data read are calculated.
[0072] Step 802: according to the first delay, the second delay and the delay design condition, a delay evaluation result is obtained.
[0073] In an illustrative embodiment, the delay design condition is that the first delay of any one of the at least two data access units in the case of all cache miss in the second level cache data read is greater than the second delay of the any one of the at least two data access units in the case of cache hit in the second level cache data read, and a hit delay ratio obtained according to the first delay and the second delay is within a preset hit delay ratio interval. The formula expression of the delay design condition is: T miss > T hit (1) (T miss - T hit ) / T miss ∈ [LR Hit ratio range] (2) Wherein, T miss is the first delay, T hit is the second delay, and [LR Hit ratio range] is the hit delay ratio interval, wherein [LR Hit ratio range] is associated with Hit ratio, and Hit ratio refers to the hit rate of the cache block in the second level cache, and the relationship between [LR Hit ratio range] and Hit ratio can be obtained through relevant tests. Meanwhile satisfying the above formula (1) and formula (2) is to satisfy the delay design condition, and not satisfying any one of formula (1) and formula (2) is to not satisfy the delay design condition.
[0074] In one specific application scenario of the data access delay evaluation system and method of the embodiments of the present disclosure, a network topology can be deployed on a hardware simulator, a real computing core and three pseudo computing cores are connected with the second level cache and the main memory through the network on chip. The computing core and the pseudo computing core perform read of the same size of two-dimensional data matrix. Figure 7A , Figure 7B For example, the large two-dimensional data matrix 710 stored in the main memory of the memory is 60×40, that is, 60 rows and 40 columns. In the exemplary embodiment, combined with Figure 4A As shown, 40 columns can represent 40 bursts in each row. Due to the limitation of computing power, the computing core and the three pseudo computing cores will each calculate an 8×8 small matrix in the large two-dimensional data matrix 710.
[0075] Figure 7A The corresponding computing task is that the two-dimensional data matrices accessed by the computing core and the three dummy computing cores do not overlap, so their initial accesses to memory will not hit the L2 cache. Assume that in this case, the latency required for the tested computing core (first data access unit 711) to read its corresponding 8×8 first data matrix 701 is 800 nanoseconds.
[0076] Figure 7B The corresponding computing task is that the two-dimensional data matrices accessed by the computing core and the three pseudo-computing cores partially overlap, so theoretically, when the computing core and the three pseudo-computing cores read data from the memory, there will be some hits in the secondary cache, resulting in a shorter reading delay for the computing core (first data access unit 711). Figure 7B In the case of the test, the delay required by the computing core (first data access unit 711) to read its corresponding 8×8 first data matrix 701 is greater than Figure 7A In this case, the delay is 800ns, which is definitely not in line with the design expectations (because there are partial hits in the L2 cache, the delay should be less than 800ns). Then, the actual data access waveform can be dumped to the L2 cache designer to observe the actual behavior of the computing core (first data access unit 711) in the L2 cache when reading data from the memory, analyze the reasons why the delay does not meet expectations, and verify whether the L2 cache behavior meets the design expectations. Figure 7B The delay required for the computing core (first data access unit 711) to read its corresponding 8×8 first data matrix 701 is 780ns, which is only slightly less than Figure 7A The 800ns delay in this case is slightly lower and does not conform to the delay of the theoretical L2 cache hit rate in the current situation. The actual data access waveform can also be dumped to the L2 cache designer to analyze and compare whether the L2 cache hit rate is consistent with the theoretical value.
[0077] By adopting the data access delay evaluation system and method of the embodiments of the present disclosure, the hit rate of the secondary cache under multi-core parallel tasks can be controlled by controlling the size of the corresponding two-dimensional data matrix read and written by each pseudo computing core and the starting address in the memory, thereby simulating the reuse of the matrix of multi-core concurrent computing tasks in real computing tasks.
[0078] Figure 9 This is a flow chart of an application scenario of the data access delay evaluation system and method according to an embodiment of the present disclosure. In this application scenario, each data access unit reads a two-dimensional data matrix in a memory to test the read data delay of a secondary cache. Figure 9 As shown, the application scenario mainly includes the following steps 901 to 914.
[0079] Step 901 : Send memory access parameters to each data access unit in the data access delay evaluation system, and then execute step 902 .
[0080] Step 902 : Each data access unit is started synchronously based on the memory access parameter, and then step 903 is executed.
[0081] Step 903 : Each data access unit reads data from the memory, and then executes step 904 .
[0082] Step 904 : For any data access unit, determine whether all L2 cache misses occur when reading the two-dimensional data matrix. If yes, execute step 905 ; otherwise, execute step 907 .
[0083] Step 905 : Read the delay data of the delay data register groups of all missed data access units, and then execute step 906 .
[0084] Step 906 : Calculate the first delay of the data access unit when all misses occur based on the read delay data, and then execute step 909 .
[0085] Step 907 : Read the delay data of the delay data register group of the data access unit where a hit occurs, and then execute step 908 .
[0086] Step 908 : Calculate the second delay of the data access unit when a hit occurs based on the read delay data, and then execute step 909 .
[0087] Step 909 : Determine whether the first delay is greater than the second delay; if so, execute step 910 ; otherwise, execute step 914 .
[0088] Step 910 : Obtain the hit rate of cache blocks in the secondary cache, and then execute step 911 .
[0089] Among them, because different data access units read and write different amounts of overlapping data between the two-dimensional data matrices, for example, Figure 7BIn the relationship between the two-dimensional data matrices shown, depending on different memory access parameters, there may be an overlap of one row of data or two rows of data between the first data matrix and the second data matrix. Therefore, various data overlap situations will result in different hit rates of cache blocks in the secondary cache. A higher hit rate means a faster reading speed for data in the memory, which in turn means a shorter second delay.
[0090] Step 911: Obtain a hit delay ratio according to the first delay and the second delay, and then execute step 912.
[0091] Step 912: Determine whether the hit delay ratio meets the expected requirement of the hit rate of the cache block in the secondary cache. If yes, execute step 913; otherwise, execute step 914.
[0092] In an exemplary embodiment, in step 912, first, a hit delay ratio range associated with the hit rate of the cache block in the L2 cache may be determined. Then, a determination is made as to whether the hit delay ratio is within the hit delay ratio range to determine whether the hit delay ratio meets the expected hit rate requirement for the cache block in the L2 cache. In an exemplary embodiment, if the hit delay ratio is within the hit delay ratio range, the hit delay ratio meets the expected hit rate requirement for the cache block in the L2 cache. If the hit delay ratio is not within the hit delay ratio range, the hit delay ratio does not meet the expected hit rate requirement for the cache block in the L2 cache.
[0093] In an exemplary embodiment, the hit delay ratio range that meets the design requirements and corresponds to different hit rates of cache blocks in the secondary cache may be obtained in advance by means of testing, and then called in step 912 .
[0094] Step 913: The test passes and the test is completed.
[0095] Step 914: Dump the data access waveform and complete the test.
[0096] In other application scenarios, the delay in writing data to the memory can also be tested. Since writing data is usually done directly to the main memory, it is slightly different from reading data. In this case, whether the design requirements are met can be directly determined based on the comparison between the write data delay and the preset write data delay threshold.
[0097] In the data access delay evaluation system and method of the embodiments of the present disclosure, a pseudo-computing core is used to implement the memory access function of the computing core. On the one hand, it is possible to enter the memory and on-chip network test phase in advance when the design of the computing core is not completed, so that the test time of the memory and on-chip network can overlap with the design time of the computing core, thereby helping to shorten the development cycle of the SoC and help improve the development efficiency of the SoC. On the other hand, because the pseudo-computing core does not have the computing function of the computing core, but only has the same memory access function as the computing core, the use of the pseudo-computing core also helps to save the time required for the computing core to perform calculations, thereby helping to shorten the test time of the memory and on-chip network and improve the test efficiency of the memory and on-chip network. At the same time, because in the related art, the SoC will set relevant registers for the memory access of the computing core, then in the data access delay evaluation system and method of the embodiments of the present disclosure, the pseudo-computing core can reuse these relevant registers required for the computing core to access the memory. Therefore, there is no need to design corresponding registers for accessing the memory for each pseudo-computing core separately, which can save the design time of the relevant registers. Moreover, since the pseudo-computing core reuses these registers for accessing the memory, the pseudo-computing core can simulate the real computing core accessing the memory, so that testers at all stages can simulate the real behavior of the computing core accessing the memory by configuring these registers without software programming, and then test the performance of the memory and the on-chip network. On this basis, the embodiment of the present disclosure can also allow the process control of multiple cores in the chip to be realized by configuring these registers, so that multiple cores and multiple processes in the chip can be run, so that the real chip behavior can be simulated during testing. Therefore, it can be directly used to test the read and write performance of the on-chip bus and inter-chip bus between the multiple processes of the chip, reducing the complexity of software programming. In addition, the embodiment of the present disclosure can also allow the maximum outstanding read and write requests and the related read and write request arbitration weights to be controlled by registers to more accurately control the read and write traffic on the bus, which is more flexible and convenient and closer to the actual scenario. The embodiment of the present disclosure also allows the memory to be read and written directly by the configuration of registers. In the process of reading performance testing of the memory, the overlap between different two-dimensional data matrices and the influence of the secondary cache are also taken into account, making the test scenario closer to the actual operation scenario of the artificial intelligence chip and making the test results more reliable.
[0098] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A data access delay evaluation system, characterized in that: include: A parameter configuration module, used for providing memory access parameters during data access delay testing; At least two data access units, coupled to the parameter configuration module and coupled to a memory via an on-chip network, for reading and writing a two-dimensional data matrix in the memory via the on-chip network based on the memory access parameters, wherein at least one of the at least two data access units is a pseudo computing core, the pseudo computing core having the same memory access function as a computing core, and the memory includes a secondary cache and a main memory; a delay data register set coupled to the at least two data access units, configured to obtain delay data of the at least two data access units during a period in which the at least two data access units access the memory; and The delay evaluation module is coupled to the delay data register group and is used to obtain a delay evaluation result according to the delay data of the at least two data access units and a preset delay design condition.
2. The data access delay evaluation system according to claim 1, wherein: At least one of the at least two data access units is a computing core.
3. The data access delay evaluation system according to claim 1, wherein: The data access delay evaluation system further includes: A data dump module is used to dump the waveform data during the reading and writing of the two-dimensional data matrix in the memory when the delay evaluation result indicates that the delay data does not meet the delay design condition requirement.
4. The data access delay evaluation system according to claim 1, wherein: The memory access parameters include at least one of a startup parameter, an abort parameter, a process identification parameter, a global process identification parameter, a read / write request arbitration weight parameter, a maximum number of outstanding write requests parameter, a maximum number of outstanding read requests parameter, a write request length parameter, a write request height parameter, a write request burst transfer length parameter, a write request step parameter, a write address parameter, a read request length parameter, a read request height parameter, a read request burst transfer length parameter, a read request step parameter, and a read address parameter.
5. The data access delay evaluation system according to claim 1, wherein: The delay data includes at least one of the number of read requests used to count read delays, the total delay of read requests, the number of completed read requests, the number of completed write requests, and the total busy time of the data access unit; The delay data register group includes at least one of a read delay request quantity register, a read delay register, a completed read request quantity register, a completed write request quantity register, and a busy cycle register; wherein, The read delay request quantity register is used to record the number of read requests used to count the read delay, the read delay register is used to record the total delay of the read requests, the completed read request quantity register is used to record the number of completed read requests, the completed write request quantity register is used to record the number of completed write requests, and the busy cycle register is used to record the total busy time of the data access unit.
6. The data access delay evaluation system according to claim 1, wherein: The delay evaluation module includes: a calculation submodule, configured to calculate, based on the delay data of the at least two data access units, a first delay of any one of the at least two data access units when all L2 cache data reads miss, and a second delay of the any one of the data access units when there is a hit in the L2 cache data read; A judgment submodule is used to obtain the delay evaluation result according to the first delay, the second delay and the delay design condition.
7. The data access delay evaluation system according to claim 1, wherein: The delay design conditions are: The first delay of any one of the at least two data access units when all secondary cache data reads miss is greater than the second delay of any one of the data access units when there is a hit in the secondary cache data read, and the hit delay ratio obtained based on the first delay and the second delay is within a preset hit delay ratio range.
8. A data access delay evaluation method, comprising: configuring memory access parameters for at least two data access units coupled to a memory via an on-chip network, wherein at least one of the at least two data access units is a pseudo computing core having the same memory access function as a computing core, and the memory includes a secondary cache and a main memory; The at least two data access units read and write the two-dimensional data matrix in the memory through the on-chip network based on the configured memory access parameters; During a period in which the at least two data access units access the memory, acquiring delay data of the at least two data access units; A delay evaluation result is obtained according to the delay data of the at least two data access units and a preset delay design condition.
9. The data access delay evaluation method according to claim 8, wherein: There is no overlap between the two-dimensional data matrices read and written by the data access units with different memory access parameter configurations, or there is overlap between the two-dimensional data matrices read and written by the data access units with different memory access parameter configurations.
10. The data access delay evaluation method according to claim 8, characterized in that: Obtaining a delay evaluation result according to the delay data of the at least two data access units and a preset delay design condition includes: Calculating, based on the delay data of the at least two data access units, a first delay of any one of the at least two data access units when all L2 cache data reads miss, and a second delay of the any one of the data access units when there is a hit in the L2 cache data read; The delay evaluation result is obtained according to the first delay, the second delay and the delay design condition.
Citation Information
Patent Citations
Hardware basic service component test system
CN119806925A
Evaluation method and evaluation system for semiconductor storage device
US20090217111A1
Memory chip, memory module and method for pseudo-accessing memory bank thereof
US20210049095A1
System and method for modeling memory devices with latency
US20220066801A1
Method and apparatus for functional testing of memory related circuits
US6341094B1