Address mapping calculation method for encrypted non-2 integer power counter of NPU memory

By using non-2 integer power counter address mapping calculation method in NPU system, the problem of frequent counter overflow is solved, and more efficient resource utilization and system performance improvement is achieved.

CN119938550APending Publication Date: 2025-05-06NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510022031.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing 2 integer power counter address mapping calculation method has the problem of frequent counter overflow in NPU systems that integrate the last-level cache, resulting in serious degradation of system performance.

Method used

Using the non-2 integer power counter address mapping calculation method, by shifting the conventional data address right to the S bit and performing division operations with L, the index value of the counter data block in the counter storage space is obtained, and the counter data block address is calculated based on the relative address.

Benefits of technology

This method can flexibly adjust the number of counters contained in a single counter block, avoid frequent overflow of counters, and improve the hardware resource utilization and system performance of the NPU system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938550A_ABST
    Figure CN119938550A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer system structures and integrated circuit design. The invention provides an address mapping calculation method for an encrypted non-2 integer power counter of an NPU memory. According to the embodiment of the invention, a conventional data address is shifted rightwards by S bits, and division operation is carried out on the conventional data address and L, so that an index value of a counter data block mapped by a conventional data block in a counter storage space is obtained; moving the index value left by S bits to obtain a relative address of a counter data block in a counter storage space; and adding the relative address and the initial address of the storage space of the counter to obtain a data block address of the counter. According to the method, the number of counters contained in a single counter block can be flexibly adjusted, and a reasonable counter bit width is set for an NPU application program, so that frequent overflow of the counters is avoided while the utilization rate of hardware resources on an NPU chip and the system performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to the technical fields of computer system structure and integrated circuit design, and in particular to a method for calculating address mapping of an NPU memory encryption non-integer power of 2 counter. Background Art

[0002] Neural networks are widely used in computer vision, natural language processing and other fields. They have great economic value and are one of the most important research areas in artificial intelligence and machine learning. Neural networks require high computing power to complete a large number of operations such as multiplication and accumulation. Domain-specific NPUs have obvious advantages over general-purpose processors such as CPUs and GPUs in running neural networks to perform inference operations, and have become an artificial intelligence processor that major semiconductor manufacturers are competing to develop. With the widespread use of NPUs, their security has received more and more attention, and there is an urgent need to build an NPU Trusted Execution Environment (TEE). The Memory Encryption Engine (MEE) ensures the data security of the main memory (Dynamic Random Access Memory, DRAM) outside the NPU chip and is an important part of TEE-related technologies. MEE is located between the NPU core and the storage controller. The NPU system structure with integrated MEE is as follows: Figure 1 As shown in the figure: The AES (Advanced Encryption Standard) engine encrypts data based on the AES algorithm to ensure the confidentiality of the data; the MAC engine generates the MAC of the data based on the hash algorithm to ensure the integrity of the data. AES encryption usually uses the counter mode, such as Figure 2 As shown in the figure, each data block in DRAM corresponds to a counter. These counter values ​​are used as leaf nodes of BMT (Bonsai Merkle Tree). The freshness of data is guaranteed based on BMT. Memory encryption operations generate a large amount of security metadata (counter data, BMT data, MAC data). These security metadata are stored in off-chip DRAM, and their memory access operations consume a large amount of memory bandwidth. MEE generally caches security metadata by integrating counter / MAC / BMT cache. However, due to factors such as chip area and energy consumption, the capacity of these caches is limited. A large amount of security metadata access will still seriously reduce the system performance of NPU.

[0003] Data access in the NPU is based on data blocks. The data block is the same size as the counter / MAC / BMT cache data block and is an integer power of 2. The commonly used granularity is 32 bytes, 64 bytes, etc. For each regular data access, its counter address needs to be calculated based on the regular data address in order to access the counter cache to obtain the counter value corresponding to the data block. The mapping relationship between regular data and counter data is as follows: Figure 3 As shown in (a), the size of the regular data block and the counter data block are both N bytes, K regular data blocks constitute a data chunk, and one chuck corresponds to one counter data block; the counter data block uses the split counter mode to record the counter values ​​of K regular data blocks, and the counter value of each regular data block is the sum of the major counter and its corresponding minor counter. The prior art requires that K is an integer power of 2, and the regular data address generally starts from 0. Based on the regular data address, the counter data block address corresponding to the regular data can be obtained by using shift and addition operations, such as Figure 3 (b). M and S are both positive integers greater than 0, and CBDA_st is the starting address of the counter storage space. Based on the index value of the regular data block in the data chunk, the corresponding minor counter can be taken out from the counter data block, and then the counter value corresponding to the regular data block can be obtained.

[0004] When a regular data block is updated by a write operation, its corresponding minor counter is incremented by 1. If a data block is written multiple times, causing its corresponding minor counter to overflow, the major counter in the counter data block is incremented by 1, all minor counters are cleared, and the regular data chunk corresponding to the counter data block is re-encrypted, and the BMT is updated at the same time. Therefore, counter overflow will cause a large performance loss. In order to reduce counter overflow, the number of regular data blocks corresponding to a single counter data block can be reduced, that is, K can be reduced, thereby increasing the bit width of the major counter and minor counter. In recent years, with the development of NPU, the last level cache (LLC) has gradually been integrated between the NPU core and the main memory DRAM to improve the system performance of the NPU, such as Figure 4As shown, for example, Huawei's Shengteng AI processor is equipped with a last-level cache shared by the CPU and AI core. Since neural networks generally have good memory access locality, the introduction of LLC can effectively reduce the number of times regular data blocks are written in DRAM, thereby reducing the number of counter overflows. At this time, the larger counter bit width causes resource waste. The method of reducing the counter bit width, that is, increasing the K value, can be used so that each counter data block can contain more counter values, thereby improving resource utilization. At the same time, it is also beneficial to improve the hit rate of the counter cache. However, since K is an integer power of 2 ( ), a small change in M ​​can cause a large change in K. For example, when M changes from 6 to 7, K changes from 64 to 128. Since the size of the counter data block remains unchanged, it is very easy to cause the counter bit width to change from too large to too small, resulting in frequent counter overflows, which in turn seriously reduces the performance of the NPU system. Therefore, the integer power of 2 counter address mapping calculation method is not suitable for NPU systems integrated with LLC. Based on publicly available literature, there is currently no record of NPUs with memory encryption functions using integer power of 2 counter address mapping calculation methods.

[0005] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0007] The purpose of the embodiments of the present disclosure is to provide a method for calculating the address mapping of a non-integer power of 2 counter for NPU memory encryption, thereby overcoming one or more problems caused by the limitations and defects of the relevant technology at least to a certain extent.

[0008] According to a first aspect of an embodiment of the present disclosure, a method for calculating an address mapping of an NPU memory encryption non-integer power of 2 counter is provided, the method comprising: Assume that the size of a data block is N bytes, a regular data chunk contains L regular data blocks, and the counter data block uses the split counter mode to store the counter values ​​corresponding to the L regular data blocks; Shift the regular data address right by S bits and divide it by L to obtain the index value of the counter data block to which the regular data block is mapped in the counter storage space; where N=2 S ; Shift the index value to the left by S bits to obtain the relative address of the counter data block in the counter storage space; Add the relative address to the starting address of the counter storage space to obtain the counter data block address.

[0009] Furthermore, L is set to any positive integer according to the memory access characteristics of the NPU.

[0010] Further, the regular data address is shifted right by S bits and divided by L to obtain the index value of the counter data block to which the regular data block is mapped in the counter storage space, including: The regular data address is shifted right by S bits and divided by L, and the integer part of the division result is taken as the index value of the counter data block to which the regular data block is mapped in the counter storage space.

[0011] Furthermore, the expression for address mapping calculation is:

[0012] in, is the counter data block address, is the regular data address, S is the number of bits of displacement, is the starting address of the counter storage space, Represents the integer part of the result of a division operation.

[0013] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects: In the embodiment of the present disclosure, through the above-mentioned NPU memory encryption non-integer power of 2 counter address mapping calculation method, on the one hand, the conventional data address is right-shifted by S bits and divided by L to obtain the index value of the counter data block mapped to the conventional data block in the counter storage space; the index value is left-shifted by S bits to obtain the relative address of the counter data block in the counter storage space; the relative address is added to the starting address of the counter storage space to obtain the counter data block address. On the other hand, this method can flexibly adjust the number of counters contained in a single counter block and set a reasonable counter bit width for the NPU application, thereby improving the utilization of NPU on-chip hardware resources and system performance while avoiding frequent overflow of the counter.

[0014] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0016] Figure 1 A schematic diagram showing an NPU system integrating MEE in an exemplary embodiment of the present disclosure is shown; Figure 2 A schematic diagram showing a Counter encryption mode in an exemplary embodiment of the present disclosure; Figure 3 A schematic diagram showing address mapping calculation based on an integer power of 2 counter in an exemplary embodiment of the present disclosure; Figure 4 A schematic diagram showing an NPU system with integrated last-level cache in an exemplary embodiment of the present disclosure; Figure 5 A method for calculating an address mapping of a non-2 integer power counter for NPU memory encryption in an exemplary embodiment of the present disclosure is shown; Figure 6 A schematic diagram showing address mapping calculation based on a non-2 integer power counter in an exemplary embodiment of the present disclosure; Figure 7 The address calculation pipeline of the non-integer power of 2 counter in the exemplary embodiment of the present disclosure is shown; Figure 8 The NPU cycle-level simulation system in the exemplary embodiment of the present disclosure is shown; Fig. 9 A schematic diagram showing an example of a counter data block structure in an exemplary embodiment of the present disclosure; Fig.10 shows a normalized NPU system operation cycle in an exemplary embodiment of the present disclosure; Fig.11 The normalized DRAM access energy consumption in an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0017] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0018] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0019] This example implementation first provides a method for calculating the address mapping of a non-2 integer power counter for NPU memory encryption, which can be applied to NPUs with memory encryption functions. NPUs are widely used in various devices, such as mobile phones, smart security cameras, etc. [flexibly adjusted according to specific circumstances, such as if the terminal device is a server, etc.]. Reference Figure 5 As shown in , the method may include the following steps: Step S101: Assume that the size of a data block is N bytes, a regular data chunk contains L regular data blocks, and a counter data block uses a split counter mode to store counter values ​​corresponding to the L regular data blocks; Step S102: Shift the regular data address rightward by S bits and perform a division operation with L to obtain the index value of the counter data block to which the regular data block is mapped in the counter storage space; wherein N=2 S ; Step S103: Shift the index value to the left by S bits to obtain the relative address of the counter data block in the counter storage space; Step S104: Add the relative address to the starting address of the counter storage space to obtain the counter data block address.

[0020] Through the above-mentioned NPU memory encryption non-integer power of 2 counter address mapping calculation method, on the one hand, the regular data address is right-shifted by S bits and divided by L to obtain the index value of the counter data block mapped to the regular data block in the counter storage space; the index value is left-shifted by S bits to obtain the relative address of the counter data block in the counter storage space; the relative address is added to the starting address of the counter storage space to obtain the counter data block address. On the other hand, this method can flexibly adjust the number of counters contained in a single counter block and set a reasonable counter bit width for the NPU application, thereby improving the utilization of NPU on-chip hardware resources and system performance while avoiding frequent overflow of the counter.

[0021] Next, we will refer to Figures 5 to 11 Each step of the above method in this example implementation is described in more detail.

[0022] In step S101 to step S104, the regular data address is shifted right by S bits and divided by L to obtain the index value of the counter data block to which the regular data block is mapped in the counter storage space; the index value is shifted left by S bits to obtain the relative address of the counter data block in the counter storage space; the relative address is added to the starting address of the counter storage space to obtain the counter data block address.

[0023] For example, for each regular data access, its counter address needs to be calculated based on the regular data address, so as to access the counter cache to obtain the counter value corresponding to the data block. Figure 6 As shown in FIG. 1 , it is a schematic diagram of address mapping calculation based on a non-2 integer power counter. Figure 6 As shown in (a), the data block size is N bytes. A regular data chunk contains L regular data blocks, corresponding to a counter data block. The counter data block uses the split counter mode to store the counter values ​​corresponding to the L regular data blocks. Different from the traditional method, L can be set to any value according to the specific memory access characteristics of the NPU. The process of calculating the counter data block address based on the regular data block address input is as follows: Figure 6 (b) As shown. The conventional data address is obtained by shifting and dividing operations to obtain the index value of the counter data block mapped to the conventional data block in the counter storage space, and then the index value is shifted left by S bits to obtain the relative address of the counter data block in the counter storage space. The relative address is added to the starting address CDBA_st of the counter storage space to obtain the required counter data block address CDBA; where N=2 S Since L can be configured as any positive integer, The result may be a fraction. Figure 6 (b) Indicates the integer part of the division result. The address calculation pipeline is as follows Figure 7 As shown, dedicated hardware circuits are used to accelerate each operation, among which the longest time-consuming operation is the integer division operation. The integer division hardware accelerator implements this part of the operation and obtains the quotient of the dividend and the divisor, that is, .

[0024] In a specific embodiment, a modeling simulation method is used to evaluate the effect of the present application, and an NPU cycle-level simulation system integrating LLC and MEE is constructed based on the widely used NPU simulator STONNE and DRAM simulator DRAMsim, such as Figure 8As shown in Figure 2, STONNE is used to model the NPU core, DRAMsim is used to model the DRAM and storage controller, LLC and MEE are used to model the last-level cache and memory encryption engine respectively. The main parameter settings of the NPU system are shown in Table 1: the data block size of all caches is 64B, and the LRU (Least Recently Used) replacement strategy is adopted; the pipeline delay of the non-integer power of 2 counter address calculation is 35 clock cycles. The NPU simulator supports three architectures, namely MAERI, SIGMA and TPU, and runs 5 neural networks (LeNet, AlexNet, MobileNet_V2, ResNet 50, YoloV4-Tiny), thus constituting 15 test vectors, as shown in Table 2. The NPU system that uses the integer power of 2 counter address mapping calculation method is the Baseline, and its counter cache data block contains 64 counter values, as shown in Figure 2. Fig. 9 As shown in (a); the NPU system that uses the non-2 integer power counter address mapping and calculation method is NPU-N2PM (NPU with non-2 integer power counter address mapping and calculation method), and its counter cache data block contains 96 counter values, such as Fig. 9 (b) Based on 15 test vectors, the performance evaluation results show that the 5-bit minorcounter size used by NPU-N2PM can ensure that the counter does not overflow. Therefore, the counter data block structure used by Baseline has certain redundancy. In addition, the system operation cycle and DRAM access energy consumption of NPU-N2PM are normalized relative to Baseline, and the results are as follows: Fig.10 and Fig.11 As shown, it can be seen that the non-2 integer power counter address mapping calculation method is beneficial to improving the system performance of the NPU.

[0025] Table 1 Main configuration parameters of NPU system

[0026] Table 2 Test vector set

[0027] Through the above-mentioned NPU memory encryption non-integer power of 2 counter address mapping calculation method, on the one hand, the regular data address is right-shifted by S bits and divided by L to obtain the index value of the counter data block mapped to the regular data block in the counter storage space; the index value is left-shifted by S bits to obtain the relative address of the counter data block in the counter storage space; the relative address is added to the starting address of the counter storage space to obtain the counter data block address. On the other hand, this method can flexibly adjust the number of counters contained in a single counter block and set a reasonable counter bit width for the NPU application, thereby improving the utilization of NPU on-chip hardware resources and system performance while avoiding frequent overflow of the counter.

[0028] It should be noted that, although the steps of the method in the present disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc. In addition, it is also easy to understand that these steps may be, for example, executed synchronously or asynchronously in multiple modules / processes / threads.

[0029] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0030] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of a module or unit described above can be further divided into multiple modules or units for concretization. The components displayed as modules or units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. Those of ordinary skill in the art can understand and implement it without paying creative work.

[0031] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

Claims

1. A method for calculating address mapping of a non-integer power of 2 counter for NPU memory encryption, characterized in that: The method includes: Assume that the size of a data block is N bytes, a regular data chunk contains L regular data blocks, and the counter data block uses the split counter mode to store the counter values ​​corresponding to the L regular data blocks; Shift the regular data address right by S bits and divide it by L to obtain the index value of the counter data block to which the regular data block is mapped in the counter storage space; where N=2 S ; Shift the index value to the left by S bits to obtain the relative address of the counter data block in the counter storage space; Add the relative address to the starting address of the counter storage space to obtain the counter data block address.

2. According to claim 1, the NPU memory encryption non-2 integer power counter address mapping calculation method is characterized in that: L is set to any positive integer according to the memory access characteristics of the NPU.

3. According to claim 2, the NPU memory encryption non-2 integer power counter address mapping calculation method is characterized in that: The regular data address is shifted right by S bits and divided by L to obtain the index value of the counter data block to which the regular data block is mapped in the counter storage space, including: The regular data address is shifted right by S bits and divided by L, and the integer part of the division result is taken as the index value of the counter data block to which the regular data block is mapped in the counter storage space.

4. According to claim 3, the NPU memory encryption non-integer power of 2 counter address mapping calculation method is characterized in that: The expression for address mapping calculation is: in, is the counter data block address, is the regular data address, S is the number of bits of displacement, L is the number of data blocks, is the starting address of the counter storage space, Represents the integer part of the result of a division operation.