DNN training data access-oriented multi-stack magnetic caching method, medium and equipment

By using three-stack magnetic random memory and cache migration algorithm to optimize memory access during deep neural network training, the memory wall and write performance bottlenecks of traditional memory are solved, and efficient DNN training performance is achieved.

CN120295941APending Publication Date: 2025-07-11NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365888.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

During the existing deep neural network training process, traditional volatile memory faces challenges such as memory wall, high latency, and high power consumption. Nonvolatile memory such as MRAM has a write performance bottleneck in high-speed application scenarios, and TLC MRAM is difficult to write in LSB, resulting in limited storage density and performance.

Method used

Three-stack magnetic random memory (sTLC MRAM) is used for memory access, data approximate storage strategies and cache migration algorithm are designed, and GPU L2 cache is optimized through physical chunking and remapping, combined with on-chip network, memory access is optimized, and memory access is reduced by using the spatio-temporal locality of data access.

Benefits of technology

It improves the storage density and access speed of deep neural network training, reduces energy consumption, ensures that the training accuracy is not affected, and realizes high-density and high-speed DNN training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295941A_ABST
    Figure CN120295941A_ABST
Patent Text Reader

Abstract

The invention provides a multi-stack magnetic caching method, medium and equipment for DNN (Deep Neural Network) training data access, belongs to the technical field of data access performance and cache management performance for a deep neural network, and reduces unit-level and interconnection-level delays through an approximate storage architecture and a cache remapping migration strategy. In a unit level, an index bit of a weight during DNN training is fixed and stored in a hard bit, power consumption and delay overhead caused by frequent write-in of the hard bit are reduced, and a compensation bit is introduced to avoid precision loss. In an interconnection level, shared caches are physically partitioned based on a path-based cache set partitioning method, so that access delay of data among the caches is reduced, and meanwhile, efficient communication among cache blocks is realized by utilizing an on-chip internet. In view of time and space locality of data access, a remapping migration strategy is further proposed to reduce overall read-write delay in storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data access performance and cache management performance for deep neural networks, and particularly relates to a multi-stack magnetic cache method, medium and device for DNN training data access. Background Art

[0002] Deep neural networks (DNNs) have shown extraordinary potential in numerous practical application scenarios. By training with large-scale datasets, their performance has been significantly improved by enhancing network depth and expanding scale. However, this is also accompanied by challenges such as the memory wall, high latency, and high power consumption. It is particularly crucial to study high-density optimized memories to meet the huge storage requirements of DNNs without sacrificing training accuracy and without increasing hardware costs. Graphic Processing Units (GPUs), with their powerful backend support and high adaptability to dense matrix operations, are widely used to accelerate the training process of DNNs, but the memory resources they provide are usually based on traditional volatile memories.

[0003] However, traditional cache devices are gradually approaching the physical limit of their nanoscale scalability and facing technical bottlenecks such as leakage current and low density. To meet the increasing demands of ultra-low power and high-speed computing systems, non-volatile memories (NVMs), such as ferroelectric random access memories (FeRAMs), phase change memories (PCMs), resistive random access memories (ReRAMs), etc., have been widely studied. Among them, magnetic random access memories (MRAMs) have become one of the most promising candidate devices due to their low power consumption, high-speed access, infinite read / write cycles, and excellent radiation resistance. Although STT-MRAM has a high storage density, it still has the problem of large write overhead. The SOT-MRAM technology introduced to solve the write performance integrates more transistors inside the cell, which instead limits their capacity expansion in high-speed application scenarios.

[0004] The multi-level cell (MLC) technology improves storage density and ensures performance by vertically stacking two storage devices in a single cell. To further increase storage density, the MRAM cell structure has evolved from MLC to TLC. The previously proposed TLC MRAM achieved storing 3 bits of data in a single cell through a three-stack structure composed of two STTs and one SOT. The recently proposed sTLC MRAM reduces the number of STTs to 1, requiring only a single-step write operation or a two-step write operation and a single-step read operation during the memory access process, which not only ensures data distinguishability but also effectively reduces power consumption and improves memory access speed.

[0005] However, sTLC MRAM still faces the challenge of difficult writing when storing the least significant bit (LSB) of STT. Experimental results show that the longest delay required to flip STT is about 10 times that of non-flipping. Although existing advanced TLC designs have been able to achieve fast writing of 2 / 3 bits and high-speed reading of the entire cell, the above problems still restrict the full-speed operation of TLC. At the same time, for memory-intensive applications, especially in the scenario of neural network training, large capacity and high performance are always the key requirements for parallel processing. Therefore, designing a cache architecture based on high-capacity TLC to optimize the training data pattern has important application value. Summary of the Invention

[0006] Aiming at the memory wall problem that becomes increasingly prominent due to the complexity of deep neural networks and the huge amount of data they process, the present invention provides a multi-stack magnetic cache method, medium and device for accessing DNN training data. Aiming at the application scenario of deep neural network training, a non-volatile, low-power and high-density storage solution based on three-stack magnetic random access memory is proposed, and a GPU non-uniform cache design based on the network-on-chip is explored.

[0007] The present invention first designs a data approximate storage strategy for DNN training, and on the premise of maintaining logical sharing, proposes physical partitioning and re-mapping of the GPU L2 cache. At the same time, based on the spatio-temporal locality of data access, a migration algorithm is designed to reduce the triple memory access latency at the cell level, array level and interconnection level.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] In the first aspect, the present invention provides a multi-stack magnetic cache method for accessing DNN training data, using sTLCMRAM for memory access, including:

[0010] For the weight data of the network, 32-bit single-precision floating-point numbers conforming to the IEEE 754 standard are used for representation and processing, including 1 sign bit, 8 exponent bits and 23 mantissa bits; at the start of training, the exponent bits of the weight data are fixed; during training, a 3-bit original code compensation is performed on the exponent of the weight data, and at the same time, 5 mantissa bits are truncated; where the LSB stores the fixed exponent bits, the CSB stores the sign bit, the CSB and the MSB store 18 mantissa bits, and the MSB stores 3 exponent compensations;

[0011] For the input data of the network, during training, 11 cell units are used to store 32-bit data, and one LSB is left empty; the input data is rearranged into consecutive cell units according to different features.

[0012] Optionally, it further includes the following cache partitioning and reorganization operations:

[0013] Partition the cache sets of the GPU L2 cache into a certain number of cache blocks (Tiles) based on ways, and combine the Tiles belonging to the same way into different clusters; among them, the data transmission between the Tiles is realized by means of the on-chip interconnection network.

[0014] Optionally, it further includes a cache migration operation, which is restricted between different ways within the same cache set.

[0015] Optionally, the cache migration operation is triggered by a direction counter, and the direction counter is a signed number represented by two's complement, and is reset in the following four cases: when the cache line is newly acquired; when the cache set is full and the cache line is evicted by the least recently used algorithm; when the value of the direction counter after migration exceeds the predefined range; after the migration operation itself occurs.

[0016] Optionally, the update of the direction counter depends on the relative position between the source Tile and the target Tile.

[0017] Optionally, when the target Tile is outside the silent radius range of the source Tile, and the value of the direction counter between the source Tile and the target Tile cannot correspond to their relative position, the cache migration operation is triggered.

[0018] In a second aspect, the present invention provides a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the multi-stack magnetic cache method for DNN training data access as described in the first aspect.

[0019] In a third aspect, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the multi-stack magnetic cache method for DNN training data access as described in the first aspect.

[0020] The beneficial effects of the present invention are as follows: The multi-stack magnetic cache architecture with enhanced DNN training data access performance proposed by the present invention is applicable to the training of all deep neural networks. The data approximate storage strategy proposed by the present invention approximately stores the weight data in the three-stack magnetic random access memory, which can ensure that the training accuracy is not affected. The way-based cache set partitioning method and the corresponding remapping strategy proposed by the present invention physically partition the L2 shared cache of the GPU, and at the same time maintain logical sharing and consistency through the on-chip interconnection network. The present invention also designs a cache migration strategy for the temporal and spatial locality of data access, reducing the overall read / write latency in storage, thereby realizing high-density and high-speed DNN training. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Schematic diagram of a multi-stack cache with enhanced DNN training data memory access performance provided in this embodiment;

[0022] Figure 2 Schematic diagram of the block-based GPU L2 shared cache mapping strategy provided in this embodiment;

[0023] Figure 3 Schematic diagram of the block-based GPU L2 shared cache migration strategy provided in this embodiment. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0025] The differences between soft bits and hard bits in the memory limit the improvement of the unit-level write performance; at the same time, in large-capacity storage, as the capacity increases, the latency is mainly affected by the interconnection line latency, thus diluting the effect of unit-level latency performance optimization. Therefore, in one embodiment, the present invention proposes a multi-stack magnetic cache method for DNN training data memory access, and the specific architecture is as Figure 1 shown, and this figure shows a multi-stack cache architecture with enhanced DNN training data memory access performance.

[0026] The hardware system uses 32-bit single-precision floating-point numbers that conform to the IEEE 754 standard to represent and process the weights in the network, where 1 bit is the sign bit, 8 bits are the exponent bits, and 23 bits are the mantissa bits. Many neural network storage schemes note that the values of the weights always vary within a limited range. Based on this observation and experimental verification, it is found that the exponent values of the weight data hardly change during the training process, so it is proposed to fix the exponent bits of the weights and approximately store them in the hard bits to alleviate the high write overhead problem of the hard bits.

[0027] In practical applications, at the beginning of each training epoch, the exponent bits of the initial weights are fixed. To improve the prediction accuracy of the DNN model, a 3-bit original code compensation is performed on the exponents of the weight data to further reduce the approximation error. At the same time, 5 bits of the mantissa are truncated to reduce the storage space consumption.

[0028] Figure 1(a) is a flow schematic diagram of the approximate storage method, using a specific example to illustrate the entire weight encoding process. First, in step 1, at the beginning of the training period, the weight value is "0.0665", which is represented in binary as "00111101100010000011000100100111". After multiple iterations, the weight is updated to "0.0621", and its binary representation is "00111101011111100101110010010010". To ensure that the data in the hard bits remains unchanged in each training period, a compensation bit of "101" is added in step 2, which means that the exponent of the weight needs to be increased by "-1" on the basis of the original data. At the same time, step 3 compresses the last 5 least significant bits of the mantissa to obtain a new 18-bit mantissa "111111001011100100". Combining the above steps, the new weight format is represented as "001111011111111001011100100101".

[0029] Next, Figure 1 (b) shows the weight storage scheme. To improve the prediction accuracy of the DNN model, according to the initial weight value of each training period, a 3-bit true value compensation is performed on the exponent of the weight data to further reduce the approximation error. The least significant bit LSB (0 - 7) stores the fixed exponent bits, the middle significant bit CSB (8) stores the sign bit, CSB (9 - 18) and the most significant bit MSB (19 - 26) store the truncated 18-bit mantissa (23 - 5 = 18), and MSB (27 - 29) stores the 3-bit exponent compensation.

[0030] Figure 1 (c) is the input data storage scheme. During the neural network training process, the input data needs to be read multiple times. To make full use of the parallel reading advantage of the sTLC structure, this embodiment uses 11 cell units to store 32-bit data and leaves one LSB empty. The feature map is stored in a traditional DRAM device. During the training process, the data is rearranged into continuous storage blocks (cell units) according to different features, thus effectively utilizing the reading and writing advantages of the sTLC.

[0031] sTLC only uses 3 transistors to store 3-bit data, and its storage density is much higher than that of SRAM, so it can achieve a smaller interconnection delay. In addition, as the scale of the DNN training model continues to expand, the demand for storage capacity is also increasing, which makes the overall cache access delay more and more affected by the interconnection level, thus further promoting the possibility of replacing SRAM with sTLC MRAM.

[0032] This embodiment proposes a cache set partitioning scheme based on ways to reduce multi-processor access conflicts and interconnection latency. The large-capacity shared cache adopts a set-associative mapping method. Based on ways, the cache set is partitioned into a certain number of cache tiles, and the cache tiles are combined into different clusters. At the same time, the on-chip network (Network on Chip, NoC) is used to ensure efficient data transmission between cache tiles. In this architecture, each computing unit (CU) can access any part of the GPU L2 cache stored on any tile through the router of the NoC. This means that although the system is physically partitioned, the logical sharing feature of the GPU L2 cache is still retained.

[0033] By physically partitioning the large-capacity GPU L2 cache into multiple small pieces, the latency at the interconnection level can be further reduced. This embodiment adopts a Tile-based NoC architecture. In this case, the new read / write access latency consists of three parts: T1 (unit level), T2 (interconnection level), and T3 (NoC level). Due to the partitioning, the interconnection-level latency of T2 has been significantly reduced. However, this also introduces a third access latency part. Therefore, to achieve an actual overall latency reduction, the optimization of T3 is indispensable.

[0034] Since the source (Tile Src) and the destination (Tile Dest) are sometimes far apart, resulting in a large T3 (Hops(Src, Dest)), this embodiment proposes a cache remapping migration strategy that utilizes the spatial and temporal locality of data access. This strategy consists of two parts: remapping and migration. As Figure 2 shown, an 8×8 GPU computing unit (CU) and a 16-way set-associative GPU L2 cache are used to illustrate the cache remapping scheme. This remapping method ensures that within a radius of 2, for all cache sets (set0, set1, set2,...), at least one of the 16 ways can be found for each cache set. The migration strategy based on the remapping scheme restricts the migration to different ways within the same cache set. For example, as Figure 2, assume that Tile Src and Tile Dest store the cachelines required for the CU calculations of Tile Src. Since Tile Dest stores the cache sets belonging to Group 3 (G3), the new migration destination (Tile new Dest) will be the Tile in the same cluster as Tile Src and store the same Group G3 as Tile Dest. That is, the migration will be completed by swapping the two-way cachelines of the same cache set stored in Tile Dest and Tile new Dest. In addition, the Silence Radius (SR) is designed to guide whether a migration should be performed. If it is outside the silence radius, migration is performed; otherwise, it remains silent without migration.

[0035] The migration strategy based on the remapping scheme proposed in this embodiment restricts the migration operation between different ways within the same cache set. Due to the Tag-based storage and lookup characteristics of the cache, this method does not require remapping the migrated cache set, thus reducing the operation complexity. Considering that migration itself incurs certain costs, the design of the migration strategy needs to be balanced and should not be too aggressive to avoid adding unnecessary overhead; at the same time, it should not be too conservative to avoid missing the opportunity to utilize the advantage of data access locality. For this reason, this embodiment introduces the concept of the silence radius to guide whether a migration should be performed or remain in the silent state.

[0036] For the migration part, see Figure 3 , Figure 3Shows the schematic diagram of the GPU L2 cache migration strategy after chunking. Migration sometimes requires historical context to guide the process. In this embodiment, a direction counter EWNS is designed to further guide whether migration should be performed. The counter is a signed number represented in two's complement. In the following four cases, the counter will be reset: (1) The cacheline is newly fetched into the cache; (2) The cache set is full, and the cacheline is evicted by the least recently used (LRU) algorithm; (3) After migration, the counter exceeds the predefined range; (4) After migration. The update of the counter is determined by the relative position between Tile Src and Tile Dest. When Tile Src needs to use the cacheline in Tile Dest for calculation, assuming that TileDest is located northwest of Tile Src, that is, Tile Dest is needed by Tile Src located in its southeast direction, then the EWcnt of the cacheline accessed in Tile Dest will increase by 1, while the NS cnt will decrease by 1. That is to say, a positive EWcnt indicates that the Tile Src to its east is more likely to need this cacheline, while a negative value indicates that the Tile Src to its west is more likely to need this cacheline. Similarly, a positive NScnt means that the Tile Src to its north is more likely to need this cacheline, while a negative value indicates that the Tile Src to its south is more likely to need this cacheline. In summary, migration will only occur when TileDest is outside the SR of TileSrc and the size relationship between the EWNScnts of TileSrc and TileDest does not correspond to the relative position.

[0037] Figure 3The migration strategy is explained through specific examples. Assume that SR is 2. For Scenario 1, Tile Dest is within the SR of Tile Src, so it remains silent and does not migrate. For Scenario 2, Tile Dest is outside the SR of Tile Src, so the EWNScnts of the two cachelines to be swapped on Tile Src and Tile Dest are continuously compared. The relative position of Tile Src is northwest of Tile Dest. Assume that the EWcnt of Tile Src is less than that of Tile Dest, and the NScnt of Tile Src is greater than that of Tile Dest. That is, the magnitude relationship between the EWNScnts of Tile Src and Tile Dest corresponds to the relative position respectively, so direct reading or writing can be performed without any migration. In Scenario 3, Tile Dest is outside the SR of Tile Src, and at the same time, the comparison between the EWNScnts of Tile Src and Tile Dest cannot correspond to the relative position, so migration is performed before direct reading and writing.

[0038] In another embodiment, the present invention proposes a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the multi-stack magnetic cache method for accessing DNN training data in the foregoing embodiment.

[0039] In another embodiment, the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the multi-stack magnetic cache method for accessing DNN training data in the foregoing embodiment is implemented.

[0040] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0041] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0042] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A multi-stack magnetic cache method for DNN training data memory access, which uses sTLC MRAM for memory access, is characterized in that Including: For the weight data of the network, 32-bit single-precision floating-point numbers compliant with the IEEE 754 standard are used for representation and processing, including 1-bit sign bit, 8-bit exponent bit, and 23-bit mantissa bit; at the start of training, the exponent bits of the weight data are fixed; during the training process, 3-bit original code compensation is performed on the exponents of the weight data, and at the same time, 5-bit mantissas are truncated; where the LSB stores the fixed exponent bits, the CSB stores the sign bit, the CSB and MSB store 18-bit mantissas, and the MSB stores 3-bit exponent compensation; For the input data of the network, during the training process, 11 cell units are used to store 32-bit data, and one LSB is left empty; the input data is rearranged into consecutive cell units according to different features.

2. The multi-stack magnetic cache method for DNN training data memory access according to claim 1, characterized in that: It also includes the following cache block division and reorganization operations: Based on the path, the cache sets of the GPU L2 cache are divided into a certain number of cache blocks (Tiles), and the Tiles belonging to the same path are combined into different clusters; Among them, the data transmission between each Tile is realized by means of the on-chip interconnection network.

3. The multi-stack magnetic cache method for DNN training data memory access according to claim 2, wherein: It also includes cache migration operations, which are restricted between different paths within the same cache set.

4. The multi-stack magnetic cache method for DNN training data memory access according to claim 3, characterized in that: The cache migration operation is triggered by a direction counter, and the direction counter is a signed number represented by two's complement and is reset in the following four cases: when the cache line is newly acquired; when the cache set is full and the cache line is evicted by the least recently used algorithm; when the value of the direction counter after migration exceeds the predefined range; After the migration operation itself occurs.

5. The multi-stack magnetic cache method for DNN training data memory access according to claim 4, characterized in that: The update of the direction counter depends on the relative position between the source Tile and the target Tile.

6. The multi-stack magnetic cache method for DNN training data memory access according to claim 5, characterized in that: When the target Tile is outside the silent radius range of the source Tile, and the value of the direction counter between the source Tile and the target Tile cannot correspond to their relative position, the cache migration operation is triggered.

7. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute the multi-stack magnetic cache method for DNN training data access as described in any one of claims 1-6.

8. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the multi-stack magnetic cache method for DNN training data access as described in any one of claims 1-6.