Heterogeneous multi-core system memory access management method for deep neural network

By adopting cache offloading and prefetching methods in heterogeneous multi-core systems, the problems of low memory utilization and large memory access latency in DNN training are solved, and higher cache hit rate and lower memory access latency are achieved, which improves training efficiency.

WO2025097560A1PCT designated stage expired Publication Date: 2025-05-15BEIJING UNIV OF TECH

Patent Information

Application Number
PCT/CN2023/140148
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-09
Filing Date
2023-12-20
Publication Date
2025-05-15

AI Technical Summary

Technical Problem

When deep neural networks (DNNs) are trained in heterogeneous multi-core systems, due to limited GPU memory capacity and low communication efficiency between cores, the memory utilization rate and obvious memory access delay are affected, which affects the computing efficiency.

Method used

A cache offload and prefetch method deployed on heterogeneous multi-core computing system is proposed. By unloading the feature map to DRAM during the forward propagation process and prefetching data back to the last-level cache (LLC) during the back propagation process, it can improve the cache hit rate and reduce the memory access delay.

Benefits of technology

The memory access efficiency is significantly optimized, the hit rate of the last-level cache is improved, the memory access delay caused by high-frequency access requests from LLC to DRAM is reduced, and the efficiency of DNN training is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140148_15052025_PF_FP_ABST
    Figure CN2023140148_15052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer storage system structures. Disclosed is a heterogeneous multi-core system memory access management method for a deep neural network (DNN). According to the present invention, a CPU-GPU heterogeneous multi-core system is used to accelerate training of a DNN model, and a memory access controller is designed on the basis of the memory access characteristics of the CPU-GPU heterogeneous multi-core system; the last level cache shared by multiple cores is unloaded, prefetched, and released in the DNN training process; the data transmission process is fine-grained; and the utilization rate of the last level cache is significantly increased. In addition, a latency hiding mechanism is also designed for the memory access controller, and the access process and the calculation process of a large amount of intermediate data of a feature extraction layer are overlapped, so that the calculation performance loss caused by the last level cache miss and waiting of a DRAM memory access response in the model calculation process is reduced, improving the training efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

A memory access management method for heterogeneous multi-core systems for deep neural networks Technical Field

[0001] The present invention belongs to the field of computer system storage system structure, and specifically relates to a storage structure based on a heterogeneous multi-core system, and a cache unloading and prefetching method deployed in the training process of a deep neural network model. Background Art

[0002] Deep neural networks (DNNs) have been widely adopted in various fields, including computer vision, speech recognition, and natural language processing, due to their outstanding performance. The surge in deep learning applications has led to the emergence of a growing number of software frameworks for analyzing and facilitating neural network processing. As developers continue to add more functionality and improve computational efficiency, the list of available frameworks continues to expand. Because GPUs can significantly accelerate the highly parallel DNN training process, these frameworks provide powerful backend support for GPU software libraries such as cuDNN. Today, nearly every group involved in training neural networks is deploying GPUs to accelerate deep learning.

[0003] A common limitation of current popular machine learning frameworks is that the memory capacity of the system's GPU ultimately limits the size of the DNNs that can be trained. DNN models trained using the stochastic gradient descent algorithm are designed as multi-layered neural networks. Training these neural networks involves a series of layer-by-layer computations, with a fixed order and spanning millions to billions of iterations. Due to the strong data dependencies of stochastic gradient descent's layer-by-layer computations, GPUs can only process computations for a single layer simultaneously during training. To accommodate these computational characteristics, current popular machine learning frameworks generally adopt a network-wide memory allocation strategy, where GPU memory backs up intermediate feature maps from all layers in the network for gradient updates. This strategy often overallocates memory to accommodate the memory usage of the entire network layer. Research by Rhu et al. indicates that this memory underutilization problem becomes more severe for deeper networks, with 53% to 79% of the allocated memory remaining unused during training. To address this memory bottleneck, machine learning practitioners must either use less optimal DNN architectures—fewer layers, smaller batch sizes, convolutional algorithms with lower performance but higher memory efficiency—or consider parallel processing of DNNs on multiple GPUs. These methods will undoubtedly hinder the speed and accuracy of training, thereby reducing the performance of the DNN model.

[0004] To address the limited memory capacity of GPUs and the efficiency of inter-core communication limited by the PCIe bus, many researchers are considering using CPU-GPU heterogeneous computing systems to accelerate deep learning. Heterogeneous systems integrate multiple computing cores and multi-level storage systems on-chip, achieving higher inter-core communication speeds and larger cache and memory spaces, thereby moderately alleviating the memory access pressure of DNN models. However, the improved computing efficiency of heterogeneous multi-core systems has led to new memory access limitations for the systems, and simply expanding available memory capacity still cannot solve the fundamental problem of low memory utilization during DNN training.

[0005] The DNN training process can be broadly divided into two steps: forward propagation and backward propagation. Forward propagation proceeds from the first (input) layer to the last (output) layer, while backward propagation proceeds in the opposite direction. Forward propagation traverses the network layer by layer, performing feature extraction and classification tasks on the given input. During forward propagation, each layer performs mathematical operations on its input feature map X and stores the results as the output feature map Y. The forward propagation computational flow is a serialized process. For a linear feedforward DNN, the Y obtained in layer n-1 will be directly used as the input X in layer n. Due to this inter-layer data dependency, the GPU can only process the computation of a single layer simultaneously during a training cycle. Therefore, the memory allocation required for each layer is determined by the input-output relationship of the layer and its activation function.

[0006] For incompletely trained DNN models, the results of a single round of inference can exhibit significant error. The backpropagation process uses a loss function to derive the magnitude of the inference error at the end of forward propagation. The gradient of the loss function is derived relative to the output of the last layer. The backpropagation process uses the loss function as the input gradient map dY for the nth layer, derives the output gradient map dX according to the chain rule, and passes this to the n-1th layer as the new dY for computation. Because the X and Y values ​​of the current layer are required for calculation, the entire X, Y, dX, and dY of the current layer typically need to be stored. These intermediate data, often collectively referred to as feature maps, occupy more than half of the system storage space in most DNN models, resulting in significant waste of storage resources.

[0007] Summary of the Invention

[0008] To address the problems of low memory utilization and significant impact of memory access latency on computational efficiency caused by the widespread adoption of network-wide memory allocation strategies when running DNNs in heterogeneous multi-core computing systems, the present invention proposes a method for offloading and prefetching the last-level cache (LLC) space used by the feature extraction layer of a DNN, deployed on a heterogeneous multi-core computing system. Feature maps that have completed forward propagation but are still residing in the cache space awaiting gradient updates are offloaded to DRAM via the bus. These data are then prefetched back to the LLC before being called during the backward propagation process, effectively hiding the prefetch latency. This improves the hit rate of the last-level cache during training and reduces the memory access latency caused by high-frequency access requests from the LLC to DRAM.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] Step 1: When the program starts, an area is allocated in the memory area. Its size matches the capacity of the last-level cache. In actual operation, it can be adjusted according to the size of the DNN model.

[0011] In step 2, during each iteration of feature extraction, the algorithm continuously tracks the feature graph X from the forward propagation process. During the forward computation of Y, X written to the LLC is copied to the memory offload area to reduce the additional latency introduced by the transfer. After the transfer is complete, the algorithm tracks inter-layer dependencies in the form of a data flow graph. If the current processing layer is the last consumer of X, X is released from the LLC to free up buffer space.

[0012] In step 3, at the beginning of the backpropagation process, the algorithm prefetches the previous layer's data from the memory offload area, overlapping the computation and transfer operations and maximizing latency hiding. The algorithm synchronously tracks the computation and transfer flows, pausing the current layer's computation if prefetching is incomplete.

[0013] Step 4: During backpropagation, when the gradient update for each layer is complete, the algorithm releases the previous data from both the LLC and the memory unload area, as the current layer's Y and dY have already contributed to the gradient descent calculation as intermediate data. The current layer's X and dX are not released at this time, as the previous layer's backpropagation still requires these values ​​for gradient derivation.

[0014] Compared with the prior art, the present invention has the following advantages:

[0015] Most existing DNN memory management technologies focus on improving communication speed and expanding memory capacity. The on-chip network architecture of the CPU-GPU heterogeneous computing system significantly improves the communication speed between cores and between cores and memory controllers. At the same time, the integrated GPU core can obtain larger cache and memory space than an independent GPU. On this basis, the present invention further optimizes the utilization and hit rate of the last-level cache. Due to the large throughput and randomness of data during DNN training, the LLC in the traditional storage structure usually maintains a high miss rate for a long time. Excessive access requests from LLC to DRAM bring relatively long communication delays, and calculations need to be stopped frequently to wait for data transmission, which reduces training efficiency. The offloading strategy for LLC deployment divides data transmission in a fine-grained manner, overlaps most of the intermediate data access processes and calculation processes of the feature extraction layer, and significantly optimizes memory access efficiency. The prefetch strategy makes the data flow in LLC more effective for the gradient descent calculation process, which can significantly improve the cache hit rate during the backpropagation process and reduce additional memory access delays. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flow chart of the memory access controller in a linear feedforward neural network.

[0017] Figure 2 is a schematic diagram of the delay hiding process.

[0018] Figure 3 is a data flow diagram of a nonlinear feedforward neural network. DETAILED DESCRIPTION

[0019] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] The present invention relates to a cache unloading and prefetching method based on the storage structure of a heterogeneous multi-core system and deployed for the training process of a deep neural network model. The process in a linear feedforward neural network is shown in Figure 1. X / Y in the figure respectively represent the input / output feature graphs of each layer of forward propagation, and dX / dY respectively represent the output / input feature gradient graphs of each layer of reverse propagation. This method adds memory access controller logic to each feature extraction layer of DNN training, which mainly includes a cache unloading process in forward propagation, a cache prefetching process and a memory release process in reverse propagation, and a priority-based random cache replacement strategy, so that DNN can more efficiently utilize the space of the last-level cache (LLC) when training on a heterogeneous multi-core system, reducing the computing performance loss caused by excessive waiting for DRAM memory access responses. The specific steps are as follows:

[0021] Step 1: When loading a training dataset, an offload region of equal size is allocated in DRAM based on the shared LLC capacity in the system. Whenever the memory controller allocates and releases data structures during DNN training, the memory manager allocates and releases memory regions from this region.

[0022] Step 2: When data preprocessing is completed and the forward propagation process begins, the memory access controller monitors the feature map data transmission and calculation process of each layer, and copies the input feature map X to the memory unloading area at the beginning of the forward calculation process of each layer (in the DNN training process, the output feature map Y of the nth layer is copied to the memory unloading area). n , input gradient map dY n Equivalent to the input feature map X of the n+1 layer n+1 , output gradient map dX n+1 , so Y and dY do not occupy additional storage space). The parallel computation and transmission flows during the delay hiding process are shown in Figure 2, which depicts the performance of three layers with different computation time-to-transmission time ratios in the delay hiding mechanism. When the offloading time exceeds the computation time, the forward computation process occupies the transmission flow resources, and the computation of the next layer needs to be paused to wait for the data to be safely offloaded. However, even though there may be a waiting delay, most of the transmission time is still overlapped with the computation process, which effectively reduces the additional transmission cost brought by the offloading mechanism. When the offloading process is completed, the memory controller releases the space of X from the LLC.

[0023] In step 3, before forward propagation begins for each subsequent layer, the memory access controller first evaluates the data dependencies between layers based on the data flow graph. When the DNN model is a feedforward linear network (such as AlexNet), the output feature map Y of the previous layer forms a unique dependency relationship with the input feature map X of the current layer, and the unloading / release process can be directly performed without additional conditions. However, when the DNN model uses a nonlinear feedforward network such as GoogleNet, the situation shown in Figure 3 occurs, where the output Y of the first layer in the network is passed to the second and third layers for calculation. In this case, the memory access controller pre-constructs the model's data flow graph and calculates the number of dependencies of each layer's output feature map Y. Since layers that depend on the same Y share X data, to maximize cache utilization, the unloading / release process is only allowed when it is determined that the current processing layer is the last dependent layer of its predecessor's output feature map Y.

[0024] In step 4, after the backpropagation process begins, the memory controller prefetches the X values ​​required by the previous layer back into the LLC while performing the reverse calculation of the input gradient graph dY at each layer. The parallel computation and transmission flows during the latency hiding process are shown in Figure 2. Similar to the offloading process, if the prefetch transmission time exceeds the computation time, the memory controller pauses the computation of the previous layer to wait for data to be safely prefetched. Because the backpropagation process requires both X and Y at each layer to be calculated, the prefetch of X in this step will not fully cover the memory access requests for the computation process until after two layers. To prevent data transmission during the reverse computation from overriding the prefetched data into the LLC, the present invention employs a random cache replacement strategy based on marked priority. A register is added to each cache block in the LLC to mark its priority. When the memory controller prefetches X, this cache block's flag bit is set to 1. Because the reuse rate of intermediate data during DNN training is very low, a cache block is randomly selected from the buffer during the cache replacement process and the flag bit is evaluated. If the prefetched data marked as 1 is selected, the random selection process is repeated; otherwise, the data is directly replaced at that location. This label replacement strategy can minimize the computation time while ensuring the security of pre-fetched feature map data.

[0025] Step 5: After the gradient update process of each layer is completed, the space occupied by Y and dY of the current layer (i.e., X and dX of the previous layer) in the cache and DRAM unloading area is released.

Claims

1. A memory access management method for a heterogeneous multi-core system for a deep neural network, characterized in that: This method adds memory access controller logic to each feature extraction layer of DNN training, including the cache unloading process in forward propagation, the cache pre-fetching process and memory release process in backward propagation, and the priority-based random cache replacement strategy, so that DNN can more efficiently utilize the space of the last level cache (LLC) when training on a heterogeneous multi-core system, and reduce the computing performance loss caused by too much waiting for DRAM memory access response; specifically, the following steps are included: Step 1: When loading the training dataset, an unloading area of ​​equal size is allocated in DRAM according to the shared LLC capacity in the system; whenever the memory controller allocates and releases data structures during the DNN training process, the memory manager allocates and releases memory areas from this area; Step 2: When data preprocessing is completed and the forward propagation process begins, the memory access controller monitors the feature map data transmission and calculation process of each layer, and copies the input feature map X of each layer to the memory unloading area when the forward calculation process is performed; during the training process of the DNN, the output feature map Y of the nth layer n , input gradient map dY n Equivalent to the input feature map X of layer n+1 n+1 , output gradient map dX n+1 , so Y and dY do not need to occupy additional storage space; when the unloading time exceeds the calculation time, since the forward calculation process occupies the transmission stream resources, the calculation of the next layer needs to be suspended to wait for the data to be safely unloaded; when the unloading process is completed, the memory controller releases the space of X from LLC; Step 3: Before the forward propagation of each subsequent layer begins, the memory access controller first evaluates the data dependency between layers according to the data flow graph; when the DNN model is a feedforward linear network, the output feature map Y of the previous layer and the input feature map X of the current layer form a unique dependency relationship, and the unloading / releasing process can be directly performed without additional conditions; however, when the DNN model uses a nonlinear feedforward network such as GoogleNet, there may be a situation where Y and X of the upper and lower layers do not form a unique dependency; at this time, the memory access controller will pre-construct the data flow graph of the model and calculate the number of dependencies of the output feature map Y of each layer. Since the layers that depend on the same Y share X data, in order to maximize cache utilization, the unloading / releasing process is allowed only when it is determined that the current processing layer is the last dependent layer of its predecessor output feature map Y; Step 4: After the reverse propagation process begins, the memory controller pre-fetches the X value required by the previous layer back to the LLC when performing the reverse calculation process on the input gradient map dY of each layer. Similar to the unloading process, when the pre-fetch transmission time exceeds the calculation time, the memory controller suspends the calculation process of the previous layer to wait for data to be safely pre-fetched. Since the reverse propagation process requires X and Y of each layer to participate in the calculation, the pre-fetch of X in this step will not be able to completely cover the memory access request of the calculation process until after 2 layers. In order to prevent the data transmission in the reverse calculation process from flushing the pre-fetched data into the LLC, a tag-based priority method is adopted. The random cache replacement strategy of the first level adds a register to each cache block in LLC to mark the priority. When the memory controller prefetches X, the cache block mark bit is assigned to 1. Since the reuse rate of intermediate data in the DNN training process is very low, a cache block is randomly selected from the buffer during the cache replacement process and the mark bit is judged. If the prefetched data marked as 1 is selected, the random selection process is repeated, otherwise it is directly replaced at this position. Step 5: After the gradient update process of each layer is completed, the space occupied by Y and dY of the current layer, that is, X and dX of the previous layer, in the cache and DRAM unloading area is released.

Citation Information

Patent Citations

  • Deep learning heterogeneous computing method and system based on layer width memory allocation

    CN109976903A

  • Memory management method and system for deep neural network

    CN116107754A

  • Methods of operating a graphics processing unit (GPU) to train a deep neural network using a GPU local memory and related articles of manufacture

    US20200302304A1

  • Information processing apparatus and memory access control method

    US20230281129A1

Cited By

  • Storage processing method, model running method, calculation module, electronic equipment, storage medium, cluster and program product

    CN120687170A