A Memory Access Management Method for Heterogeneous Multi-core Systems Targeting Deep Neural Networks
By performing cache unloading and prefetching on the feature extraction layer of a deep neural network on a heterogeneous multi-core system, the problems of low memory utilization and memory access latency are solved, thereby improving computational efficiency and training speed.
Patent Information
- Application Number
- CN202311487081.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-11-09
AI Technical Summary
In existing technologies, heterogeneous multi-core computing systems suffer from low memory utilization and memory access latency affecting computational efficiency during deep neural network training. This is especially true when GPU memory capacity is limited and communication efficiency between cores is restricted by the PCIe bus, leading to a decrease in training speed and accuracy.
On a heterogeneous multi-core system, a method is adopted to unload and prefetch the feature extraction layer of a deep neural network using the final level cache (LLC) space. By unloading the feature map to DRAM during forward propagation and prefetching data during backward propagation, the prefetch latency is hidden, thereby optimizing cache utilization and hit rate.
It significantly improved cache hit rate during training, reduced memory access latency, improved computational efficiency and training speed, and optimized memory access efficiency.
Smart Images

Figure CN117591274B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer architecture storage system structure, specifically relating to a storage structure based on a heterogeneous multi-core system, and a cache offloading and prefetching method deployed during the training process of a deep neural network model. Background Technology
[0002] Deep neural networks (DNNs) have been widely applied in various fields such as computer vision, speech recognition, and natural language processing due to their outstanding performance. The surge in deep learning applications has led to the emergence of numerous software frameworks for analyzing and enhancing neural networks. As developers continuously add more features and improve computational efficiency, the list of available frameworks continues to expand. Because GPUs can significantly accelerate the highly parallel training process of DNNs, these frameworks provide powerful backend support for GPU software libraries such as cuDNN. Today, almost every group involved in training neural networks deploys GPUs to accelerate deep learning.
[0003] A common limitation of current popular machine learning frameworks is that the memory capacity of the GPU in the system ultimately limits the size of the DNN that can be trained. DNN models trained using stochastic gradient descent are designed as multi-layered neural networks. The training of these networks involves a series of layer-by-layer computations in a statically fixed order, undergoing millions to billions of iterations throughout the training process. Due to the strong data dependency of the layer-by-layer computations in stochastic gradient descent, GPUs can only process single-layer computations simultaneously during training. To accommodate this computational characteristic, current popular machine learning frameworks generally employ network-wide memory allocation strategies, allowing the GPU to back up intermediate feature maps of all layers in the network for gradient updates. To accommodate the memory usage of the entire network layer, such strategies often over-allocate memory. Research by Rhu et al. indicates that this underutilization of memory becomes more severe for deeper networks, with 53% to 79% of the allocated memory remaining unused during training. To address the memory capacity bottleneck, machine learning practitioners must either use a less-than-ideal DNN architecture—lower layers, smaller batch sizes, and less efficient but more memory-efficient convolutional algorithms—or consider parallel processing of DNNs across multiple GPUs. These methods will undoubtedly hinder the speed and accuracy of training, thereby reducing the performance of the DNN model.
[0004] To address the limitations of GPU memory capacity and the PCIe bus-dependent communication efficiency between cores, many researchers have considered using CPU-GPU heterogeneous computing systems to accelerate deep learning. Heterogeneous systems integrate multiple computing cores and multi-level storage systems on-chip, achieving higher inter-core communication speeds and larger cache and memory spaces, thus moderately alleviating the memory access pressure on DNN models. However, the improved efficiency of heterogeneous multi-core computing introduces new memory access limitations, and simply expanding the available memory still cannot solve the fundamental problem of low memory utilization during DNN training.
[0005] The training process of a DNN can be broadly divided into two processes: forward propagation and backward propagation. Forward propagation proceeds from the first (input) layer to the last (output) layer, while backward propagation proceeds in the opposite direction. Forward propagation traverses the network layer by layer, performing feature extraction and classification tasks on the given input. During forward propagation, each layer performs mathematical operations on its input feature map X and stores the result as the output feature map Y. The computational flow of forward propagation is a sequential process; for a linear feedforward DNN, the Y obtained from the (n-1)th layer will be directly used as the input X by the nth layer. Due to this inter-layer data dependency, a GPU can only process the computation of a single layer at a time during the training cycle. Therefore, the memory allocation required for each layer is determined by the layer's input-output relationship and its activation function.
[0006] For DNN models that are not fully trained, the result of a single round of inference contains a large error. The backpropagation process uses a loss function to derive the magnitude of the inference error at the end of forward propagation; the gradient of the loss function is derived relative to the output of the last layer. The backpropagation process uses the loss function as the input gradient map dY of the nth layer, derives the output gradient map dX according to the chain rule, and passes it to the (n-1)th layer as a new dY for computation. Since the derivation requires the X and Y values of the current layer for calculation, it is usually necessary to store all the X, Y, dX, and dY values of the current layer. This intermediate data is typically referred to as feature maps, and in most DNN models, they occupy more than half of the system storage space, leading to a significant waste of storage resources. Summary of the Invention
[0007] To address the issues of low memory utilization and significant computational efficiency caused by network-wide memory allocation strategies commonly used in heterogeneous multi-core computing systems running DNNs, this invention proposes a method for offloading and prefetching the final-level cache (LLC) space used by the feature extraction layer of a DNN deployed on a heterogeneous multi-core computing system. Feature maps that have completed forward propagation but are still residing in the cache space awaiting gradient updates are offloaded to DRAM via a bus. Before being accessed during backpropagation, this data is prefetched back to the LLC, effectively hiding the prefetch latency. This improves the hit rate of the final-level cache during training and reduces the memory access latency caused by high-frequency access requests from the LLC to DRAM.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] Step 1: When the program starts, it allocates a region in the memory area, the size of which matches the capacity of the last-level cache. In actual operation, it can be adjusted according to the size of the DNN model.
[0010] Step 2: During the feature extraction process of each iteration, the algorithm continuously tracks the feature map X of the forward propagation process. Overlaying this tracking with the time taken to compute Y, X written to the LLC is copied to the memory unloading area to reduce additional latency caused by transmission. After transmission, the algorithm tracks inter-layer dependencies in the form of a data flow graph. When the current processing layer is the last consumer of X, X is released from the LLC to free up buffer space.
[0011] Step 3: At the start of the backpropagation process, the algorithm prefetches data from the previous layer from the memory offload area, overlapping the computation and transport operations to maximize the benefits of latency hiding. The algorithm synchronously tracks the computation and transport flows to pause computation of the current layer if prefetching is incomplete.
[0012] Step 4: During backpropagation, when the gradient update of each layer is complete, since Y and dY of the current layer have already contributed to the gradient descent calculation as intermediate data, the algorithm releases their previous data from the LLC and memory unloading areas. X and dX of the current layer are not released at this time because the backpropagation of the previous layer still requires these values for gradient derivation.
[0013] Compared with the prior art, the present invention has the following advantages:
[0014] Existing DNN memory management technologies mostly focus on improving communication speed and expanding memory capacity. The on-chip network architecture of CPU-GPU heterogeneous computing systems significantly improves communication speed between cores and between cores and the memory controller. Simultaneously, integrated GPU cores can obtain larger cache and memory space compared to discrete GPUs. Building upon this, this invention further optimizes the utilization and hit rate of the final-level cache. Due to the high throughput and randomness of data during DNN training, the LLC in traditional storage structures typically maintains a high miss rate for extended periods. Excessive access requests from LLC to DRAM lead to relatively long communication latency, requiring frequent computation pauses to wait for data transfer, thus reducing training efficiency. The offloading strategy for LLC deployment finely divides data transfer, overlapping most intermediate data access and computation processes in the feature extraction layer, significantly optimizing memory access efficiency. Furthermore, the prefetch strategy makes the data flow in the LLC more efficient for gradient descent computation, significantly improving cache hit rate during backpropagation and reducing additional memory access latency. Attached Figure Description
[0015] Figure 1 This is a flowchart of the memory access controller in a linear feedforward neural network.
[0016] Figure 2 This is a diagram illustrating the delayed hiding process.
[0017] Figure 3 This is the data flow graph for a nonlinear feedforward neural network. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0019] This invention relates to a storage structure based on heterogeneous multi-core systems, and a cache offloading and prefetching method deployed during the training process of deep neural network models. The process in a linear feedforward neural network is as follows: Figure 1 As shown in the figure, X / Y represent the input / output feature maps of each forward propagation layer, and dX / dY represent the output / input feature gradient maps of each backward propagation layer. This method adds memory access controller logic to each feature extraction layer during DNN training. This mainly includes the cache unloading process during forward propagation, the cache prefetching process and memory release process during backward propagation, and a priority-based random cache replacement strategy. This allows the DNN to utilize the last-level cache (LLC) space more efficiently during training on heterogeneous multi-core systems, reducing the computational performance loss caused by excessive waiting for DRAM memory access responses. The specific steps are as follows:
[0020] Step 1: When loading the training dataset, an unloaded region of equal size is allocated in DRAM based on the shared LLC capacity in the system. Whenever the memory access controller allocates and releases data structures during DNN training, the memory manager allocates and releases memory from this region.
[0021] Step 2: When data preprocessing is complete and the forward propagation process begins, the memory access controller monitors the feature map data transfer and computation process of each layer. At the beginning of the forward computation process of each layer, the input feature map X is copied to the memory offload region (during the training of the DNN, the output feature map Y of the nth layer). n Input gradient map dY n The input feature map X is equivalent to layer n+1. n+1 Output gradient map dX n+1 Therefore, Y and dY do not require additional storage space. The parallel computation and transport streams during the delay hiding process are as follows: Figure 2 As shown in the figure, the performance of three layers with different computation time-transfer time ratios in the latency hiding mechanism is illustrated. When the offloading time exceeds the computation time, the computation of the next layer needs to be paused to wait for the data to be safely offloaded because the forward computation process occupies transport stream resources. However, even with the potential waiting delay, most of the transmission time is still overlapped in the computation process, which effectively reduces the additional transmission cost brought by the offloading mechanism. After the offloading process is completed, the memory access controller releases the space of X from the LLC.
[0022] Step 3: Before each subsequent forward propagation layer begins, the memory access controller first evaluates the data dependencies between layers based on the data flow graph. When the DNN model is a feedforward linear network (such as AlexNet), the output feature map Y of the previous layer and the input feature map X of the current layer form a unique dependency, and the unloading / release process can be performed directly without additional conditions. However, when the DNN model uses a non-linear feedforward network like GoogleNet, the following will occur: Figure 3 As shown, the output Y of the first layer in the network is fed into the second and third layers for computation. At this time, the memory access controller will pre-construct the data flow graph of the model and calculate the number of dependencies of the output feature map Y of each layer. Since layers that depend on the same Y share X data, in order to maximize cache utilization, the unloading / release process can only be allowed when it is determined that the current processing layer is the last dependent layer of its predecessor output feature map Y.
[0023] Step 4: After the backpropagation process begins, when the memory access controller performs the backpropagation calculation on the input gradient map dY of each layer, it prefetches the X values required by the previous layer back into the LLC. The parallel computation and transport flows during the delay hiding process are as follows: Figure 2As shown. Similar to the unloading process, when the prefetch transfer time exceeds the computation time, the memory access controller pauses the computation process of the previous layer to wait for safe data prefetching. Since the backpropagation process requires the joint computation of X and Y in each layer, the prefetching of X in this step will not fully cover the memory access requests of the computation process until after layer 2. To prevent the data transfer during the backpropagation process from washing away the data prefetched into the LLC, this invention adopts a random cache replacement strategy based on priority marking. A register bit is added to each cache block in the LLC to mark the priority. When the memory access controller prefetches X, the mark bit of the cache block is assigned to 1. Since the reuse rate of intermediate data in the DNN training process is very low, a cache block is randomly selected from the buffer during the cache replacement process, and the mark bit is judged. If the prefetched data marked as 1 is selected, the random selection process is repeated; otherwise, the replacement is performed directly at that position. This mark replacement strategy can ensure the safety of the prefetched feature map data while minimizing the computation time.
[0024] Step 5: After the gradient update process of each layer is completed, release the space occupied by the cache and DRAM offload area of the current layer's Y and dY (i.e., X and dX of the previous layer).
Claims
1. A memory access management method for heterogeneous multi-core systems oriented towards deep neural networks, characterized in that, This method adds memory access controller logic to each feature extraction layer during DNN training, including cache unloading during forward propagation, cache prefetching and memory release during back propagation, and a priority-based random cache replacement strategy. This enables DNNs to utilize the last-level cache (LLC) space more efficiently when trained on heterogeneous multi-core systems, reducing the computational performance loss caused by excessive waiting for DRAM memory access responses. Specifically, it includes the following steps: Step 1: When loading the training dataset, allocate an unloaded region of equal space in DRAM based on the shared LLC capacity in the system; whenever the memory access controller allocates and releases data structures during DNN training, the memory manager will allocate and release memory regions from this region. Step 2: When data preprocessing is complete and the forward propagation process begins, the memory access controller monitors the data transfer and computation process of feature maps for each layer. During the forward computation of the input feature map X for each layer, it copies it to the memory offload region. During the DNN training process, the output feature map Y of the nth layer... n Input gradient map dY n The input feature map X is equivalent to layer n+1. n+1 Output gradient map dX n+1 Therefore, Y and dY do not require additional storage space; when the unloading time exceeds the computation time, the computation of the next layer needs to be paused to wait for the data to be safely unloaded because the forward computation process occupies the transport stream resources; after the unloading process is completed, the memory access controller releases the space of X from the LLC; Step 3: Before each subsequent forward propagation begins, the memory access controller first evaluates the data dependencies between layers based on the data flow graph. When the DNN model is a feedforward linear network, the output feature map Y of the previous layer and the input feature map X of the current layer form a unique dependency relationship, and the unloading / release process can be performed directly without additional conditions. However, when the DNN model uses a non-linear feedforward network like GoogleNet, there may be cases where Y and X of the upper and lower layers do not form a unique dependency relationship. In this case, the memory access controller will pre-construct the model's data flow graph and calculate the number of dependencies of the output feature map Y of each layer. Since layers that depend on the same Y share X data, in order to maximize cache utilization, the unloading / release process can only be allowed when it is determined that the current processing layer is the last dependent layer of its predecessor output feature map Y. Step 4: After the backpropagation process begins, the memory access controller prefetches the required X values from the previous layer back into the LLC during the backpropagation process of the input gradient map dY at each layer. Similar to the unloading process, when the prefetch transmission time exceeds the computation time, the memory access controller pauses the computation process of the previous layer to wait for safe data prefetching. Since the backpropagation process requires the joint participation of X and Y at each layer, the prefetching of X in this step will not fully cover the memory access requests of the computation process until after two layers. To prevent the data transmission during the backpropagation process from washing away the data prefetched into the LLC, a random cache replacement strategy based on priority marking is adopted. A register bit is added to each cache block in the LLC to mark the priority. When the memory access controller prefetches X, the mark bit of the cache block is set to 1. During the cache replacement process, a cache block is randomly selected from the buffer and the mark bit is judged. If the prefetched data marked as 1 is selected, the random selection process is repeated; otherwise, the replacement is performed directly at that position. Step 5: After the gradient update process of each layer is completed, release the space occupied by the cache and DRAM offload area of Y and dY of the current layer, i.e. X and dX of the previous layer.
Citation Information
Patent Citations
Neural network method and device, electronic equipment and storage medium
CN114118397A
Calculation unloading and resource allocation method suitable for CPU-GPU heterogeneous cluster
CN115442851A