A Deep Learning Data Management Method on a Hybrid Memory Multi-Core CPU System
A mixed memory system with Cache-ZVC compression effectively manages deep learning data by storing short-term data in DRAM and compressing long-term data in non-volatile memory, addressing memory constraints and enhancing training efficiency for complex deep neural networks.
Patent Information
- Application Number
- CN202111510365.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The memory demand of deep learning training systems is too high, making it difficult to support the ever-increasing deep neural network model, resulting in inefficient training.
The hybrid memory multi-core CPU system is adopted to store short-term data using DRAM memory, and the non-volatile memory stores long-term data, and the long-term data is compressed through the Cache-ZVC compression algorithm to reduce the amount of data written in non-volatile memory.
Improves the scalability of deep neural networks, enables more efficient training of more complex models, and reduces bandwidth pressure of non-volatile memory.
Smart Images

Figure CN114218127B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management, and particularly to a deep learning data management method on a hybrid memory multi-core CPU system. Background Art
[0002] In recent years, with the increasing popularity of the research and application of artificial intelligence, deep learning, as its core technology, has played an increasingly important role in academic research and practical applications based on deep neural network models. With the development of deep learning, the number of layers and complexity of deep neural network models have continuously increased. At the same time, the scale of the training sample data set has also gradually increased. The increase in the number of layers and complexity of the neural network model and the increase in the scale of the training sample data set have led to a significant increase in data such as model parameters, intermediate results between layers, and temporary buffers implemented by network layers, resulting in the problem of excessive memory overhead in deep learning training. Therefore, it has gradually become an urgent problem that the limited memory of the deep learning training system is difficult to support the deep neural network model with an increasingly large memory requirement. Summary of the Invention
[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a deep learning data management method on a hybrid memory multi-core CPU system, which alleviates the pressure of the increasing memory demand brought about by the rapid development of deep learning, and makes it easier to efficiently train more complex deep neural networks under the background of the continuous increase in the scale of the data set.
[0004] The technical solution adopted by the present invention is: a deep learning data management method on a hybrid memory multi-core CPU system, including the following steps:
[0005] In the forward propagation stage of model training, each network layer reads the data in the temporary memory buffer, performs the forward calculation of the layer, and obtains the forward propagation output data;
[0006] Divide the forward propagation output data into short-term data and long-term data;
[0007] Save the short-term data into the temporary memory buffer;
[0008] Perform Cache-ZVC compression on the long-term data and save the compressed data into the non-volatile memory.
[0009] Furthermore, it further includes:
[0010] In the backward propagation stage of model training, each network layer reads the backward propagation error in the temporary memory buffer;
[0011] Read the compressed data and perform the backward calculation of the layer to obtain the backward propagation output data;
[0012] Save the backpropagation output data in a temporary memory buffer.
[0013] Further, the step of saving the short-term data to the temporary memory buffer specifically includes:
[0014] Create a temporary memory buffer in the DRAM memory and save the short-term data to this temporary memory buffer;
[0015] The short-term data output by the current network layer serves as the input data for the next network layer.
[0016] Further, the step of performing Cache-ZVC compression on the long-term data and saving the compressed data to the non-volatile memory specifically includes:
[0017] The long-term data includes a mask and a data block to be compressed;
[0018] Set the size of the data block to be compressed in the long-term data to be the same as the cache block size;
[0019] When the data block to be compressed is still in the cache after the network layer calculation is completed, compress the data block to be compressed to obtain a compressed data block;
[0020] Use the non-temporal store instruction in the CPU instruction set to write the mask and the compressed data block to the non-volatile memory.
[0021] Further, the mask is used to record the positions of non-zero values in the data block, and the compressed data block contains the non-zero values of the data block.
[0022] Further, the step of reading the compressed data and performing the backward calculation of this layer to obtain the backpropagation output data specifically includes:
[0023] Read the compressed data and perform Cache-ZVC decompression to obtain decompressed data;
[0024] According to the decompressed data and the backpropagation error in the temporary memory buffer, perform the backward calculation of this layer to obtain the backpropagation output data;
[0025] The backpropagation output data output by the current network layer serves as the input data for the next network layer.
[0026] The beneficial effects of the method of the present invention are as follows: Through the hybrid memory system composed of non-volatile memory and DRAM memory, in the context of the continuous growth of the training sample data set scale, the scalability of the deep neural network can be increased, and more complex deep neural network models can be trained more efficiently. By using the efficient Cache-ZVC compression algorithm to compress the long-term data, the amount of data written to the non-volatile memory can be reduced. Brief Description of the Drawings
[0027] Figure 1 is a flowchart of the steps of a deep learning data management method on a hybrid memory multi-core CPU system of the present invention;
[0028] Figure 2 is a schematic diagram of the data to be compressed in a specific embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of forward propagation compression in a specific embodiment of the present invention;
[0030] Figure 4 is a schematic diagram of backpropagation decompression in a specific embodiment of the present invention;
[0031] Figure 5 is a schematic diagram of long-term data compression in a specific embodiment of the present invention;
[0032] Figure 6 is a schematic diagram of long-term data compression in parallel computing of a multi-core CPU in a specific embodiment of the present invention;
[0033] Figure 7 is a schematic diagram of a deep learning data management method on a hybrid memory system in a specific embodiment of the present invention. Detailed Description of the Embodiments
[0034] The following further elaborates the present invention in detail with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0035] A deep neural network model is composed of multiple network layers, including layers with trainable parameters such as convolutional layers and fully connected layers, and also includes layers without parameters such as pooling layers and normalization layers. There are intermediate results between the layers that make up the neural network, which save the calculation results of the previous layer and are used for the calculation of the layer after that. The deep learning training process includes two stages: forward propagation and backpropagation. Forward propagation starts from the training sample data being input into the neural network model, and an error is obtained after the forward calculation of each network layer; then backpropagation begins, and this error is input back into the neural network model in reverse, and each network layer performs reverse calculation and updates the trainable parameters of this network layer.
[0036] Referring to Figure 1 and Figure 7 , the present invention provides a deep learning data management method on a hybrid memory multi-core CPU system, which stores short-term data in DRAM memory and compresses and stores long-term data in non-volatile memory during the deep learning training process, including:
[0037] Forward propagation stage of deep neural network model training:
[0038] S1. Each network layer reads the data in the temporary memory buffer, performs the forward calculation of this layer, and obtains the forward propagation output data;
[0039] S2. Divide the forward propagation output data into short-term data and long-term data;
[0040] The deep learning framework first parses the definition file of the neural network model, and then constructs a data structure describing the neural network model in memory. At this time, it is necessary to specify which data are long-term data and which are short-term data. Long-term data refers to data that will not be immediately used or will only be used once in a short period of program operation, but will be used after a certain interval; while short-term data are data that will be used soon during the program's calculation process, and this type of data is generally also used frequently.
[0041] S3. Save the short-term data to the temporary memory buffer;
[0042] Specifically, create a temporary memory buffer in the DRAM memory and save the short-term data to this temporary memory buffer. The short-term data output by the current network layer serves as the input data for the next network layer.
[0043] S4. Perform Cache-ZVC compression on the long-term data and save the compressed data to non-volatile memory.
[0044] S4.1. The long-term data includes a mask and a data block to be compressed;
[0045] The mask is used to record the positions of non-zero values in the data block, and the compressed data block contains the non-zero values of the data block. Among them, referring to Figure 2 , the mask consists of the same number of bits as the number of floating-point numbers in the data block to be compressed. A bit set to 1 indicates that the floating-point number corresponding to this bit is non-zero, and a bit set to 0 indicates that the corresponding floating-point number is a zero value. Since all zero values in the data block to be compressed are compressed into a 0 bit in the mask, the zero values in the data block to be compressed can be removed, leaving only the non-zero values, and then saved to the compressed data block.
[0046] S4.2. Set the size of the data block to be compressed in the long-term data to be the same as the cache block;
[0047] S4.3. When the data block to be compressed is still in the cache after the network layer calculation is completed, compress the data block to be compressed to obtain the compressed data block;
[0048] S4.4. Write the mask and the compressed data block to the non-volatile memory using the non-temporal store instruction in the CPU instruction set.
[0049] In view of the computing characteristics of the multi-core CPU system, and the characteristics that the read and write performance of the Optane non-volatile memory is asymmetric and the write performance is relatively poor in a multi-threaded parallel environment, the present invention designs a Cache-ZVC (ZeroValue Compressing in Cache) compression algorithm for long-term data. Before writing the long-term data to the Optane non-volatile memory, the long-term data is compressed to reduce the amount of data written to the Optane non-volatile memory and relieve the write bandwidth pressure of the Optane non-volatile memory.
[0050] During the process of the CPU executing network layer calculations, data is stored in the CPU cache cache in blocks, and the data within the block is reused using the temporal locality and spatial locality of program calculations, so as to reduce memory access, reduce data access latency, and thus accelerate the core calculations of the CPU. The present invention selects the timing of data compression after the execution of the layer calculation of the data block is completed and before the calculation result (to be compressed) is written to the memory; the purpose of selecting the timing in this way is to avoid reading the data back to the CPU for compression again after it is written back to the memory. Frequent data movement will result in low efficiency and make the compression not worthwhile. The compression of long-term data is as Figure 6 shown.
[0051] The compression of long-term data in multi-core CPU parallel computing is as Figure 7 shown. For the compression algorithm in a multi-threaded parallel environment, the needs of multiple CPU cores for compressing data simultaneously need to be taken into account. The present invention adopts the method of separately storing the mask and the compressed data block in the Optane non-volatile memory. Since the mask is fixed-length data and the compressed data block is variable-length data, and its size depends on the specific number of non-zero values in the long-term data, the block composed of the mask and the compressed data block is variable-length data as a whole, which is not conducive to parallel reading and writing. Therefore, for the convenience of compression and decompression, the mask and the compressed data block are stored separately. The mask is stored in the fixed-length part of the compressed data, and the compressed data block is stored after the fixed-length part of the compressed data. During multi-threaded compression, each thread can process a data block to be compressed, and then write the mask and the compressed data block of the block to the Optane non-volatile memory respectively; during multi-threaded decompression, each thread reads the mask of its own block, and then reads the compressed data block from the Optane non-volatile memory according to the mask and restores it to the data block before compression.
[0052] The Cache-ZVC compression algorithm of the multi-core CPU system designed by the present invention makes full use of the characteristic that there are relatively many zero values in the long-term data during the deep learning training process, and meets the relatively high requirements for the data compression processing speed in the form of Cache compression, which can reduce the amount of data written to the Optane non-volatile memory in a multi-threaded parallel environment and reduce its bandwidth pressure.
[0053] The backpropagation stage of the deep neural network model training:
[0054] S5. Each network layer reads the backpropagation error in the temporary memory buffer;
[0055] S6. Read the compressed data and perform the backward calculation of this layer to obtain the backpropagation output data;
[0056] S6.1. Read the compressed data and perform Cache-ZVC decompression to obtain the decompressed data;
[0057] S6.1.1. Read the mask;
[0058] S6.1.2. According to the number of bits with value 1 in the mask, read the corresponding number of floating-point numbers from the data;
[0059] S6.1.3. Restore the compressed data to the data before compression according to the positions of non-zero values recorded by the mask.
[0060] For decompression, each physical core of the CPU reads the mask and the compressed data block from the Optane non-volatile memory respectively, performs Cache-ZVC decompression to obtain a cache block, and then performs the backward calculation on this cache block.
[0061] S6.2. According to the decompressed data and the backpropagation error in the temporary memory buffer, perform the backward calculation of this layer to obtain the backpropagation output data;
[0062] S6.3. The backpropagation output data output by the current network layer is used as the input data of the next network layer.
[0063] S7. Save the backpropagation output data in the temporary memory buffer.
[0064] Specifically, during the backpropagation process, the network layer performs the backward calculation, which is also carried out in units of cache blocks. Before the backward calculation is executed, the mask and the compressed data block need to be read from the Optane non-volatile memory into the cache, the data is decompressed to obtain the cache block, and then together with the backpropagation error, it is used as the input data for the network layer backward calculation.
[0065] The specific process of deep learning training is as follows: read a batch of data from the training sample dataset of deep learning, input it into the deep neural network model, and then start the forward propagation stage. Each network layer reads its input data, that is, the calculation result of the previous layer, and then performs calculations to obtain the output data of the layer. Since the calculations of network layers in the CPU system are carried out in units of cache blocks, the scheduling of long-term data is as Figure 3 shown. The network layer reads a cache block, performs layer calculations and then compresses it, and saves the compressed cache block into the Optane non-volatile memory; at the same time, since the next network layer in the forward propagation still needs to read the output data of the current layer, the uncompressed output data of the layer is also saved into a temporary memory buffer in the DRAM memory in units of cache blocks. This temporary memory buffer is reusable. After the current layer's calculations are completed, the next layer can reuse this temporary memory buffer when performing calculations.
[0066] When the forward propagation reaches the last layer of the deep neural network model, an error data will be obtained. In the second half of the deep learning training, the backpropagation needs to reverse this error data into the deep neural network model, perform the reverse calculations of the network layers, and then update the trainable parameters of the network layers. The reverse calculations of network layers in the CPU system are also carried out in units of cache blocks, and the data scheduling during the calculation process is as Figure 4 shown. Since the reverse calculations of network layers need to use long-term data, it is necessary to read cache blocks from the Optane non-volatile memory and decompress them. At the same time, read the backpropagation error from the next layer saved in the temporary memory buffer, perform the reverse calculations of the network layers, and save the results into the temporary memory buffer.
[0067] As Figure 2 shown, a deep learning data management system on a hybrid memory multi-core CPU system includes:
[0068] A forward propagation module. In the forward propagation stage of model training, each network layer reads the data in the temporary memory buffer, performs the forward calculations of the layer, and obtains the forward propagation output data; divides the forward propagation output data into short-term data and long-term data; saves the short-term data into the temporary memory buffer; performs Cache-ZVC compression on the long-term data and saves the compressed data into the non-volatile memory.
[0069] A backpropagation module. In the backpropagation stage of model training, each network layer reads the backpropagation error in the temporary memory buffer; reads the compressed data and performs the reverse calculations of the layer to obtain the backpropagation output data; saves the backpropagation output data into the temporary memory buffer.
[0070] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented in the system embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0071] A deep learning data management device on a hybrid memory system:
[0072] At least one processor;
[0073] At least one memory for storing at least one program;
[0074] When the at least one program is executed by the at least one processor, the at least one processor implements the deep learning data management method on a hybrid memory multi-core CPU system as described above.
[0075] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented in the device embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0076] A storage medium storing instructions executable by a processor, characterized in that: the instructions executable by the processor are used to implement the deep learning data management method on a hybrid memory multi-core CPU system as described above when executed by the processor.
[0077] The content in the above method embodiments is applicable to the storage medium embodiments of the present invention. The functions specifically implemented in the storage medium embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0078] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A deep learning data management method on a hybrid memory multi-core CPU system, characterized in that It includes the following steps: In the forward propagation stage of model training, each network layer reads the data in the temporary memory buffer, performs the forward calculation of this layer, and obtains the forward propagation output data; Divide the forward propagation output data into short-term data and long-term data; Save the short-term data into the temporary memory buffer; Perform Cache-ZVC compression on the long-term data and save the compressed data into the non-volatile memory; The step of performing Cache-ZVC compression on the long-term data and saving the compressed data into the non-volatile memory specifically includes: The long-term data includes a mask and a data block to be compressed; Set the size of the data block to be compressed in the long-term data to be the same as the cache block; When the data block to be compressed is still in the cache after the network layer calculation is completed, compress the data block to be compressed to obtain the compressed data block; Use the non-temporal store instruction in the CPU instruction set to write the mask and the compressed data block to the non-volatile memory.
2. The deep learning data management method on a hybrid memory multi-core CPU system according to claim 1, wherein It also includes: In the backward propagation stage of model training, each network layer reads the backward propagation error in the temporary memory buffer; Read the compressed data and perform the backward calculation of this layer to obtain the backward propagation output data; Save the backward propagation output data in the temporary memory buffer.
3. The deep learning data management method on a hybrid memory multi-core CPU system according to claim 2, wherein The step of saving the short-term data into the temporary memory buffer specifically includes: Create a temporary memory buffer in the DRAM memory and save the short-term data into this temporary memory buffer; The short-term data output by the current network layer is used as the input data of the next network layer.
4. The deep learning data management method on a hybrid memory multi-core CPU system according to claim 3, characterized in that The mask is used to record the positions of non-zero values in the data block, and the compressed data block contains the non-zero values of the data block.
5. The deep learning data management method on a hybrid memory multi-core CPU system according to claim 4, wherein, The step of reading the compressed data and performing the backward calculation of this layer to obtain the backward propagation output data specifically includes: Read the compressed data and perform Cache-ZVC decompression to obtain the decompressed data; According to the decompressed data and the backward propagation error in the temporary memory buffer, perform the backward calculation of this layer to obtain the backward propagation output data; The backward propagation output data output by the current network layer is used as the input data of the next network layer.
6. The deep learning data management method on a hybrid memory multi-core CPU system according to claim 5, wherein, The step of reading the compressed data and performing Cache-ZVC decompression specifically includes: Read the mask; According to the number of bits with a value of 1 in the mask, read the corresponding number of floating-point numbers from the data; Restore the compressed data to the data before compression according to the positions of non-zero values recorded by the mask.
Citation Information
Patent Citations
Method and device for managing heterogeneous hybrid memory
CN104239225A
Intelligent model training memory allocation method and apparatus, and computer-readable storage medium
WO2020248365A1