Key value storage and data prefetching method based on Caffe application features
By adopting key-value storage and data prefetching methods based on Caffe application features in Caffe model training, the problems of resource waste and overfitting are solved, and more efficient resource utilization and higher testing accuracy are achieved.
Patent Information
- Application Number
- CN202510425804.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In the prior art, there are problems of resource waste and overfitting during the training of Caffe model, resulting in poor testing accuracy.
The key-value storage and data prefetching method based on Caffe application characteristics is adopted. By adjusting the data block size and the dual-thread parallel execution strategy, disk I/O and CPU resource utilization are optimized, and data is read sequentially by using a global out-of-order reading mechanism and iterator pointer pointer to ensure that the order of input instances is different for each round.
It improves the resource utilization rate of computer systems, reduces data prefetching time, reduces the risk of model overfitting, and improves the test accuracy and training performance of the model.
Smart Images

Figure CN119937935A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer storage, and in particular to a key-value storage and data pre-fetching method based on Caffe application features. Background Art
[0002] Modern data storage design is driven by three fundamental trends: 1) applications are storing more and more data; 2) much data performs more inserts than reads during its lifetime; and 3) the popularity of cloud-based data management has driven the development of immutable systems. Combined with these trends, a large number of applications choose to adopt LSM-tree-based key-value storage systems in their storage layer, such as LevelDB, RocksDB, and Cassandra. The main idea of the LSM tree is to maintain a batch of random writes in a memory buffer, and when the buffer is full, these key-value pairs are flushed to the storage device in the form of files to achieve high write throughput. In the storage device, the LSM tree maintains a hierarchical structure, and users can add levels and components as needed. Unlike traditional indexes that perform in-place updates, LSM trees perform out-of-place updates, convert multiple random disk I / Os into sequential disk I / Os, and continuously reorganize data through merge operations to achieve better space utilization.
[0003] Data reading in the LSM tree includes point query and range query operations. The search order of point query is similar to the order of data writing. First, search in the memory table, then search in the immutable memory table, and finally search from L0 to L6 layer by layer until the required data is found or the data does not exist. Compared with point query, range query operation is more complicated. First, read the first data block of all SST files in the L0 layer and the first SST file in the L1 to L6 layers into the memory, and build iterator pointers for them and the memory table and immutable memory table, pointing to the first data of each file, and then create the minimum pointer and the current pointer. Then, starting from the pointer pointing to the memory table, compare with the minimum pointer in hierarchical order. In each comparison process, the minimum pointer always points to the data with a smaller key. After the comparison is completed, the current pointer points to the minimum pointer, and the user reads the data pointed by the current pointer. The comparison process is repeated until all the data in the query range is read.
[0004] The deep learning framework Caffe is used for model training and uses LevelDB as its underlying storage engine. Because of the sequential data storage feature of LevelDB, the image dataset is stored in the key-value storage system in key order (i.e., write order). During model training, Caffe uses two threads for parallel execution, one thread for the forward and reverse transfer of the model, and the other thread for data prefetching. During data prefetching, the processing of each image data mainly includes three steps: reading key-value pairs, parsing key-value pairs, and converting data formats. Reading key-value pairs refers to reading a key-value pair from LevelDB, parsing key-value pairs refers to deserializing the value from a string into a Datum object, and converting data formats refers to data conversion of Datum objects, such as mirroring, scaling, and cropping. Since the disk I / O unit of LevelDB is only one data block (64MB), and the key-value pairs corresponding to image data are large in size, such as the key-value pair size corresponding to a 28*28 pixel image is 3KB, so a disk I / O can only read a small amount of data, so multiple disk I / O operations are required for a round of model training. Caffe uses global range query operations to read all data. After constructing the iterator pointer, it needs to compare the key value size level by level starting from the memory table. However, because the smallest key value is always at the highest level, each read key-value pair operation requires multiple pointer comparisons. In the worst case, a read operation requires 12 comparisons. At the beginning of each training round, Caffe calls the SeektoFirst function of LevelDB to reposition the pointer to the first data of each level and read the data in key order. Therefore, during the model training process, the order of input training instances in each round is the same. The model may use this input mode as a method to reduce training loss, which will be more likely to cause overfitting, resulting in poor test accuracy. At the same time, because the three steps of reading key-value pairs, parsing key-value pairs, and converting data formats are executed serially, only one of the I / O resources and CPU resources can be used at the same time, resulting in serious resource waste and increasing the waiting time of the data pre-fetching process. Summary of the invention
[0005] The main purpose of the present invention is to overcome the above-mentioned defects in the prior art and propose a key-value storage and data prefetching method based on Caffe application characteristics, which improves the resource utilization of computer system, reduces data prefetching time, and inputs training instances in different orders in each round of Caffe training model, reduces model overfitting, improves model testing accuracy, and achieves improved model training performance.
[0006] The present invention adopts the following technical solution:
[0007] A key-value storage and data pre-fetching method based on Caffe application characteristics, applied in a key-value storage system, comprising:
[0008] Set the data block size in the sort string table SST file according to the bandwidth size of different storage devices;
[0009] When image data is written into the key-value storage system, the image data is converted into key-value pairs and saved in the memory table. If the capacity of the memory table reaches the preset threshold, the image data is converted into an immutable memory table and refreshed to the disk, organized into an SST file. Each SST file contains multiple data blocks for storing key-value pair data.
[0010] When performing data pre-fetch operations, the disk I / O unit size to be read is matched with the bandwidth of the storage device. At the same time, when the image resolution is greater than the preset resolution, a dual-thread parallel execution strategy is adopted, in which one thread is responsible for reading and parsing key-value pairs and the other thread performs data format conversion. When reading key-value pairs, a global out-of-order reading mechanism is adopted.
[0011] The implementation process of the global out-of-order reading mechanism is as follows:
[0012] Read the first data block of all SST files in the L0 layer into memory;
[0013] Use a random sorter to randomly sort the SST files of each level from L1 to L6 in the LSM tree and record them in the random order table. Then load the first data block of the first SST file of each level from L1 to L6 in the random order table into the memory, and create corresponding iterator pointers and current pointers for all data blocks in the memory; the iterator pointer points to the first key-value pair in the data block of each level in turn, and the current pointer points to the iterator pointer cyclically; among them, each iterator pointer corresponds to number 1~n;
[0014] Caffe reads data according to the data pointed to by the current pointer; when all data in an SST is read, the next SST file of the same level is read from the random order table, and the reading operation continues until all data in the key-value storage system is read;
[0015] For the next round of data reading, the iterator and random order table are rebuilt to ensure that the order of input instances in each round is different.
[0016] Preferably, one of the two threads is a first thread, and the other is a second thread; the implementation process of the dual-thread parallel execution strategy is as follows:
[0017] When the first thread completes reading and parsing image i, the data will be written into the first buffer, and then the second thread will convert the data in the first buffer; at the same time, the first thread starts to read and parse image i+1 and write it into the second buffer, and the second thread starts to convert the data in the second buffer after processing the data of image i; the two buffers are used alternately, and the mutex lock is used to ensure the safe access to the buffer.
[0018] Preferably, the first buffer and the second buffer are both located in the memory, and are used to temporarily store the image data read and parsed by the first thread, and the data is cyclically allocated between the two buffers through the index variable j. When a thread is using a buffer, a mutex lock is used to ensure that other threads cannot modify the data during access to the buffer.
[0019] Preferably, the random sorter is located in the memory, and is used to read the index information of all SST files in layers L1 to L6, and disrupt the order of files in each layer through a random number generation mechanism.
[0020] Preferably, the random order table is located in the memory and includes two fields: level number and file order; wherein the level number is used to record the numbers of layers L1 to L6, and the file order is used to save the order of files at each level after random sorting.
[0021] Preferably, the preset resolution is 64×64 pixels.
[0022] Preferably, the data pre-fetching operation specifically includes:
[0023] S101, the first thread sends a read request to the key-value storage system, to S102;
[0024] S102, determine whether there is a memory table and an immutable memory table, if so, go to S103, otherwise, go to S104;
[0025] S103, create table iterators for the memory table and the immutable memory table, and then go to S104;
[0026] S104, create table iterators for all SST files in L0, and read the first data block of all files into memory, and then go to S105;
[0027] S105, respectively read the index information of all SST files in layers L1 to L6, the random sorter randomly sorts the order of files in layers L1 to L6, and writes them into the random order table, and then goes to 106;
[0028] S106, create a level iterator for L1 to L6 layers according to the file order in the random order table, and read the first data block in the first SST file of each layer into the memory, and then go to S107;
[0029] S107, merge all table iterators and level iterators into a global range query iterator, and create n iterator pointers and a current pointer, the iterator pointers point to the data blocks in sequence, the current pointer points to the first iterator pointer, and then points to S108;
[0030] S108, the first thread reads the key-value pair pointed to by the current pointer, parses it and writes it into the buffer, and then goes to S109;
[0031] S109, the second thread converts the format of the data in the buffer and writes it into the array. At the same time, the first thread executes the Next operation, and the iterator pointer moves backward, and then moves to S110;
[0032] S110, determine whether there is any unread data in the system, if so, go to S111, otherwise go to S114;
[0033] S111, point the current pointer to the next iterator pointer with unread data, and then point to S112;
[0034] S112, determine whether the number of data in the array reaches a batch size, if so, go to S113, otherwise go to S108;
[0035] S113, forward pass the array to other layers in the model, and create a new array, and then go to S108;
[0036] S114, determine whether the next round of data reading needs to be performed, if so, go to S101, otherwise go to S115;
[0037] S115, data pre-fetching is completed, and the global range query iterator is deleted.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] (1) The present invention adjusts the data block size so that the disk I / O operation can more effectively utilize the computer system bandwidth; at the same time, dual threads are used to perform data pre-fetching in parallel, thereby avoiding CPU idle waiting during the reading operation and disk I / O idle waiting when performing data conversion, thereby achieving full utilization of computer system resources and improving computer system resource utilization;
[0040] (2) The present invention increases the data block size so that more data can be read in one disk I / O, thereby reducing the number of disk I / Os. In addition, dual threads are used to perform data pre-fetching operations in parallel, and the reading and parsing steps are parallelized with the data conversion steps, which significantly shortens the data pre-fetching time. At the same time, iterator pointers are used to read data sequentially, avoiding unnecessary pointer comparisons, thereby further improving the data pre-fetching efficiency.
[0041] (3) The present invention realizes random input of data when training the Caffe model by randomly selecting SST files and using iterator pointers to read data sequentially, thereby improving the stability of the model, reducing the risk of overfitting, and ultimately improving the test accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A system structure diagram of a key-value storage and data pre-fetching method based on Caffe application features of the present invention;
[0043] Figure 2 A schematic diagram of a data pre-fetching process of a key-value storage and data pre-fetching method based on Caffe application features of the present invention;
[0044] Figure 3 A schematic diagram of out-of-order reading of a key-value storage and data pre-fetching method based on Caffe application features of the present invention. DETAILED DESCRIPTION
[0045] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.
[0046] The key-value storage and data pre-fetching method based on Caffe application characteristics proposed in the present invention includes the following steps.
[0047] First, bandwidth tests are conducted for different storage devices, and the data block size in the sorted string table (SST) file is set according to the bandwidth size. When image data is written to the key-value storage system, it is first converted into a key-value pair and saved in the memory table. If the memory table capacity reaches the preset threshold, it is converted into an immutable memory table and refreshed to the disk to organize it into an SST file. Each SST file contains multiple data blocks for storing key-value pair data. When performing data prefetch operations, the read disk I / O unit size is matched with the bandwidth of the storage device, so as to make full use of the system bandwidth and transmit more data in the same transmission time. At the same time, the data prefetch operation adopts a dual-thread parallel execution strategy: one thread is responsible for reading and parsing key-value pairs, and the other thread performs data format conversion. Specifically, when the first thread (thread 1) completes the reading and parsing of image i, the data will be written to the first buffer (buffer 1), and then the second thread (thread 2) converts the data in the first buffer; at the same time, the first thread starts to read and parse image i+1 and writes it to the second buffer (buffer 2). After processing the data of image i, the second thread starts to convert the data in the second buffer. The two buffers are used alternately, and a mutex lock is used to ensure safe access to the buffers. When reading key-value pairs, a global out-of-order reading mechanism is used. First, the first data block in all SST files in the L0 layer is read into the memory. Then the random sorter is used to randomly sort the SST files of each level from L1 to L6 in the LSM tree and record them in the random order table. Then the first data block of the first SST file of each level from L1 to L6 in the random order table is loaded into the memory, and an iterator pointer (1-n) and a current pointer are created for all data blocks in the memory. The iterator pointer points to the first key-value pair in the data block of each level in turn, and the current pointer points to the iterator pointer cyclically. Caffe reads according to the data pointed to by the current pointer when training the model. When all the data in an SST is read, the next SST file of the same level is read from the random order table, and the reading operation continues until all the data in the key-value storage system is read. For the next round of data reading, the iterator and random order table will be rebuilt to ensure that the order of input instances is different in each round. Finally, in order to adapt to image datasets of different resolutions, an adaptive optimization mechanism is used. When the image resolution is less than 64×64 pixels, the dual-thread parallel strategy is not used; otherwise, dual-thread parallel execution of data prefetching operations is enabled.
[0048] It should be noted that Caffe is a deep learning framework that is oriented to image classification and image segmentation and supports CNN, RCNN, LSTM and fully connected neural network designs. Since the invention point involved in the present invention is the key-value storage system called by Caffe and the way it reads data, the present invention can achieve out-of-order reading as long as it is a model that can be deployed and trained in Caffe. The specific model can be LeNet, ResNet50, CaffeNet, CIFAR10_full, etc.
[0049] See also Figure 1 As shown, this is a system structure diagram of a key-value storage and data pre-fetching method based on Caffe application characteristics proposed by the present invention, including a series of memory and disk components, such as a random sorter, a random sequence table, a buffer, an iterator pointer, a current pointer, an SST file, and a data block. When training a model, data pre-fetching is required to speed up the model training efficiency. The data pre-fetching step includes four steps: 1) The first thread (thread 1) initiates a key-value pair read operation on the key-value storage system; 2) Build a global range query iterator, read the first data block of all SST files in the L0 layer into memory, and use a random sorter to sort the SST files from L1 to L6, and then read the first data block of the first SST file after sorting at each level into memory, and then create an iterator pointer and a current pointer. The current pointer points to the iterator pointer in a loop; 3) Return the key-value pair data pointed to by the current pointer, the first thread parses it, and writes it into the buffer (buffer 1 and buffer 2); 4) The second thread (thread 2) converts the data format of the data parsed by the first thread (thread 1). At the same time, the first thread (thread 1) continues to perform the next key-value pair data reading and parsing operation.
[0050] See also Figure 2 As shown, it is a data pre-fetching flow chart of the key-value storage and data pre-fetching method based on Caffe application characteristics, which includes the following steps.
[0051] 101 , the first thread sends a read request to the key-value storage system, to 102 .
[0052] 102, determine whether a memory table and an immutable memory table exist, if so, go to 103, otherwise go to 104.
[0053] 103, create table iterators for the memory table and the immutable memory table, and go to 104 after creation.
[0054] 104, create a table iterator for all SST files in L0, and read the first data block of all files into memory, and then go to 105.
[0055] 105, read the index information of all SST files in layers L1 to L6 respectively, the random sorter randomly sorts the order of files in layers L1 to L6, and writes them into the random order table, and then goes to 106.
[0056] 106, according to the file order in the random order table, create a level iterator for L1 to L6 layers, and read the first data block in the first SST file of each layer into the memory, and then go to 107 after reading.
[0057] 107, all table iterators and level iterators are merged into a global range query iterator, and n iterator pointers and a current pointer are created. The iterator pointers point to data blocks in sequence, and the current pointer points to iterator pointer 1 (that is, the first iterator pointer), and then points to 108.
[0058] 108, the first thread reads the key-value pair pointed to by the current pointer, writes it into the buffer after parsing, and then goes to 109 after writing.
[0059] 109 , the second thread converts the format of the data in the buffer and writes it into the array. Meanwhile, the first thread executes the Next operation, and the iterator pointer moves backward to 110 .
[0060] 110, determine whether there is any unread data in the system, if so, go to 111, otherwise go to 114.
[0061] 111, point the current pointer to the next iterator pointer with unread data, and then point to 112.
[0062] 112, determine whether the number of data in the array reaches a batch size, if so, go to 113, otherwise go to 108.
[0063] 113, forward pass the array to other layers in the model and create a new array, which is then passed to 108.
[0064] 114, determine whether the next round of data reading needs to be executed, if so, go to 101, otherwise go to 115.
[0065] 115, data prefetching completed, delete global range query iterator.
[0066] See also Figure 3 As shown, it is a schematic diagram of out-of-order reading of key-value storage and data pre-fetching method based on Caffe application characteristics.
[0067] Figure 3The process of out-of-order reading of model training instances in two training rounds of the Caffe training model is shown. In the first round, after the L1 to L6 layers are processed by the random sorter, the file order is: L1 (SST13, SST11, SST12), ..., L6 (SST2, SST4, SST3, SST1). According to this order, the range of data blocks read into the memory includes: u130-u134, u140-u144, u120-u124, ..., u10-u14. Among them, u130-u134 and u140-u144 belong to the L0 layer. In the initial stage of training, the first data block of all SST files in the L0 layer needs to be loaded into the memory. Therefore, the input order of the model training instances in the first round is: 130, 140, 120, ..., 10, then 131, 141, 121, ..., 11, and so on. In the second round, the order of files from L1 to L6 layers after the random sorter becomes: L1 (SST11, SST13, SST12), ..., L6 (SST3, SST1, SST2, SST4). The corresponding read data block range includes: u130-u134, u140-u144, u100-u104, ..., u20-u24. Therefore, the input order of the second round model training instance is: 130, 140, 100, ..., 20, then 131, 141, 101, ..., 21, and so on. If the model training includes multiple rounds, the operation of each round is the same as the above method.
[0068] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.
Claims
1. A key-value storage and data pre-fetching method based on Caffe application characteristics, applied in a key-value storage system, characterized in that: include: Set the data block size in the sort string table SST file according to the bandwidth size of different storage devices; When image data is written into the key-value storage system, the image data is converted into key-value pairs and saved in the memory table. If the capacity of the memory table reaches the preset threshold, the image data is converted into an immutable memory table and refreshed to the disk, organized into an SST file. Each SST file contains multiple data blocks for storing key-value pair data. When performing data pre-fetch operations, the disk I / O unit size to be read is matched with the bandwidth of the storage device. At the same time, when the image resolution is greater than the preset resolution, a dual-thread parallel execution strategy is adopted, in which one thread is responsible for reading and parsing key-value pairs and the other thread performs data format conversion. When reading key-value pairs, a global out-of-order reading mechanism is adopted. The implementation process of the global out-of-order reading mechanism is as follows: Read the first data block of all SST files in the L0 layer into memory; Use a random sorter to randomly sort the SST files of each level from L1 to L6 in the LSM tree and record them in the random order table. Then load the first data block of the first SST file of each level from L1 to L6 in the random order table into the memory, and create corresponding iterator pointers and current pointers for all data blocks in the memory; the iterator pointer points to the first key-value pair in the data block of each level in turn, and the current pointer points to the iterator pointer cyclically; among them, each iterator pointer corresponds to number 1~n; Caffe reads data according to the data pointed to by the current pointer; when all data in an SST is read, the next SST file of the same level is read from the random order table, and the reading operation continues until all data in the key-value storage system is read; For the next round of data reading, the iterator and random order table are rebuilt to ensure that the order of input instances in each round is different.
2. The key-value storage and data pre-fetching method based on Caffe application features according to claim 1, characterized in that: One of the two threads is a first thread, and the other is a second thread; the implementation process of the dual-thread parallel execution strategy is as follows: When the first thread completes reading and parsing image i, the data will be written into the first buffer, and then the second thread will convert the data in the first buffer; at the same time, the first thread starts to read and parse image i+1 and write it into the second buffer, and the second thread starts to convert the data in the second buffer after processing the data of image i; the two buffers are used alternately, and the mutex lock is used to ensure the safe access to the buffer.
3. The key-value storage and data pre-fetching method based on Caffe application features according to claim 2 is characterized in that: The first buffer and the second buffer are both located in the memory, and are used to temporarily store the image data read and parsed by the first thread, and the data is cyclically allocated between the two buffers through the index variable j. When a thread is using a buffer, a mutex lock is used to ensure that other threads cannot modify the data during the access of the buffer.
4. The key-value storage and data pre-fetching method based on Caffe application features according to claim 1, characterized in that: The random sorter is located in the memory and is used to read the index information of all SST files in the L1 to L6 layers, and disrupt the order of files in each layer through a random number generation mechanism.
5. The key-value storage and data pre-fetching method based on Caffe application features according to claim 1, characterized in that: The random order table is located in the memory and includes two fields: level number and file order; wherein the level number is used to record the numbers of the layers L1 to L6, and the file order is used to save the order of files of each level after random sorting.
6. The key-value storage and data pre-fetching method based on Caffe application features according to claim 1, characterized in that: The preset resolution is 64×64 pixels.
7. The key-value storage and data pre-fetching method based on Caffe application features according to claim 1, characterized in that: The data pre-fetching operation specifically includes: S101, the first thread sends a read request to the key-value storage system, to S102; S102, determine whether there is a memory table and an immutable memory table, if so, go to S103, otherwise, go to S104; S103, create table iterators for the memory table and the immutable memory table, and then go to S104; S104, create table iterators for all SST files in L0, and read the first data block of all files into memory, and then go to S105; S105, respectively read the index information of all SST files in layers L1 to L6, the random sorter randomly sorts the order of files in layers L1 to L6, and writes them into the random order table, and then goes to 106; S106, create a level iterator for L1 to L6 layers according to the file order in the random order table, and read the first data block in the first SST file of each layer into the memory, and then go to S107; S107, merge all table iterators and level iterators into a global range query iterator, and create n iterator pointers and a current pointer, the iterator pointers point to the data blocks in sequence, the current pointer points to the first iterator pointer, and then points to S108; S108, the first thread reads the key-value pair pointed to by the current pointer, parses it and writes it into the buffer, and then goes to S109; S109, the second thread converts the format of the data in the buffer and writes it into the array. At the same time, the first thread executes the Next operation, and the iterator pointer moves backward, and then moves to S110; S110, determine whether there is any unread data in the system, if so, go to S111, otherwise go to S114; S111, point the current pointer to the next iterator pointer with unread data, and then point to S112; S112, determine whether the number of data in the array reaches a batch size, if so, go to S113, otherwise go to S108; S113, forward pass the array to other layers in the model, and create a new array, and then go to S108; S114, determine whether the next round of data reading needs to be performed, if so, go to S101, otherwise go to S115; S115, data pre-fetching is completed, and the global range query iterator is deleted.
Citation Information
Patent Citations
Transaction execution method and device, computing equipment and storage medium
CN115098537A
Deep learning heterogeneous computing method based on layer-wide memory allocation and system thereof
US20200272907A1
Computational data storage systems
US20220188028A1
Data processing method and apparatus, electronic device, and readable storage medium
US20240419405A1
Cited By
Data acquisition method for track defect detection
CN120852384A