Data deduplication method using GPU
By utilizing a GPU to perform data deduplication operations and dynamically selecting the processor based on usage, the method addresses the CPU overhead issue in cloud storage systems, enhancing efficiency and performance.
Patent Information
- Application Number
- PCT/KR2023/018732
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-30
AI Technical Summary
Existing data deduplication methods in cloud storage systems require significant CPU resources, leading to increased overhead and inefficiencies in storage performance, network bandwidth, and system efficiency.
A data deduplication method that leverages a GPU to perform deduplication operations, reducing CPU overhead by offloading computation-intensive tasks to the GPU, and dynamically selecting the processor (CPU or GPU) based on usage monitoring.
The method significantly reduces CPU overhead by performing computationally intensive deduplication tasks on the GPU, thereby enhancing storage system efficiency, reducing network bandwidth usage, and improving overall system performance.
Smart Images

Figure KR2023018732_30052025_PF_FP_ABST
Abstract
Description
Data deduplication using GPUs
[0001] The present invention relates to a method for removing data duplication, and more particularly, to a method for removing data duplication using a GPU, which is a hardware accelerator.
[0002]
[0003] The data processed in cloud storage services is growing exponentially. Data processed through cloud storage services is replicated across distributed storage systems for high availability, reliability, and failure management. This redundant data impacts the overall system, including storage performance, storage system efficiency, and network bandwidth. Therefore, cloud storage service providers utilize data deduplication to eliminate redundant data stored by users.
[0004] Data deduplication technology can be divided into inline deduplication and offline deduplication, based on the deduplication point. Inline deduplication is performed before write-requested data is stored in storage, while offline deduplication removes duplicated data after the write-requested data has been stored in storage.
[0005] Data deduplication requires resources separate from those used to store data, which increases CPU overhead. Therefore, a data deduplication method that can reduce CPU overhead is needed.
[0006]
[0007] The present invention provides a data deduplication method capable of reducing CPU overhead resulting from data deduplication operations.
[0008]
[0009] According to one embodiment of the present invention for achieving the above-described purpose, a data deduplication method is provided, including: when a write request occurs, performing a deduplication operation on data requested to be written in a GPU; and storing, in a CPU, the data requested to be written in storage according to a result of the deduplication operation of the GPU.
[0010] In addition, according to another embodiment of the present invention for achieving the above-mentioned purpose, a data deduplication method is provided, including: a step of monitoring CPU usage and GPU usage when a write request occurs, and generating monitoring data including the CPU usage and GPU usage; a step of selecting a processor among the CPU and the GPU to perform deduplication on a write-requested file according to the monitoring data; and a step of performing the deduplication using the selected processor.
[0011]
[0012] According to one embodiment of the present invention, data deduplication operations requiring a significant amount of computation are performed on a GPU, thereby reducing CPU overhead.
[0013]
[0014] FIG. 1 is a drawing for explaining a storage server according to one embodiment of the present invention.
[0015] FIG. 2 is a drawing for explaining a data deduplication method according to one embodiment of the present invention.
[0016] FIG. 3 is a drawing for explaining a data deduplication method according to another embodiment of the present invention.
[0017] FIG. 4 is a drawing for explaining a data deduplication method according to another embodiment of the present invention.
[0018]
[0019] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0020]
[0021] The present invention proposes a data deduplication method using a hardware accelerator to reduce CPU overhead due to data deduplication operations, and one embodiment of the present invention can use CUDA so that a GPU, which is one of the hardware accelerators, can perform data deduplication operations.
[0022] CUDA (Compute Unified Device Architecture) is a GPGPU platform supporting GPGPU technology. It provides an API that allows computations performed on the GPU to be written in C. The APIs provided by CUDA can be broadly categorized into the Runtime API, Driver API, and Math API.
[0023] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings.
[0024]
[0025] FIG. 1 is a drawing for explaining a storage server according to one embodiment of the present invention.
[0026] Referring to FIG. 1, a storage server according to one embodiment of the present invention includes a CPU (110), a first memory (120), a GPU (130), a second memory (140), and storage (150).
[0027] When an input / output request, such as a write request, occurs, the CPU (110) requests a deduplication operation from the GPU (130). Then, the CPU (110) receives the deduplication operation result from the GPU (130) and stores the write-requested data in the storage (150).
[0028] The GPU (130) performs a deduplication operation on the data requested to be written in response to a deduplication operation request from the CPU, and determines whether the data requested to be written is already stored in the storage (150).
[0029] The GPU (130) divides the data requested to be written into data blocks of a preset size, inputs each data block into a preset hash function, and calculates a hash value for each data block. Then, the GPU (130) checks whether the hash value for the data block is included in a deduplication table, thereby checking whether each data block is stored in duplicate. The deduplication table is a table in which hash values and storage address values for data stored in storage (150) are recorded, and if the hash value for the data block is already recorded in the deduplication table, the GPU (130) can determine that the corresponding data block is data already stored in the storage (150).
[0030] After performing a deduplication operation, the GPU (130) transmits the deduplication operation result, i.e., information on whether the data requested to be written is already stored in storage (150), to the CPU (110).
[0031] The CPU (110) stores data blocks that are not stored in the storage (150) among the data requested to be written in the storage (150), and does not store data blocks that are already stored in the storage (150) in the storage (150).
[0032] As described above, the data deduplication operation is performed through tens of thousands of function calls, including hash value calculation, confirmation of duplicate storage, etc., and according to one embodiment of the present invention, the data deduplication operation, which requires a significant amount of computation, is performed on the GPU, thereby reducing the overhead of the CPU.
[0033] Meanwhile, depending on the embodiment, the storage server may selectively perform data deduplication operations on the GPU (130), and the CPU (110) may select the GPU as the processor to perform deduplication based on CPU usage and GPU usage.
[0034]
[0035] FIG. 2 is a drawing for explaining a data deduplication method according to one embodiment of the present invention, and FIG. 2 explains a data deduplication method performed in the storage server described above as one embodiment.
[0036] According to an embodiment of the present invention, when a write request occurs, a storage server performs a deduplication operation on the data requested to be written on the GPU (S210). The storage server may use the GPU to divide the data requested to be written into multiple data blocks and perform a deduplication operation to determine whether each data block is data already stored in the storage. The storage server may divide the data requested to be written into data blocks of a size set by the user and perform the deduplication operation.
[0037] The storage server can then select a GPU as the processor to perform deduplication based on CPU and GPU usage, and perform deduplication operations on the GPU. When CPU usage is low and GPU usage is high, performing deduplication operations on the CPU may be more advantageous than performing them on the GPU. Therefore, the storage server can monitor CPU and GPU usage and, based on the monitoring results, select the processor to perform deduplication.
[0038] Furthermore, since deduplication operations are affected by the size of the data block and the size and type of the requested data (i.e., the file), the storage server can select a GPU as the processor to perform deduplication based on at least one of the following: the size of the data block, the size and type of the requested data, or both. For example, as the size of the data block increases, the cost of hash value calculations also increases, which may lead to an increase in the cost of deduplication operations.
[0039] In one embodiment, the storage server may select the GPU as the processor to perform deduplication when the GPU usage is below a threshold usage, and the threshold usage may be adaptively adjusted depending on the size of the data block. As described above, since the deduplication operation cost also increases as the size of the data block increases, the threshold usage may be adjusted to decrease. In other words, the threshold usage may be adjusted to be inversely proportional to the size of the data block. Accordingly, if the size of the data block is small while the GPU usage is high, the GPU may be selected as the processor to perform deduplication, and if the size of the data block is large, the CPU may be selected as the processor to perform deduplication.
[0040] Additionally, the storage server may utilize a pre-trained machine learning model to select a processor to perform deduplication based on, for example, CPU usage, GPU usage, data block size, and the size and type of data requested to be written.
[0041] According to one embodiment of the present invention, a storage server stores data requested to be written by a CPU in storage (S220) based on the results of a deduplication operation performed by a GPU. As described above, the storage server stores data blocks not yet stored in storage among the data requested to be written, and does not store data blocks already stored in storage.
[0042]
[0043] FIG. 3 is a drawing for explaining a data deduplication method according to another embodiment of the present invention, and in FIG. 3, a data deduplication method in a software layer is explained as an example.
[0044] When an application running in user space requests data input / output (I / O), it requests data deduplication on the GPU using the CUDA API.
[0045] And file systems in the kernel space, such as ZFS (ZettaByte File System), request data deduplication operations from the GPU according to the application's input / output requests. The GPU performs the deduplication operation according to the ZFS data deduplication operation request and passes the data deduplication operation result to ZFS. And ZFS stores data blocks that have not been stored in the storage among the data requested for I / O in the storage, and does not store data blocks that have already been stored in the storage.
[0046]
[0047] FIG. 4 is a drawing for explaining a data deduplication method according to another embodiment of the present invention.
[0048] Referring to FIG. 4, a storage server according to an embodiment of the present invention monitors CPU usage and GPU usage when a write request occurs, and generates monitoring data including the CPU usage and GPU usage (S410). According to an embodiment, the storage server may further generate monitoring data including at least one of the size of a data block used for deduplication, and the size and type of a file requested to be written.
[0049] The storage server then selects a processor to perform deduplication on the requested file, among the CPU and GPU, based on monitoring data (S420). As in the aforementioned embodiment, the computing device can select a processor to perform deduplication by comparing CPU or GPU usage with a threshold usage.
[0050] In one embodiment, the storage server compares CPU usage and GPU usage with a threshold usage. If the GPU usage is less than the threshold usage or both the CPU usage and GPU usage are greater than the threshold usage, the storage server may select the GPU as the processor to perform deduplication. Furthermore, if the CPU usage is less than the threshold usage and the GPU usage is greater than the threshold usage, the storage server may select the CPU as the processor to perform deduplication. The threshold usages compared to the CPU usage and GPU usage may be set differently.
[0051] Alternatively, the storage server can input monitoring data into a pre-trained machine learning model, and select a processor to perform deduplication by selecting one of the preset classes. The machine learning model may be, for example, a logistic regression, random forest, or naive Bayes model. In the reference storage server, when the processor performing data deduplication is a CPU or a GPU, training for the machine learning model can be performed using the acquired monitoring data.
[0052] A machine learning model may be trained to classify input monitoring data into one of three classes. The first class may be a class that performs deduplication on the CPU, the second class may be a class that performs deduplication on the GPU, and the third class may not perform deduplication. In some cases, for example, when the size of the requested write file is very small, it may be advantageous to not perform deduplication. Therefore, the classes used to train the machine learning model may include the third class.
[0053] The storage server performs deduplication using the processor selected in step S420 (S430). If the machine learning model classifies the monitoring data into the first class, the storage server performs the deduplication operation on the CPU. If the machine learning model classifies the monitoring data into the second class, the storage server performs the deduplication operation on the GPU. If the machine learning model classifies the monitoring data into the third class, the storage server stores the write-requested file in the storage, regardless of whether the write-requested data is stored in duplicate.
[0054] Machine learning models can be retrained using data collected during the data deduplication process on storage servers.
[0055]
[0056] The technical contents described above may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiments, or may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.
[0057]
[0058] Although the present invention has been described with reference to specific details such as specific components and limited embodiments and drawings, these have been provided only to help a more general understanding of the present invention, and the present invention is not limited to the above embodiments, and those with ordinary skill in the art to which the present invention pertains can make various modifications and variations based on these descriptions. Therefore, the spirit of the present invention should not be limited to the described embodiments, and all things that are equivalent or equivalent to the following claims as well as the claims are considered to fall within the scope of the spirit of the present invention.
Claims
1. When a write request occurs, the GPU performs a deduplication operation on the data requested to be written; and A step of storing the requested write data in storage in the CPU according to the result of the deduplication operation of the GPU. A method for removing data duplication comprising:
2. In paragraph 1, The step of performing deduplication on the above requested write data is A step of dividing the above-mentioned write-requested data into multiple data blocks; and A step of determining whether each of the above data blocks is data already stored in the storage. A method for removing data duplication comprising:
3. In paragraph 2, The step of storing the above requested write data in storage is Among the above data blocks, a data block that is not data already stored in the storage is stored in the storage. How to deduplicate data.
4. In paragraph 1, The step of performing deduplication on the above requested write data is Depending on CPU usage and GPU usage, the GPU is selected as the processor to perform the deduplication. How to deduplicate data.
5. In paragraph 4, The step of performing deduplication on the above requested write data is Selecting the GPU as a processor to perform the deduplication according to at least one of the size of the data block, the size and type of the data requested to be written. How to deduplicate data.
6. In paragraph 5, The step of performing deduplication on the above requested write data is If the above GPU usage is less than the threshold usage, the GPU is selected as the processor to perform the above deduplication. The above threshold usage is Adaptively adjusted according to the size of the above data block How to deduplicate data.
7. When a write request occurs, a step of monitoring CPU usage and GPU usage and generating monitoring data including the CPU usage and GPU usage; A step for selecting a processor to perform deduplication for a write-requested file among CPU and GPU according to the above monitoring data; and A step of performing the deduplication using the above-mentioned selected processor A method for removing data duplication comprising:
8. In paragraph 7, The steps for generating the above monitoring data are: Generating the monitoring data further including at least one of the size of the data block used for the deduplication, the size and type of the file requested to be written, How to remove data duplication.
9. In paragraph 8, The step of selecting a processor to perform the above deduplication is By inputting the above monitoring data into a pre-trained machine learning model, one of the preset classes is selected. The above class is, A first class that performs the deduplication on the above CPU; A second class that performs the deduplication on the GPU; and A third class that does not perform the above deduplication A method for removing data duplication comprising:
10. In paragraph 9, The steps for performing the above duplicate removal are If the machine learning model classifies the class of the monitoring data into the third class, the file requested to be written is stored in storage regardless of whether the file requested to be written is stored in duplicate. How to remove data duplication.
Citation Information
Patent Citations
Data parallel deduplication system
KR101229851B1
Method and System to manage and schedule GPU memory resource in Container-based virtualized environment
KR102092459B1
Method for storing data in virtual environment
KR102484914B1
Data input and output method using storage node based key-value srotre
KR102599116B1
Cited By
Leveraging host accelerator resources to reduce storage system computational load
US12737111B2
Leveraging host accelerator resources to reduce storage system computational load
US20250298507A1