Storage control device, storage control method, and storage control program

The storage control device employs a secondary hash value and update-only hash table to efficiently manage deduplication in VDI environments, addressing high computational loads and speeding up update processes.

JP7701500B1Active Publication Date: 2025-07-01NEC PLATFROMS LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024035131
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-07-01
Estimated Expiration
2044-03-07

AI Technical Summary

Technical Problem

In VDI environments, the computational load for deduplication processes is high due to the geometric increase in hash value calculation as bit length increases, leading to prolonged update times for duplicate elimination.

Method used

A storage control device and method that utilizes a secondary hash value with a shorter bit length, combined with an update-only hash table, to compare and update file information efficiently by referencing the main hash value and physical address of the latest data file, reducing the need for extensive hash entry searches and calculations.

Benefits of technology

This approach enables high-speed deduplication updates by minimizing the computational load and reducing the need for extensive hash entry searches and calculations, thereby optimizing storage capacity and processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701500000001_ABST
    Figure 0007701500000001_ABST
Patent Text Reader

Abstract

Reduce the computational load of the file management system. 【Solution】 Compare the secondary hash value obtained from the data file written to the storage device with the secondary hash value included in the file information stored in the directory. If the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, read the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device from the update-only hash table, and use the main hash value and hash entry address of the latest data file included in the secondary hash entry, and the physical address of the latest data file to execute a process of updating the file information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a storage control device, a storage control method, and a storage control program, and more particularly, to a storage control device, a storage control method, and a storage control program having a deduplication function implemented therein.

Background Art

[0002] In a VDI (Virtual Desktop Infrastructure) environment, a large number of users use a common OS (Operating System) and common applications on virtual PCs (Personal Computers).

[0003] In an OS update in a VDI environment, for each user, an operation for update is performed on a common OS image. When the OS images before update in the VDI environment of each user are common, the OS images after update are also common among users.

[0004] When a deduplication function is implemented in a storage device storing a VDI environment, in related technologies, the OS and applications of each user are processed as duplicate data to improve storage capacity efficiency and speed up read processing.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] In related technologies, each user using a VDI environment calculates the hash value of updated data every time an update operation is performed, and performs processing such as comparison and duplication elimination of the hash values, and subtraction of the number of duplicates of the old OS image.

[0007] The hash value of data requires a certain bit length. However, as the bit length of the hash value increases, the computational load of the hash value increases geometrically. As a result, the update process of duplicate elimination takes a long time.

[0008] The present disclosure has been made in view of the above problems, and an object thereof is to reduce the computational load of a file management system.

Means for Solving the Problems

[0009] A storage control device according to an aspect of the present disclosure is a storage control device that controls a storage device. The storage device is connected to a directory, a hash table, a dedicated update hash table, and a disk. The disk stores actual data files. The directory stores file information including the physical address of the actual data file, the main hash value used in the deduplication process of the actual data file, and a secondary hash value having a shorter bit length than the main hash value. The hash table stores hash entries and a hash entry search database for searching for the hash entries from the main hash value. The dedicated update hash table stores secondary hash entries corresponding to the secondary hash values. The secondary hash entry includes the main hash value (new) of the latest data file, the hash entry address, the main hash value (old) of the data file before update, and the physical address of the latest data file. Compare the secondary hash value obtained from the data file written to the storage device with the secondary hash value included in the file information stored in the directory. If the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, read the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device from the dedicated update hash table, and compare the main hash value (old) of the data file before update with the main hash value before update described in the directory from the sub-entries included in the secondary hash entry. If it is detected that the main hash value (old) of the data file before update matches the main hash value described in the directory, compare the data file written to the storage device with the latest data file. If it is confirmed that the data file written to the storage device matches the latest data file, it is determined that the data is duplicate, and the main hash value and the hash entry address of the latest data file, andExecute a process of updating the file information by using the physical address of the latest data file.

[0010] A storage control method according to one aspect of the present disclosure is a storage control method for controlling a storage device. The storage device is connected to a directory, a hash table, a dedicated update hash table, and a disk. The disk stores actual data files. The directory stores file information including the physical address of the actual data file, the main hash value used in the deduplication process of the actual data file, and a secondary hash value having a shorter bit length than the main hash value. The hash table stores hash entries and a hash entry search database for searching for the hash entries from the main hash value. The dedicated update hash table stores secondary hash entries corresponding to the secondary hash values. The secondary hash entry includes the main hash value (new) of the latest data file, the hash entry address, the main hash value (old) of the data file before update, and the physical address of the latest data file. The storage control method includes a process of comparing the secondary hash value obtained from the data file written to the storage device with the secondary hash value included in the file information stored in the directory. When the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device is read from the dedicated update hash table, and the main hash value (old) of the data file before update is compared with the main hash value before update described in the directory from the sub-entries included in the secondary hash entry. When the main hash value (old) of the data file before update and the main hash value described in the directory match, the data file written to the storage device is compared with the latest data file. When the data file written to the storage device and the latest data file match, it is determined that the data is duplicate.Cause the computer to execute a process of updating the file information using the main hash value and the hash entry address of the latest data file, and the physical address of the latest data file.

[0011] A storage control program according to an aspect of the present disclosure is a storage control program for controlling a storage device. The storage device is connected to a directory, a hash table, a dedicated update hash table, and a disk. The disk stores actual data files. The directory stores file information including the physical address of the actual data file, the main hash value used in the deduplication process of the actual data file, and a secondary hash value having a shorter bit length than the main hash value. The hash table stores hash entries and a hash entry search database for searching for the hash entries from the main hash value. The dedicated update hash table stores secondary hash entries corresponding to the secondary hash values. The secondary hash entry includes the main hash value (new) of the latest data file, the hash entry address, the main hash value (old) of the data file before update, and the physical address of the latest data file. The storage control program compares the secondary hash value obtained from the data file written to the storage device with the secondary hash value included in the file information stored in the directory. When the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, the storage control program reads the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device from the dedicated update hash table, and compares the main hash value (old) of the data file before update included in the secondary entry with the main hash value before update described in the directory. When the coincidence of the main hash value (old) of the data file before update and the main hash value described in the directory is detected, the storage control program compares the data file written to the storage device with the latest data file. When the coincidence of the data file written to the storage device and the latest data file is confirmed, it is determined that the data is duplicate.Causes the computer to execute a process of updating the file information using the main hash value and the hash entry address of the latest data file, and the physical address of the latest data file.

Advantages of the Invention

[0012] According to one aspect of the present disclosure, it becomes possible to perform the deduplication update process at high speed.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Embodiments for Carrying Out the Invention

[0014] As implementation methods of deduplication in a storage device, there are two implementation methods: a file unit method that performs deduplication for each file, and a block unit method that performs deduplication for each fixed-length or variable-length block. In the present disclosure, the deduplication method at the file unit is targeted.

[0015] 〔Embodiment 1〕 Embodiment 1 of the present disclosure will be described with reference to FIGS. 1 to 16.

[0016] (File system) FIG. 1 shows an implementation example of a file system according to an embodiment. The storage device 2 shown in FIG. 1 is different in that an update-only hash table 313 is added to the disk 31 as compared with a storage device (not shown) according to related art. The storage control device 23 is an example of a storage control device.

[0017] As shown in FIG. 1, the storage control device 23 is connected to a plurality of host computers 111 to 114. Data read or data write requests from the host computers 111 to 114 are sent to the storage control device 23 via the host interface 21 of the storage device 2. The storage control device 23 accesses the disk 31 via the disk interface 22 based on requests from the host computers 111 to 114.

[0018] In FIG. 1, the disk 31 is shown as a single physical disk, but it may be implemented as a virtual disk formed by combining a plurality of disks. In addition to using a magnetic disk as the physical disk, a non-volatile semiconductor memory such as an SSD (Solid State Drive) may be used.

[0019] In FIG. 1, in the disk 31, in addition to the directory 311 and the actual data 3151 and 3152, a hash table 312 used in the deduplication process is stored. Note that the actual data after 3152 (for example, 3153, 3154, etc.) is also recorded on the disk 31.

[0020] (Storage control device 23) FIG. 2 is a block diagram showing a configuration example of the storage control device 23 in the storage device 2.

[0021] As shown in FIG. 2, functions such as a file system 231, a RAID engine 232, deduplication 233, data compression 234, and thin provisioning 235 are implemented in the storage control device 23.

[0022] Note that the file system 231 and the deduplication 233 are essential, but other functions may or may not be implemented. Each function may be implemented in software or in hardware. Furthermore, functions other than those described above may be further implemented in the storage control device 23.

[0023] (Directory 311) FIG. 3 shows an example of the structure of the directory 311. In FIG. 3, file information 3111 and 3112 are recorded in the directory 311.

[0024] In the file information 3111, an identifier 31111 for identifying the file, a logical address 31112 indicating the recording position of the file in the storage device 2, a physical address 31113 indicating the actual data storage position, a main hash value 31114 used in the deduplication process for each file, and control information 31115 used for other file controls are recorded. Further, in the file information 3111, a storage column for a sub-hash value 31116 with a lower bit length than the main hash value 31114 is added.

[0025] In the file information 3112, an identifier 31121 for identifying the file, a logical address 31122 indicating the recording position of the file in the storage device 2, a physical address 31123 indicating the actual data storage position, a main hash value 31124 used in the deduplication process for each file, and control information 31125 used for other file controls are recorded. Further, in the file information 3112, a storage column for the sub-hash value 31126 is added.

[0026] In the directory 311, information after the file information 3112 (such as 3113, 3114, etc.) is also recorded on the disk 31. The directory 311 may have a tree structure that incorporates other directories 311 inside, and the file system 231 of the storage control device 23 also supports the tree structure of the directory 311.

[0027] (Structure of the hash table 312) Figure 4 shows the structure of the hash table 312. The hash table 312 includes a hash entry search database 3121 and hash entries 31221, 31222, 31223.

[0028] In the hash entry search database 3121, information for searching the hash entry 31221 of file A from the main hash value 312211 of file A is recorded. Generally, the hash entry search database 3121 has a tree structure.

[0029] In the hash entry 31221, the main hash value 312211 of file A, the duplication number 312212, the physical address 312213, and the control information 312214 are stored. In the hash entry 31221, a pointer to the hash value of the next search target may also be stored.

[0030] (Hash entry search database 3121) Figure 5 shows an implementation example of the hash entry search database 3121.

[0031] When searching for the hash entries 31221, 31222, 31223 using the hash entry search database 3121 shown in Figure 5, first, the search starts from the hash table root 31211. The main hash value to be searched is compared with the hash values of the hash value ranges p, q,... of the hash table root 31211. Next, the upper branch 312121 associated with the hit hash value range p is searched.

[0032] Next, search for lower branches associated with the hash value ranges pp, qq, ··· of the upper branch 312121.

[0033] In one example, when comparing the hash value ranges pp, qq, ··· of the upper branch 312121 with the main hash value to be searched, if a hit occurs in the hash value range pp, next, search for the lower branch 312131.

[0034] In one example, when a hit occurs in the hash value range ppp of the lower branch 312131, refer to the physical address 313114 and compare the hash entries 31221, 31222, 31223, etc. with the main hash value to be searched. If there is a hit hash entry, perform a duplicate elimination process based on the information of that hash entry. If there is no hit hash entry, add a new hash entry.

[0035] Note that in the hash entry search database 3121, upper branches after the upper branch 312122 (for example, 312123, 312124, etc.) and lower branches after the lower branch 312131 (for example, 312132, 312133, etc.) are also recorded.

[0036] (Update - only hash table 313) FIG. 6 shows an example of the update - only hash table 313. In FIG. 6, sub - hash entries 3131 to 3133 are implemented in the update - only hash table 313. Each sub - hash entry 3131 is stored at a relative address obtained arithmetically from the sub - hash values 31116, 31126 (FIG. 1) stored in the directory 311. A plurality of sub - entries 31311, 31312, 31313 are stored within the sub - hash entry 3131.

[0037] ​The sub-entry 31311 stores the main hash value (new) 313111 of the updated file A', the main hash value (old) 313112 of the file A before update, the hash entry address 313113 of the updated file A', the physical address 313114 of the updated file A', and the control information 313115.

[0038] Since the sub-hash value used in the update-only hash table 313 has a short bit length, there may be cases where the sub-hash values match but the data does not match.

[0039] In general deduplication processing, a case where the hash values match but the data is different is called a collision. In the update-only hash table 313, assuming such a collision may occur, multiple sub-entries are prepared for one sub-hash entry. For example, in FIG. 6, three sub-entries 31311, 31312, and 31313 are implemented in the sub-hash entry 3131.

[0040] Note that the number of sub-entries prepared in the sub-hash entry 3131 varies depending on the implementation of the storage device 2.

[0041] (Expansion of the capacity of the update-only hash table 313) Referring to FIG. 7, a method for expanding the capacity of the sub-hash entry when there are many different files having the same sub-hash value and the sub-entry 31311 in the sub-hash entry 3131 is exhausted will be described.

[0042] First, update the content of the sub-entry 31311 in the sub-hash entry 3131. Save the description content of the sub-entry 31311 in advance, invalidate all of the main hash value (new) 313111, the main hash value (old) 313112, and the hash entry address 313113 in the sub-entry 31311, and store the physical address 314114 of the additional sub-hash entry 3141 in the physical address 313114.

[0043] Finally, rewrite the control information 313115 (add an entry) and retain the link from the sub-entry 31311 to the newly added secondary hash entry 3141.

[0044] The newly added secondary hash entry 3141 secures an additional area 314 in the update-only hash table 313. Copy the information of the sub-entry 31311, which has been saved in advance, to the first sub-entry 31411 of the newly added secondary hash entry 3141.

[0045] Note that the sub-entries after the newly added secondary hash entry 3141 are also recorded on the disk 31 in the additional area 314 of the update-only hash table 313.

[0046] (Update procedure for the update-only hash table 313) With reference to FIGS. 8 to 11, the modification procedure of the directory 311 and the update-only hash table 313 in the case of performing data update processing will be described.

[0047] (Directory 311 before updating File A) In FIG. 8, it is assumed that a process of updating File A to File A' has occurred in the VDI environments of Users A to D. Before data update, the file information regarding File A for each user is stored as file information 3111, 3112, 3113, and 3114 on the directory 311 as different files for each user.

[0048] As shown in FIG. 8, before the update of the file information 3111 to 3114, the pre-update secondary hash value A is stored in the secondary hash value 31116. From the pre-update secondary hash value A, the relative address of the secondary hash entry 3131 is arithmetically determined.

[0049] In the file information 3111 before update, the physical address A of the file A before update is stored at the physical address 31113. By comparing the physical address A stored at the physical address 313114 of the sub-entry 31311 in the sub-hash entry 3131 with the physical address A of the file A before update, it is confirmed that the main hash value (new) 313111 of the sub-entry 31311 matches the data of the file A before update.

[0050] In the file system according to the present disclosure, since deduplication processing is performed, a sub-entry that matches the physical address A of the file A before update always exists in the sub-hash entry 3131. If it does not exist, it is determined that the directory 311 or the sub-hash entry 3131 is broken, and a critical failure (file system stop) process is performed.

[0051] Since the files A of each user A to D are common, the directory 311 is described so as to refer to the entity of the file A with the same physical address A.

[0052] For example, the physical address A of the file A is registered at the physical address 31113 of the file information 3111, and the main hash value A, which is the hash value of the file A, is registered at the main hash value 31114.

[0053] The sub-hash value 31116 stores the data of the sub-hash value A calculated by the algorithm for the update-only hash table 313. Regarding other file information, since the actual data is the same, the same main hash value A, the same physical address A, and the same sub-hash value A are stored.

[0054] FIG. 9 shows the update process of the directory 311 and the update-only hash table 313 when the user A updates the file A to the file A'. In FIG. 9, when the user A updates the file A to the file A', the updated file A' is written from the host computer 111 to the storage device 2.

[0055] The storage control device 23 obtains the secondary hash value A' from the written file A', reads the secondary hash value 31116 of the file information 3111, and compares it with the secondary hash value A' obtained from the file A'. If the comparison result is inconsistent, it is determined that it is a file update process, and the secondary hash entry 3132 corresponding to the secondary hash value A' obtained from the file A' is read from the update-only hash table 313.

[0056] In FIG. 9, since the sub-entry in the secondary hash entry 3132 is "empty", a new sub-entry 31321 is created. A new main hash value A' is calculated from the updated file A', and the hash entry A' and the physical address A' are obtained from the hash table 312 (FIG. 4).

[0057] Then, the obtained hash entry A' and the physical address A' are stored in the hash entry address 313213 and the physical address 313214 of the secondary hash entry 3132.

[0058] The calculated main hash value A' is registered in the main hash value (new) 313211, and the hash value A is read from the main hash value 31114 described in the file information 3111 before the update and registered in the main hash value (old) 313212. Finally, the control information 313215 of the new sub-entry 31321 is updated to indicate that the information in the sub-entry 31321 is valid.

[0059] Although FIG. 9 does not describe the update process of the secondary hash entry 3131 corresponding to the file A before the update, the update process of the secondary hash entry 3131 will be described later with reference to FIG. 11.

[0060] After the update processes of the secondary hash entry 3132 (FIG. 9) and the secondary hash entry 3131 (FIG. 11) are completed, the file information 3111 corresponding to the identifier A of the host computer 111 is updated.

[0061] In the example shown in FIG. 9, the contents of the physical address 31113, the main hash value 31114, and the sub-hash value 31116 are updated from the physical address A to the physical address A', from the hash value A to the hash value A', and from the sub-hash value A to the sub-hash value A', respectively. If necessary, the control information 31115 is also updated.

[0062] Subsequently, when user B updates file A to file A', file A' which is the update of file A is written to the storage device 2.

[0063] The storage control device 23 obtains the sub-hash value A' from the written file A', reads out the sub-hash value 31126 of the file information 3112 corresponding to the identifier B of the host computer 112, and compares it with the sub-hash value A' of the written file A'. If the comparison result is inconsistent, it is determined that it is a data update process, and the sub-hash entry 3132 corresponding to the sub-hash value A' is read out from the update-only hash table 313 (FIG. 6).

[0064] As shown in FIG. 10, a sub-entry 31321 created when user A updates file A to file A' already exists in the sub-hash entry 3132. When the hash value A of the main hash value (old) 313212 in the sub-entry 31321 is compared with the main hash value A before the update of the file information 3112 and it is confirmed that the main hash value before the update matches, it is determined that file A has been updated to file A', and the actual data after the update is compared.

[0065] The file A' transferred from the host computer 112 to the storage device 2 is compared with the actual data stored in the area pointed to by the physical address 313214 in the sub-entry 31321 of the sub-hash entry 3132 to confirm that the data matches.

[0066] When the data matches, without using the hash entry search database 3121 (Figure 4), the file information 3112 of the directory 311 and the duplicates 312212 of the hash entry 31221 are updated using the hash entry address A' stored at the hash entry address 313213 in the sub-entry 31321.

[0067] Regarding the update process of the sub-hash entry 3131 corresponding to the file A before the update, it will be described with reference to Figure 11.

[0068] After the update processes of the sub-hash entry 3132 and the sub-hash entry 3131 are completed, the file information 3112 corresponding to the identifier B of the host computer 112 is updated.

[0069] In Figure 10, the physical address 31123, the main hash value 31124, and the sub-hash value 31126 of the file information 3112 are each updated. If necessary, the control information 31125 is also updated.

[0070] (Update Process of Sub-Hash Entry 3131) With reference to Figure 11, the update process of the sub-hash entry 3131 corresponding to the file A before the update will be described. As described above, when updating the file information 3111 or the file information 3112, the sub-hash entry 3131 is specified from the sub-hash value 31116 before the update or the sub-hash value 31126 before the update.

[0071] Next, the sub-entry 31311 of the sub-hash entry 3131 is specified from the physical address 31113 or the physical address 31123 before the update. Then, from the hash entry address 313113 of the specified sub-entry 31311, the hash entry 31221 (Figure 4) to be updated is specified in the hash table 312 (Figure 4).

[0072] In the processes shown in FIGS. 9 and 10, since File A has been replaced by File A', the number of users referring to the actual data of File A has decreased.

[0073] Within the hash entry 31221 of the hash table 312 shown in FIG. 4, a duplicate count 312212 is managed. Each time the number of users referring to File A before the update decreases, the duplicate count 312212 included in the hash entry 31221 of File A in the hash table 312 shown in FIG. 4 is decremented by 1.

[0074] (Invalidation of sub-entry 31311) FIG. 11 shows the state after the update process from File A to File A' by Users A to D is completed. The physical addresses 31113 of the file information 3111 to 3114 have all been updated to the physical address A'. The update process from File A to A' means that the number of users referring to File A decreases. At this time, the duplicate count 312212 included in the hash entry 31221 of File A has become zero (none).

[0075] When the duplicate count 312212 of the hash entry 31221 becomes 0, File A is no longer referred to from anywhere, so the hash entry is deleted and the area of the actual data of File A is freed.

[0076] As a trigger for the deletion process of the hash entry 31221, since the sub-hash entry 3131 is being referred to, when it is detected that the duplicate count of the hash entry 31221 has become 0, the invalidation process of the sub-hash entry 3131 can be immediately executed.

[0077] After the duplicate count 312212 of the hash entry 31221 becomes 0 and the hash entry 31221 is deleted, the sub-entry 31311 of the sub-hash entry 3131 that refers to the same physical address A is invalidated. Then, information indicating entry invalidation is written to the control information 313115 of the sub-entry 31311.

[0078] (Data update process) Referring to FIGS. 12 to 16, the flow of the data update process will be described. In the data update process other than the first time, the flow order is FIGS. 12, combiner A, FIG. 13, combiner D, FIG. 14, combiner G, and FIG. 16.

[0079] (Data writing process) Referring to FIG. 12, the data writing process will be described.

[0080] As shown in FIG. 12, when the storage device 2 receives data writing from any of the host computers 111 to 114, a secondary hash value is calculated from the received data (S101). As the secondary hash value, it may be calculated by a dedicated algorithm, or it may be substituted with an existing check code such as CRC-16.

[0081] When using an existing check code, it is expected that the calculation result will be a specified value (usually 0) when all data including the check code is received. Therefore, as the secondary hash value, the code itself attached as the check code is used.

[0082] The directory 311 is read, and a comparison is made between the identifier or logical address of the write data and the identifier or logical address of the file information in the directory 311 (S102). Then, depending on the presence or absence of file information that hits the data, it is determined whether it is a new write or an update process of an existing file (S103).

[0083] Taking the configuration of the directory 311 in FIG. 3 described above as an example, the identifier 31111 of the file information 3111 is compared with the identifier of the written data. If they match (Yes in S103), it is determined to be an update process of the file information 3111. On the other hand, if they do not match (No in S103), a comparison is made between the identifier 31121 of the next file information 3112 and the identifier of the written data (S105). Thereafter, the file information in the directory 311 and the identifier of the written data are sequentially compared.

[0084] When it is determined that the data writing is an update process of an existing file (Yes in S103), the sub-hash value calculated at the time of data reception is compared with the sub-hash value recorded in the file information 3111, and a primary determination is made as to whether the data writing is an overwriting of the same data (S104).

[0085] Taking the configuration of the directory 311 in FIG. 3 as an example, when the identifier 31111 of the file information 3111 matches the identifier of the writing data (Yes in S103), the sub-hash value of the writing data is compared with the sub-hash value 31116 of the file information 3111 (S104). If they do not match (No in S104), the process proceeds to the update process of the update-only hash table 313 (connector A). If they match (Yes in S104), the written data is compared with the actual data (S105).

[0086] The physical address A is obtained from the physical address 31113 of the file information 3111, the actual data is read from the physical address A, and the comparison between the actual data and the written data is performed (S105). When the data comparison result matches (Yes in S106), since it is an overwriting and saving of the same data, update processing such as the update history is performed on the control information 31115 of the file information 3111, and the directory 311 is updated (S107). However, since the actual data and the written data match, the change of the actual data and the update process for duplicate elimination are not performed.

[0087] (Update process of the update-only hash table 313) Referring to FIG. 13, the flow of the search and update process of the update-only hash table 313 newly implemented in the present disclosure will be described. In FIG. 13, descriptions other than the numbered steps are omitted.

[0088] As shown in FIG. 13, in the case of updating an existing file (branching from the combiner A in FIG. 12), and in the case of creating a new file (branching from the combiner B in FIG. 12), in either case, the update-only hash table 313 is read, and a search for a sub-hash entry 3131 for the sub-hash value of the write data is performed (S201). Here, the description of the branch starting from the combiner B is omitted.

[0089] If the sub-entry in the sub-hash entry 3131 corresponding to the sub-hash value of the written file is "empty" (no valid sub-entry exists) (Yes in S202), a new sub-entry creation process is performed (S204). If the sub-entry in the sub-hash entry 3131 is not "empty" (No in S202), it is determined that there was a previous file update request from another user's VDI environment and an update process using the sub-hash entry was performed, and it is determined whether the main hash value (old) matches the main hash value in the directory before the update (S203). By comparing the main hash value (old) with the main hash value in the directory before the update, the number of actual data comparisons can be reduced when a collision occurs in the sub-hash entry 3131. The description of some of the flows below step S203 is omitted.

[0090] In the branch starting from the combiner A in FIG. 13, if the data does not match even after searching all valid sub-entries in the sub-hash entry 3131 (Yes in S205), a sub-entry addition process is performed (S206) (combiner C). Also, if a sub-entry addition is necessary and the sub-entry in the sub-hash entry 3131 is already "full" (Yes in S207), the addition process of the additional sub-hash entry 3141 described in FIG. 7 is performed (combiner F). Note that the flow after the combiner F, that is, the description of the addition process of the additional sub-hash entry 3141, is omitted.

[0091] If consistent data is detected, a process of incrementing the number of duplicates of the hash entry is performed (S208). Then, the process proceeds to the flow of the hash entry update process for the pre-update data shown in FIG. 14 (connector D).

[0092] (Hash Entry Update Process for Pre-Update Data) Referring to FIG. 14, the flow of the hash entry update process for the pre-update data will be described.

[0093] First, the pre-sub hash value and the pre-physical address are extracted from the pre-update file information (S301).

[0094] Next, referring to the update-only hash table 313, the sub-hash entry on the update-only hash table 313 is specified (S302). Then, the physical address on the sub-entry within the specified sub-hash entry is referred to and compared with the pre-physical address (S303).

[0095] If the pre-physical address matches (Yes in S303), the hash entry is specified from the hash entry address within the sub-entry, and the number of duplicates of the hash entry is decremented by 1 (S304). Thereafter, the flow proceeds to the process of writing new data shown in FIG. 16 (connector G).

[0096] When searching for the sub-entry of the sub-hash entry, if there is no sub-entry that matches the previous physical address (Yes in S305), it means that the file registered in the directory 311 does not exist, so it is determined as a critical failure. In this case, since the probability of other data cannot be guaranteed, the service of the storage device 2 should be stopped.

[0097] (New Data Writing Process) FIG. 15 shows the flow of the new data writing process continuing from connector E in FIG. 13.

[0098] As shown in FIG. 15, in the process of writing new data, first, the main hash value is calculated (S401). Next, the deduplication hash table is searched (S402), and it is determined whether there is an entry whose main hash value matches (S403).

[0099] If there is no entry whose main hash value matches (No in S403), the process of adding a hash entry is executed (S404), and the process proceeds to the end process flow of data writing (combiner H) shown in FIG. 16.

[0100] On the other hand, if there is an entry whose main hash value matches (Yes in S403), the write data and the actual data are compared (S405). If the write data and the actual data do not match (No in S406), collision processing is executed. In the present disclosure, a detailed description of the collision processing is omitted.

[0101] On the other hand, if the write data and the actual data match (Yes in S406), the hash entry is updated (+1) (S407). Then, the flow proceeds to the end process flow of data writing (combiner I) shown in FIG. 16.

[0102] (End process of data writing) Referring to FIG. 16, the flow of the end process of data writing will be described. As shown in FIG. 16, in the branch starting from combiner G, the process when the previously updated data is updated in the VDI environment of another user is performed.

[0103] If the hash entry of the data before update has not been deleted (No in S501), the update dedicated hash table 313 itself does not perform any update processing in particular, and only the update processing of the directory 311 is performed (S502).

[0104] On the other hand, if the hash entry of the data before update has been deleted (Yes in S501), the update dedicated table is read and the sub-hash entry is referred to (S503). Then, the sub-entry of the sub-hash entry of the update dedicated hash table 313 is deleted (S504).

[0105] In the branch starting from the combiner H, a new data creation process is performed. First, the update-only hash table 313 is read, and the sub-hash entry is referenced (S505). In the sub-entry of the sub-hash entry in the update-only hash table 313, the registration process of new data is performed. For the registration process of new data, the registration of the main hash value (new), the hash entry address, and the physical address is performed. The main hash value (old) is invalidated (S506). Then, the update process of the directory 311 is performed (S502).

[0106] In FIG. 16, the branch starting from the combiner I represents the first flow of the data update process. First, the update-only hash table 313 is read, and the sub-hash entry is referenced (S507). In the sub-entry of the sub-hash entry in the update-only hash table 313, the registration of the main hash value (new), the main hash value (old), the hash entry address, and the physical address is performed (S508). Then, the update process of the directory 311 is performed (S502).

[0107] When a large number of users update the same file in a VDI environment or the like, the process of searching the hash entry search database 3121 of the hash table 312 (the process with the largest file load in the duplicate elimination process) and the process of calculating the main hash value of the file (the process with the largest CPU load) can be omitted except for the first time.

[0108] In addition, since the duplicate elimination process can be implemented only by adding and subtracting multiple hash entries, and the hash entry can be referenced from the update-only hash table 313, the duplicate elimination process of the present disclosure can reduce the load of searching the hash entry search database compared with related technologies.

[0109] (Effect of this embodiment) According to the configuration of this embodiment, the secondary hash value obtained from the data file written to the storage device 2 is compared with the secondary hash value included in the file information stored in the directory 311. If the secondary hash value of the data file written to the storage device 2 does not match the secondary hash value included in the file information stored in the directory 311, the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device 2 is read from the update-only hash table 313, and the main hash value (old) of the data file before update is compared with the main hash value before update described in the directory 311 from the sub-entries included in the secondary hash entry. If it is detected that the main hash value (old) of the data file before update matches the main hash value described in the directory 311, the data file written to the storage device 2 is compared with the latest data file. If it is confirmed that the data file written to the storage device 2 matches the latest data file, it is determined that it is duplicate data, and the file information is updated using the main hash value and hash entry address of the latest data file, and the physical address of the latest data file.

[0110] Thereby, since it is possible to omit the process of searching the hash entry search database 3121 of the hash table 312 and the process of calculating the main hash value of the file except for the first time, the calculation load of the file management system 1 can be reduced.

[0111] [Embodiment 2] Embodiment 2 of the present disclosure will be described with reference to FIGS. 17 to 19. In the present embodiment 2, the same reference numerals as those in the first embodiment are given to the components common to the first embodiment, and the description thereof is omitted.

[0112] (File management systems 1A to 1C) FIG. 17 is an implementation example of the file management system 1A according to an embodiment. In the file management system 1 of the above-described Embodiment 1, the directory 311, the hash table 312, the update-only hash table 313, the additional area 314 of the update-only hash table, and the actual data 3151 and 3152 are all recorded on a single physical disk or virtual disk 32.

[0113] In this Embodiment 2, the actual data 3251 and 3252 are recorded on the physical disk 31 and the virtual disks 32, 33, ···.

[0114] In FIG. 17, the directory 311, the hash table 312, the update-only hash table 313, and the additional area 314 of the update-only hash table 313 are recorded on a single physical disk 31 or virtual disk 32, but there may also be an implementation in which these are recorded on separate physical disks or virtual disks.

[0115] In FIG. 17, the physical disk or virtual disks 32 to 33 are shown, but the number of disks is not limited to three, and disks after the fourth may also be implemented.

[0116] FIG. 18 is an implementation example of the file management system 1B according to an embodiment. FIG. 18 is an example in which the update-only hash table 313 is resident in the cache memory in the storage device 2 and used as the on-cache update-only hash table 23313. Since the update-only hash table 313 has a small capacity, it can be resident in the cache memory when the cache memory capacity of the storage device is large.

[0117] As a result, the update process can be processed on the cache, enabling further speedup.

[0118] FIG. 19 is an implementation example of the file management system 1C according to an embodiment. FIG. 19 is an example of accessing a common physical or virtual disk 31 from a plurality of storage devices 2. In FIG. 19, host computers 111 to 114 are connected to the storage device 2 and the storage device 4. The storage device 2 and the storage device 4 share the disk 31.

[0119] When the storage device 2 and the storage device 4 execute deduplication processing, they access the directory 311, the hash table 312, the update-only hash table 313, and the actual data 3151, 3152, etc., respectively.

[0120] When accessing data on a common disk from a plurality of storage devices, processing for maintaining data consistency, such as exclusive control, is required, and these are realized by well-known techniques.

[0121] In FIG. 19, two storage devices, the storage device 2 and the storage device 4, share the disk 31, but three or more storage devices may share the disk 31. Also, a plurality of disks 31 may be shared among a plurality of storage devices.

[0122] (Effect of this embodiment) According to the configuration of this embodiment, the sub-hash value obtained from the data file written to the storage devices 2 and 4 is compared with the sub-hash value included in the file information stored in the directory 311. If the sub-hash value of the data file written to the storage devices 2 and 4 does not match the sub-hash value included in the file information stored in the directory 311, the sub-hash entry corresponding to the sub-hash value of the data file written to the storage devices 2 and 4 is read from the update-dedicated hash table 313, and the main hash value (old) of the data file before update is compared with the main hash value before update described in the directory 311 from the sub-entries included in the sub-hash entry. If it is detected that the main hash value (old) of the data file before update matches the main hash value described in the directory 311, the data file written to the storage device 2 is compared with the latest data file. If it is confirmed that the data file written to the storage device 2 matches the latest data file, it is determined that it is duplicate data, and the file information is updated using the main hash value and hash entry address of the latest data file, and the physical address of the latest data file.

[0123] As a result, the process of searching the hash entry search database 3121 of the hash table 312 and the process of calculating the main hash value of the file can be omitted except for the first time, so that the calculation load of the file management systems 1A, 1B, 1C, and 1D can be reduced.

[0124] [Embodiment 3] Referring to FIG. 20, Embodiment 3 of the present disclosure will be described. In this Embodiment 3, the same reference numerals as those in Embodiments 1 and 2 are given to the constituent elements common to Embodiments 1 and 2, and the description thereof is omitted.

[0125] (Server 5) FIG. 20 shows an example in which the duplicate elimination processing program described in the first embodiment is implemented in the server 5, and the built-in server disk is provided as the storage device 2 described in the first and second embodiments. As shown in FIG. 20, the server 5 reads the storage control program 3261 from the virtual disk 32, expands it in the memory in the server 5, and executes the above-described duplicate elimination processing as the storage control (program) 53.

[0126] Thereby, in a VDI environment or the like, when a large number of users use the same or similar environments, it becomes possible to reduce the storage capacity by eliminating duplicates of common files.

[0127] (Differences from Related Technologies) In related technologies, when a user updates a common file, duplicate elimination processing has to be performed individually.

[0128] On the other hand, according to the configuration of the present embodiment, paying attention to the fact that when updating files shared by a large number of users such as periodic updates of the OS, the updated files are often shared, the update process can be simplified. As a result, the load of duplicate elimination processing can be reduced compared to related technologies.

[0129] Each component of the storage control device 23 described in the first to third embodiments indicates a block of a functional unit. Some or all of these components are realized by an information processing device as shown in FIG. 21, for example. FIG. 21 is a block diagram showing an example of the hardware configuration of the information processing device.

[0130] As shown in FIG. 21, the computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These components are connected to each other via a bus 121 so as to be capable of data communication with each other. Note that the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to or instead of the CPU 111.

[0131] The CPU 111 expands the programs (codes) in the present embodiment stored in the storage device 113 into the main memory 112 and executes various operations by executing these in a predetermined order. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory). Further, the programs in the present embodiment are provided in a state stored in a computer-readable recording medium 120. Note that the programs in the present embodiment may be distributed on the Internet connected via the communication interface 117.

[0132] Specific examples of the storage device 113 include a hard disk drive and a semiconductor storage device such as a flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to a display device 119 and controls the display on the display device 119.

[0133] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, and reads programs from the recording medium 120 and writes the processing results in the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0134] Also, specific examples of the recording medium 120 include general-purpose semiconductor memory devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as Flexible Disk, or optical recording media such as CD-ROM (Compact Disk Read Only Memory).

[0135] (Supplementary Note) Some or all of the above-described embodiments may be described as follows in the supplementary note, but are not limited thereto.

[0136] (Supplementary Note 1) A storage control device for controlling a storage device, The storage device is connected to a directory, a hash table, an update-only hash table, and a disk, Actual data files are stored in the disk, The directory stores file information including the physical address of the actual data file, the main hash value used in the deduplication process of the actual data file, and a secondary hash value having a shorter bit length than the main hash value, The hash table stores hash entries and a hash entry search database for searching for the hash entries from the main hash value, The update-only hash table stores secondary hash entries corresponding to the secondary hash values, The secondary hash entry includes the main hash value (new) of the latest data file, the hash entry address, the main hash value (old) of the data file before update, and the physical address of the latest data file, Compare the secondary hash value obtained from the data file written to the storage device with the secondary hash value included in the file information stored in the directory, If the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, Read the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device from the update-only hash table, and compare the main hash value (old) of the data file before the update with the main hash value before the update described in the directory from the sub-entries included in the secondary hash entry. If it is detected that the main hash value (old) of the data file before the update matches the main hash value described in the directory, compare the data file written to the storage device with the latest data file. If it is confirmed that the data file written to the storage device matches the latest data file, it is determined that there is duplicate data, and the file information is updated using the main hash value and the hash entry address of the latest data file, and the physical address of the latest data file. Storage control device.

[0137] (Appendix 2) In the case where the secondary hash value of the data file written to the storage device matches the secondary hash value included in the file information stored in the directory, Compare the data file written to the storage device with the data file stored in the storage device. Even if the data file written to the storage device does not match the data file stored in the storage device, Execute the process of updating the file information. The storage control device according to Appendix 1, characterized in that.

[0138] (Appendix 3) A plurality of sub-entries are implemented in the secondary hash entry. The storage control device according to appended claim 1, characterized in that...

[0139] (Appended claim 4) When there is no sub-entry in the sub-hash entry corresponding to the sub-hash value of the data file written to the storage device, A new sub-entry corresponding to the data file written to the storage device is created in the sub-hash entry. The storage control device according to appended claim 1, characterized in that...

[0140] (Appended claim 5) When the actual data file stored on the disk is updated by the data file written to the storage device, The sub-hash entry stores a new main hash value corresponding to the updated data file and an old main hash value corresponding to the actual data file before the update. The storage control device according to appended claim 1, characterized in that...

[0141] (Appended claim 6) The directory, the hash table, and the update-only hash table are stored on a physical disk, The actual data file is stored on one or more virtual disks. The storage control device according to appended claim 1, characterized in that...

[0142] (Appended claim 7) When the sub-entries in the sub-hash entry are full, an additional area for adding the sub-entries is secured in the update-only hash table. The storage control device according to appended claim 1, characterized in that...

[0143] (Appended claim 8) The hash entry includes the main hash value corresponding to the actual data file, a plurality of duplicates indicating the number of host computers referring to the actual data file, the physical address of the actual data file, and control information used for file control. The storage control device according to appended claim 1, characterized in that.

[0144] (Appended claim 9) A storage control method for controlling a storage device, The storage device is connected to a directory, a hash table, a dedicated update hash table, and a disk, The disk stores actual data files, The directory stores file information including the physical address of the actual data file, the main hash value used in the deduplication process of the actual data file, and a secondary hash value having a shorter bit length than the main hash value. The hash table stores hash entries and a hash entry search database for searching for the hash entries from the main hash value. The dedicated update hash table stores secondary hash entries corresponding to the secondary hash values. The secondary hash entry includes the main hash value (new) of the latest data file, the hash entry address, the main hash value (old) of the data file before update, and the physical address of the latest data file. The storage control method includes: A process of comparing the secondary hash value obtained from the data file written to the storage device with the secondary hash value included in the file information stored in the directory; When the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, Read the sub-hash entry corresponding to the sub-hash value of the data file written to the storage device from the dedicated update hash table, and compare the main hash value (old) of the data file before the update with the main hash value before the update described in the directory from the sub-entries included in the sub-hash entry. When the main hash value (old) of the data file before the update matches the main hash value described in the directory, compare the data file written to the storage device with the latest data file. When the data file written to the storage device matches the latest data file, it is determined that there is duplicate data, and the file information is updated using the main hash value and the hash entry address of the latest data file, and the physical address of the latest data file. A storage control method for causing a computer to execute.

[0145] (Appendix 10) A storage control program for controlling a storage device, The storage device is connected to a directory, a hash table, a dedicated update hash table, and a disk. The disk stores actual data files. The directory stores file information including the physical address of the actual data file, the main hash value used for deduplication processing of the actual data file, and a sub-hash value having a shorter bit length than the main hash value. The hash table stores hash entries and a hash entry search database for searching for the hash entries from the main hash value. The dedicated update hash table stores sub-hash entries corresponding to the sub-hash values. The secondary hash entry includes the main hash value (new) of the latest data file, the hash entry address, the main hash value (old) of the data file before update, and the physical address of the latest data file. The storage control program A process of comparing the secondary hash value obtained from the data file written to the storage device with the secondary hash value included in the file information stored in the directory. When the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory. Read the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device from the update-only hash table, and compare the main hash value (old) of the data file before update with the main hash value before update described in the directory from the sub-entries included in the secondary hash entry. When the match between the main hash value (old) of the data file before update and the main hash value described in the directory is detected, compare the data file written to the storage device with the latest data file. When the match between the data file written to the storage device and the latest data file is confirmed, it is determined that there is duplicate data, and a process of updating the file information using the main hash value of the latest data file, the hash entry address, and the physical address of the latest data file. A storage control program that causes a computer to execute.

[0146] Also, some or all of the configurations described in Appendices 2 to 8 subordinate to Appendix 1 (ex. storage control device) described above can also be subordinate to Appendix 9 (ex. storage control method) and Appendix 10 (ex. storage control program) in the same subordinate relationship as Appendices 2 to 8. Furthermore, within the scope not departing from the above-described embodiments, for various hardware, software, various recording means for recording software, or systems, some or all of the configurations described as appendices can similarly be made subordinate.

[0147] As described above, the present disclosure has been explained with reference to several embodiments. However, the present disclosure is not limited to the above embodiments. Each embodiment can be combined with other embodiments as appropriate. Also, various changes that can be understood by those skilled in the art can be made to the configurations and details of the above embodiments within the scope of the present disclosure.

Industrial Applicability

[0148] The present disclosure assumes a case of providing a deduplication function in a storage device connected to a server or cloud where a large number of users such as in VDI use a virtual environment.

Explanation of Signs

[0149] 1 File management system 1A File management system 2 Storage device 23 Storage control device 31 Disk 111 Host computer 112 Host computer 113 Host computer 114 Host computer 311 Directory 312 Hash table 313 Update-only hash table 314 Expansion area 3121 Hash entry search database 3131 Sub-hash entry 3132 secondary hash entry 3133 secondary hash entry 3141 additional secondary hash entry 31311 sub-entry 31312 sub-entry 31313 sub-entry

Claims

1. A storage control device that controls a storage device, the storage device is connected to a directory, a hash table, an update-only hash table, and a disk; The disk has actual data files stored thereon, the directory stores file information including a physical address of the real data file, a main hash value used in a deduplication process of the real data file, and a sub hash value having a bit length shorter than that of the main hash value; the hash table stores hash entries and a hash entry search database for searching the hash entries from the main hash value; the update-only hash table stores a secondary hash entry corresponding to the secondary hash value, the secondary hash entry includes a main hash value (new) and a hash entry address of the latest data file, a main hash value (old) of the data file before the update, and a physical address of the latest data file; comparing a secondary hash value calculated from the data file written in the storage device with the secondary hash value included in the file information stored in the directory; If the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, The secondary hash entry corresponding to the secondary hash value of the data file written to the storage device is read from the update-only hash table, and a comparison is made between the main hash value (old) of the data file before the update and the main hash value before the update written in the directory based on a sub-entry included in the secondary hash entry. If a match is detected between the main hash value (old) of the data file before the update and the main hash value written in the directory, the data file written to the storage device is compared with the latest data file. If a match is confirmed between the data file written to the storage device and the latest data file, it is determined that the data is duplicate data, and a process of updating the file information is performed using the main hash value and the hash entry address of the latest data file, and the physical address of the latest data file. Storage control device.

2. When the secondary hash value of the data file written to the storage device matches the secondary hash value included in the file information stored in the directory, comparing the data file written to the storage device with the data file saved in the storage device; Even if the data file written to the storage device does not match the data file saved in the storage device, Executing the process of updating the file information The storage control device according to claim 1 .

3. Within the secondary hash entry, multiple sub-entries are implemented. The storage control device according to claim 1 .

4. If there is no sub-entry in the secondary hash entry that corresponds to the secondary hash value of the data file written to the storage device, A new sub-entry corresponding to the data file written to the storage device is created in the secondary hash entry. The storage control device according to claim 1 .

5. When the actual data file stored on the disk is updated by the data file written to the storage device, The secondary hash entry stores a new primary hash value corresponding to the updated data file and an old primary hash value corresponding to the real data file before the update. The storage control device according to claim 1 .

6. the directory, the hash table, and the update-only hash table are stored on a physical disk; The actual data files are stored on one or more virtual disks. The storage control device according to claim 1 .

7. If the sub-entries in the secondary hash entry are full, an expansion area for expanding the sub-entries is secured in the update-only hash table. The storage control device according to claim 1 .

8. The hash entry includes the main hash value corresponding to the real data file, a duplication number indicating the number of host computers that reference the real data file, a physical address of the real data file, and control information used in file control. The storage control device according to claim 1 .

9. A storage control method for controlling a storage device, comprising: the storage device is connected to a directory, a hash table, an update-only hash table, and a disk; The disk has actual data files stored thereon, the directory stores file information including a physical address of the real data file, a main hash value used in a deduplication process of the real data file, and a sub hash value having a bit length shorter than that of the main hash value; the hash table stores hash entries and a hash entry search database for searching the hash entries from the main hash value; the update-only hash table stores a secondary hash entry corresponding to the secondary hash value, the secondary hash entry includes a main hash value (new) and a hash entry address of the latest data file, a main hash value (old) of the data file before the update, and a physical address of the latest data file; The storage control method includes: a process of comparing a secondary hash value calculated from the data file written in the storage device with the secondary hash value included in the file information stored in the directory; If the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, a process of reading from the update-only hash table the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device, comparing the primary hash value (old) of the data file before the update with the primary hash value before the update written in the directory from a sub-entry included in the secondary hash entry, and if a match is detected between the primary hash value (old) of the data file before the update and the primary hash value written in the directory, comparing the data file written to the storage device with the latest data file, and if a match is confirmed between the data file written to the storage device and the latest data file, determining that the data is duplicate data, and updating the file information using the primary hash value and hash entry address of the latest data file and the physical address of the latest data file; A storage control method for causing a computer to execute the above.

10. A storage control program for controlling a storage device, the storage device is connected to a directory, a hash table, an update-only hash table, and a disk; The disk has actual data files stored thereon, the directory stores file information including a physical address of the real data file, a main hash value used in a deduplication process of the real data file, and a sub hash value having a bit length shorter than that of the main hash value; the hash table stores hash entries and a hash entry search database for searching the hash entries from the main hash value; the update-only hash table stores a secondary hash entry corresponding to the secondary hash value, the secondary hash entry includes a main hash value (new) and a hash entry address of the latest data file, a main hash value (old) of the data file before the update, and a physical address of the latest data file; The storage control program includes: a process of comparing a secondary hash value calculated from the data file written in the storage device with the secondary hash value included in the file information stored in the directory; If the secondary hash value of the data file written to the storage device does not match the secondary hash value included in the file information stored in the directory, a process of reading from the update-only hash table the secondary hash entry corresponding to the secondary hash value of the data file written to the storage device, comparing the primary hash value (old) of the data file before the update with the primary hash value before the update written in the directory from a sub-entry included in the secondary hash entry, and if a match is detected between the primary hash value (old) of the data file before the update and the primary hash value written in the directory, comparing the data file written to the storage device with the latest data file, and if a match is confirmed between the data file written to the storage device and the latest data file, determining that the data is duplicate data, and updating the file information using the primary hash value and hash entry address of the latest data file and the physical address of the latest data file; A storage control program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Duplicate file detection device

    JP2012198832A

  • Memory system and control method

    JP2019057178A

  • Thin provisioning Virtual Desktop Infrastructure virtual machines in cloud environments that do not support thin clones

    JP2020531969A

  • Systems and methods for sketch computation

    JP2023510134A

  • Image Retrieval Method Based on Variable-Length Deep Hash Learning

    US20180276528A1