Storage device, computer system, and data transfer program

By determining the fingerprints of the differential part and the differential segmented data unit of the data block in the processor, deciding whether to send the differential part or its fingerprints, the problem of poor network traffic reduction between the core and the edge in the prior art is solved, and more efficient network transmission is achieved.

CN114924687BActive Publication Date: 2025-06-10HITACHI VANDALA CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110986459.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-02-12
Filing Date
2021-08-26
Publication Date
2025-06-10
Estimated Expiration
2041-08-26

AI Technical Summary

Technical Problem

The prior art is not effective in reducing network traffic between the core and the edge, especially when there are more stubbed areas or fewer update areas within the data block.

Method used

By determining the fingerprints of the differential part of the data block and the differential segmentation data unit in the processor, it is determined whether to send the differential part or its fingerprint to other storage devices to reduce network traffic.

Benefits of technology

It effectively reduces network traffic between storage devices and optimizes network transmission efficiency in file virtualization function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114924687B_ABST
    Figure CN114924687B_ABST
Patent Text Reader

Abstract

The present invention provides a storage device, a computer system, and a data transfer program. An edge storage device (100) is connected to other storage devices capable of performing deduplication and management on file data in units of predetermined data blocks via a network, and includes a CPU (111). When the CPU (111) transfers data of a newly created file or an updated file to other storage devices, it determines whether to send the differential part or the fingerprint of the differential data block, which is the data block containing the differential part, to the core storage device based on the size of the differential part and the size of the fingerprint of the differential data block, and sends the differential part or the fingerprint to the core storage device based on the determination result. According to the present invention, the communication volume between storage devices for transferring data can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for transferring data between multiple storage devices. Background Art

[0002] There is an increasing demand for platforms that are united between bases such as a hybrid cloud or an Edge-Core union. File virtualization functionality is a technology corresponding to such a demand. In file virtualization functionality, there are functions such as detecting data of a file generated / updated by an Edge memory in byte units and asynchronously migrating it to a Core memory, a stubbing function of deleting a file that has not been accessed from a client from the memory, and a recall function of obtaining object data in byte units from a core storage device when it is referred to again from the client.

[0003] The biggest bottleneck in file virtualization functionality is network transfer between bases (for example, WAN (Wide Area Network) transfer), and a technology for reducing network traffic is required.

[0004] The following technique (referred to as network deduplication) is disclosed in Patent Document 1: Fingerprint (FP) of an edge computing data block is generated, the FP is transferred to the core, and duplicate determination is performed at the core. Thus, transfer of a data block when there is a data block with the same FP is reduced at the core.

[0005] Prior Art Documents

[0006] Patent Documents

[0007] Patent Document 1: U.S. Patent No. 6,928,526 Specification. Summary of the Invention

[0008] Problems to be Solved by the Invention

[0009] For example, according to the technique disclosed in Patent Document 1, it may not always be effective in reducing the traffic in the network between the core and the edge. For example, in the technique of Patent Document 1, an FP of a data block of a transfer object that has been updated is calculated. Data of the data block is required for the calculation of the FP. In a case where most of the data block is a stubbed area (stub area), data of the stub area needs to be recalled, and as a result, network traffic may increase. In addition, in a case where the updated area in the data block is small, since communication is performed on the FP, network traffic may increase.

[0010] The present invention has been made in view of the above circumstances, and an object thereof is to provide a technique capable of reducing the traffic between storage devices that transfer data.

[0011] Technical Means for Solving the Problems

[0012] To achieve the above object, an aspect of the present invention provides a storage device having a processor and connected via a network to another storage device capable of performing deduplication and management on data of a specified data unit in units of specified divided data units. Wherein, when the processor transmits data of a data unit including a difference to the other storage device, based on the size of the difference part and the size of the fingerprint of the divided data unit including the difference part, that is, the differential divided data unit, it determines whether to send the difference part to the other storage device or send the fingerprint of the differential divided data unit to the other storage device, and based on the determination result, sends the difference part or the fingerprint to the other storage device.

[0013] Advantages of the Invention

[0014] According to the present invention, the communication volume between storage devices for transmitting data can be reduced. Brief Description of the Drawings

[0015] Figure 1 It is a diagram showing an example of the structure of a computer system according to the first embodiment.

[0016] Figure 2 It is a diagram showing an example of the structure of an edge storage device and a core storage device according to the first embodiment.

[0017] Figure 3 It is a structural diagram of file virtualization management information according to the first embodiment.

[0018] Figure 4 It is a structural diagram of a deduplication status management table according to the first embodiment.

[0019] Figure 5 It is a structural diagram of a deduplicated data block management table according to the first embodiment.

[0020] Figure 6 It is a structural diagram of a deduplication determination table according to the first embodiment.

[0021] Figure 7 It is a structural diagram of a transmitted deduplicated data block list according to the first embodiment.

[0022] Figure 8 It is a flowchart showing an example of file migration processing according to the first embodiment.

[0023] Figure 9 It is a flowchart showing an example of differential transmission processing according to the first embodiment.

[0024] Figure 10 It is a flowchart showing an example of network deduplication processing according to the first embodiment.

[0025] Figure 11It is a flowchart of an example of the differential reflection process of the first embodiment.

[0026] Figure 12 It is a flowchart of an example of the differential transmission process of the second embodiment.

[0027] Figure 13 It is a flowchart of an example of the duplicate exclusion process of the stub file of the third embodiment. Detailed implementation manners

[0028] Several embodiments will be described with reference to the accompanying drawings. In addition, the embodiments described below do not limit the technical solutions recited in the claims, and furthermore, all the elements and combinations described in the embodiments are not necessarily essential for the solution of the invention.

[0029] In the following description, information is sometimes described by the expression "AAA table", but the information can also be expressed in any data structure. That is, in order to indicate that the information is independent of the data structure, the "AAA table" can be referred to as "AAA information".

[0030] In addition, in the following description, the "program" is sometimes used as the acting subject to describe the process. However, since the program is executed by a processor (such as a CPU (Central Processing Unit)) and appropriately uses a storage unit (such as a memory) and / or an interface device, etc. to perform a prescribed process, the acting subject of the process can also be set as the processor (or a device or system having the processor). In addition, the processor may also include a hardware circuit that performs part or all of the process. The program can also be installed from a program source into a device such as a computer. The program source can be, for example, a program distribution server or a computer-readable storage medium (such as a portable storage medium). In addition, in the following description, two or more programs can be implemented as one program, and one program can also be implemented as two or more programs.

[0031] Embodiment 1

[0032] In the computer system 1 of the first embodiment, in the core site 20, a variable-length duplicate exclusion process is performed on a file (an example of a data unit). Here, variable-length duplicate exclusion means a process of dividing the data of the file into variable-length data blocks (an example of a divided data unit) by a prescribed method and performing duplicate exclusion in units of variable-length data blocks.

[0033] Figure 1 It is a diagram showing an example of the structure of the computer system of the first embodiment.

[0034] The computer system 1 has edge sites 10-1, 10-2 and a core site 20. The edge sites 10-1, 10-2 and the core site 20 are connected via a network 30. The network 30 is, for example, a WAN. In addition, the network 30 is not limited to this, and various networks can be used. Also, in Figure 1 the shown computer system 1, there are two edge sites 10-1, 10-2, but the number of edge sites is not limited to this.

[0035] The edge site 10-1 has an edge storage device 100, a client 400, and a management terminal 500 as an example of a storage device (second storage device). The edge storage device 100, the client 400, and the management terminal 500 are interconnected, for example, via a LAN (Local Area Network).

[0036] The edge storage device 100 accesses the core storage device 200, for example, using a protocol such as HTTP (Hypertext Transfer Protocol). The specific structure of the edge storage device 100 will be described later.

[0037] The client 400 is an information processing device such as a computer capable of performing various information processes. The client 400 uses protocols such as NFS (Network File System), CIFS (Common Internet File System), or HTTP (HyperText Transfer Protocol) to utilize the storage service provided by the edge storage device 100.

[0038] The management terminal 500 manages the edge storage device 100 and gives various operation instructions to the edge storage device 100 when an abnormality occurs in the edge storage device 100, etc.

[0039] The edge site 10-2 has an edge storage device 100 and a client 400 as an example of a storage device (second storage device). In addition, Figure 1 the shown hardware structures of the edge sites 10-1, 10-2 are only examples. As long as the structure has at least one edge storage device 100 and one client 400 respectively, the number of them is not limited, and other hardware can also be included.

[0040] The core site 20 has a core storage device 200 and a client 400 as an example of a storage device (first storage device). The core storage device 200 functions as a backup target for files (an example of a data unit) stored in the edge storage devices 100 of the edge sites 10-1, 10-2.

[0041] Figure 2FIG. 0 is an example of the structures of an edge storage device and a core storage device according to a first embodiment. Further, in Figure 2 an example is shown in which the edge storage device 100 and the core storage device 200 are configured as memories having a function that can operate regardless of which is selected.

[0042] The memories (the edge storage device 100 and the core storage device 200) include a controller 110 and a storage system 130.

[0043] The controller 110 includes a CPU 111 as an example of a processor, a memory 112, a cache 113, a LAN interface (I / F) 114, a WAN I / F 115, and an I / F 116. These structures are interconnected through a communication channel such as a bus.

[0044] The CPU 111 controls the operations of the controller 110 and the entire memory.

[0045] The memory 112 is, for example, a RAM (RANDOM ACCESS MEMORY), and temporarily stores programs and data for performing the operation control of the CPU 111. The memory 112 stores a network storage program P1, an I / O hook program P3, a local storage program P5, a data mover program P7, a differential reflection program P9, and a data capacity reduction program P11. Further, each program and information stored in the memory 112 may also be stored in a storage device 134 described later.

[0046] When the network storage program P1 is executed by the CPU 111, it receives various requests such as read / write from a client 400 or the like, and processes the protocols included in the requests. For example, the network storage program P1 processes protocols such as NFS (Network FileSystem), CIFS (Common Internet File System), and HTTP (HyperText Transfer Protocol).

[0047] When the I / O hook program P3 is executed by the CPU 111, it detects operations on files stored in the storage system 130 performed by the network storage program P1. Further, the I / O hook program P3 may also detect operations on directories or objects (an example of a data unit).

[0048] When the local storage program P5 is executed by the CPU 111, it provides a file system or an object memory to the network storage program P1.

[0049] The data transfer program P7 migrates, stubifies, or restores the files in the storage device 134 detected by the IO hook program P3 to the main storage device 200.

[0050] The differential reflection program P9, when executed by the CPU 111, receives the data updates of the files from the data transfer program P7 of other bases (other sites) and reflects the data updates to the storage device 134.

[0051] The data capacity reduction program P11, when executed by the CPU 111, performs duplicate elimination on the user files stored in the storage device 134 using inlining or post-processing.

[0052] In addition, in the first embodiment, when the Figure 2 shown memory is used as the edge storage device 100, the differential reflection program P9 and the data capacity reduction program P11 are not executed, that is, duplicate elimination or differential reflection in the edge storage device 100 is not performed. Therefore, when the memory is used only as the edge storage device 100, the differential reflection program P9 and the data capacity reduction program P11 may also be absent.

[0053] The cache 113 is, for example, a RAM that temporarily stores the data written from the user terminal 400 or the data read from the storage system 130. The LAN I / F 114 communicates with other devices (such as the user terminal 400) within the site. The WAN I / F 115 communicates with the devices of other sites (other edge sites 10, core site 20) via the network 30. The I / F 116 communicates with the storage system 130. In addition, in Figure 1 , the storage system 130 is connected to the I / F 116, but the storage device 134 may also be connected.

[0054] The storage system 130 provides, for example, a block-form memory function such as FC-SAN (Fibre Channel Storage Area Network) to the controller 110. The storage system 130 has a CPU 131 as an example of a processor, a memory 132, a cache 133, a storage device 134, and an I / F 136.

[0055] The CPU 131 controls the operation of the storage system 130. The memory 132 is, for example, a RAM that temporarily stores the programs and data for controlling the operation of the CPU 131. The cache 133 temporarily stores the data written from the controller 110 or the data read from the storage device 134. The I / F 136 communicates between the storage device 134 and the controller 110. The storage device 134 is, for example, a hard disk or a flash memory, etc., and stores various files including the user files used by the user of the user terminal 400.

[0056] In addition, the storage device 134 stores file virtualization management information T10, a deduplication status management table T20, a deduplicated data block management table T30, and a deduplication determination table T40. Details of these information and tables will be described later. In addition, in the first embodiment, when the Figure 2 shown memory is used as the edge storage device 100, since the edge storage device 100 does not perform deduplication, the storage device 134 may not store the deduplication status management table T20, the deduplicated data block management table T30, and the deduplication determination table T40 either.

[0057] Next, the file virtualization management information T10 will be described.

[0058] Figure 3 is a structural diagram of the file virtualization management information of the first embodiment.

[0059] The file virtualization management information T10 is created for each user file. In addition, in this example, an example in which the file virtualization management information T10 is stored separately from the user file is shown, but the file virtualization management information T10 may also be stored in the user file. The file virtualization management information T10 has user file management information T11 and partial management information T12.

[0060] The user file management information T11 includes an access path C11, a file status C12, and a file ID C13.

[0061] The access path C11 is an address (access path) on the core storage device 200 where the user file corresponding to the file virtualization management information T10 is stored. The file status C12 indicates the status of the user file corresponding to the file virtualization management information T10. Among the statuses of the file, there are: "Dirty", which indicates that there is differential data in the user file that has not been reflected in the core storage device 200; "Cached", which indicates that the user file is stored in the core file 200; and "Stub", which indicates that at least a part of the user file area is stubbed. The file ID C13 is an identifier (file ID) indicating the main data of the user file corresponding to the file virtualization management information T10. The file ID is used for operations on the main data.

[0062] The partial management information T12 stores entries corresponding to each part when updates, appends, etc. are performed in the user file. The entries of the partial management information T12 include fields of an offset C16, a length C17, and a partial status C18.

[0063] The start position of the corresponding part when the user file is updated or the like is stored in the offset C16. The data length from the start position of the part corresponding to the entry is stored in the length C17. The state (partial state) of the part corresponding to the entry is stored in the partial state C18. In the partial state, there is "Dirty", which indicates that the partial data has not been reflected in the core storage device 200, that is, an update (including addition) has been performed after the last file migration process (refer to Figure 8 ); "Cached", which indicates that the partial data is stored in the core storage device 200, that is, there is data in both the local (edge storage device 100) and the core storage device 200; and "Stub", which indicates that the partial data is stubbed, that is, it is cleared from the local (edge storage device 100) after the file migration process (refer to Figure 8 ).

[0064] Next, the duplicate state management table T20 will be described.

[0065] Figure 4 It is a structural diagram of the duplicate state management table of the first embodiment.

[0066] The duplicate state management table T20 is created for each user file. The duplicate state management table T20 includes a field for the file ID C21 and a field group (C22 to C27) for each data block within the user file.

[0067] The file ID of the main data of the user file corresponding to the duplicate state management table T20 is stored in the file ID C21.

[0068] Each field group of the data block includes fields for the in-file offset C22, the data block length C23, the data reduction process completion flag C24, the data block state C25, the duplicate data block save file ID C26, and the reference offset C27.

[0069] The start position within the user file where the data block corresponding to the field group is stored is saved in the in-file offset C22. The data length of the data block corresponding to the field group is saved in the data block length C23. The data deletion process completion flag indicating whether the data block corresponding to the field group has completed the data reduction process is saved in the data reduction process completion flag C24. The data deletion process completion flag is set to False (false) when the data within the data block is updated, and set to true (true) after the data reduction process. The status of the data block corresponding to the field group is saved in the data block status C25. As the status of the data block, there are non-duplicate indicating that the data block does not duplicate with other data blocks and duplicate indicating that the data block duplicates with other data blocks. In addition, for a data block with a non-duplicate data block status, values are not set in the duplicate data block save file ID C26 and the reference offset C27. The ID (file ID) of the file (duplicate data block save file) that stores the data of the data block that duplicates with the data block corresponding to the field group is saved in the duplicate data block save file ID C26. The offset within the duplicate data block save file is saved in the reference offset C27, and the data of the data block (duplicate data block) that duplicates with the data block corresponding to the field group is stored in the above duplicate data block save file.

[0070] Next, the duplicate data block management table T30 will be described.

[0071] Figure 5 It is the structural diagram of the duplicate data block management table of the first embodiment.

[0072] The duplicate data block management table T30 is a table for managing the reference count of the duplicate data blocks stored in the duplicate data block save file, and stores entries for each duplicate data block save file. The entries of the duplicate data block management table T30 include fields of file ID C31, offset C32, data block length C33, and reference count C34.

[0073] The ID (file ID) of the main data of the duplicate data block save file corresponding to the entry is saved in the file ID C31. The start position of each duplicate data block within the duplicate data block save file corresponding to the entry is saved in the offset C32. The data length of each duplicate data block is saved in the data block length C33. The number of references (reference count) from the user files of each duplicate data block is saved in the reference count C34.

[0074] Next, the duplicate determination table T40 will be described.

[0075] Figure 6 It is the structural diagram of the duplicate determination table of the first embodiment.

[0076] The duplicate determination table T40 is a table that stores information for duplicate determination of data blocks, and stores entries for each data block stored in the core storage device 200. The entries of the duplicate determination table T40 include fields for fingerprint C41, file ID C42, offset C43, and data block length C44.

[0077] The fingerprint C41 stores the fingerprint of the data block corresponding to the entry. The fingerprint is the value obtained by applying a hash function to the data in the data block and is used to confirm the identity (duplication) of the data block. As a method for calculating the fingerprint, for example, MD5 (message digest algorithm 5) or SHA-1 (Secure Hash Algorithm 1) can also be used. The file ID C42 stores the file ID of the file that stores the data block corresponding to the entry. The offset C43 stores the offset within the file of the data block corresponding to the entry. The data block length C44 stores the data length of the data block corresponding to the entry.

[0078] Next, the transfer duplicate data block list T50 used in the file migration process described later will be described.

[0079] Figure 7 It is a structural diagram of the transfer duplicate data block list of the first embodiment.

[0080] The transfer duplicate data block list T50 manages the save destination on the core storage device 200 of the data blocks that are duplicates on the core storage device 200 among the data blocks having differences (differential data blocks: an example of a differential division data unit) within the file of the migration target (transfer target). The transfer duplicate data block list T50 stores entries for each data block. The entries of the transfer duplicate data block list T50 include fields for in-file offset C51, data block length C52, duplicate data block possible file ID C53, and reference offset C54.

[0081] The in-file offset C51 stores the offset within the file of the migration target of the data block corresponding to the entry. The data block length C52 stores the data length of the data block. The duplicate data block save file ID C53 stores the file ID of the duplicate data block save file on the core storage device 200, and the duplicate data block save file on the core storage device 200 stores the data block that is a duplicate of the data block corresponding to the entry. The reference offset C54 stores the offset within the duplicate data block save file, and the duplicate data block save file stores the data block that is a duplicate of the data block corresponding to the entry.

[0082] Next, the file migration process for migrating a file in the computer system 1 of the first embodiment will be described.

[0083] Figure 8It is a flowchart of an example of the file migration process of the first embodiment.

[0084] In each edge storage device 100, the file migration process is performed by the CPU 111 of the controller 110 executing the data transfer program P7. The file migration process can be performed regularly or irregularly, for example, under the condition of meeting a specified condition, or can be executed when the client 400 performs an I / O operation on the edge storage device 100. In the file migration process, the data transfer program P7 acquires a file that includes a newly created or updated differential part (differential part), that is, a file with a file status C12 of Dirty (dirty data) as the file to be processed (object file). The method of acquiring a file with a file status C12 of Dirty (dirty data) can be a method of crawling the file system or a method of extracting from an operation log that records the operations of the file system.

[0085] S101: The data transfer program P7 executes the differential transfer process (refer to Figure 9 ). In the differential transfer process, a process of transferring the differential part of the object file and reflecting it in the core storage device 200 is executed.

[0086] S102: The data transfer program P7 changes the file status C12 of the object file and the partial status C18 corresponding to the differential part of the object file to Cached (cached), and ends the file migration process.

[0087] Next, the differential transfer process of step S101 will be described.

[0088] Figure 9 It is a flowchart of an example of the differential transfer process of the first embodiment.

[0089] Basically, the differential transfer process is as follows: The differential parts of the files created and updated in the edge storage device 100 are transferred to the core storage device 200 that performs variable-length duplicate elimination and are reflected in the core storage device 200.

[0090] S201: The data transfer program P7 acquires entries with a partial status C18 of Dirty (dirty data) (corresponding to the transfer target part) from the partial management information T12 of the object file as the transfer part list.

[0091] S202: The data transfer program P7 confirms whether the target file is a stub file containing a stubbed area, that is, whether the file status C12 of the target file is Stub. As a result, if the target file is a stub file (S202: Yes), the data transfer program P7 does not perform network deduplication processing (S206) on the data of the transfer target portion, and the processing proceeds to step S210. Here, if the target file is a stub file, the reason for not performing network deduplication processing, etc. is that the amount of data to be transferred in the recall processing of stubbed data, etc. cannot be easily estimated. On the other hand, if the target file is not a stub file (S202: No), the data transfer program P7 proceeds to step S203.

[0092] S203: The data transfer program P7 uses rolling hashing or the like to divide the object file into variable-length data blocks (variable-length data blocks). Here, rolling hashing refers to a process of calculating the hash value of the data of a window of a specified length for the object file while staggering the window as a method of dividing the object file into variable-length data blocks. When the calculated hash value is a specified value, there is a method of using the portion as a dividing point of the variable-length data block. For example, the Rabin-Karp algorithm can be used as a specific process for dividing the object file into variable-length data blocks. In addition, the process for dividing the object file into variable-length data blocks is not limited to the above, and any process may be used.

[0093] The following processes of steps S204 to S209 are executed for each data block (variable-length data block) divided in S203.

[0094] S204: The data transfer program P7 uses one of the divided data blocks as a data block to be processed (target data block), and obtains management information of the target data block (file offset, data block length, etc.)

[0095] S205: The data transfer program P7 determines whether the size of the part (differential part) in the target data block whose partial status is Dirty (dirty data) (dirty data size within the data block) is larger than the fingerprint size of the data block for the network duplicate elimination process (S206) described later. Here, a data block that contains a part (differential part) with a partial status of Dirty (dirty data) within the target data block is equivalent to a differential data block. As a result, when the dirty data size within the data block is larger than the fingerprint size (S205: Yes), by performing the network duplicate elimination process, it is possible to obtain the effect of reducing the transmission of duplicate data. Therefore, the data transfer program P7 advances the process to step S206. On the other hand, when the dirty data size within the data block is not larger than the fingerprint size (S205: No), the amount of data transmitted by the party sending the dirty data (differential data) within the data block is smaller than when performing the network duplicate elimination process. Therefore, the data transfer program P7 does not perform the network duplicate elimination process and advances the process to step S209. In addition, in this case, in subsequent processing, the differential data within the data block is sent to the core storage device 200.

[0096] In addition, when it is estimated that the probability (network duplicate elimination rate: prediction ratio) of the existence of data identical to the data block of the target file in the core storage device 200 is 100% or a value close thereto, the determination of whether to execute the network duplicate elimination process based on the above-mentioned dirty data size within the data block and the fingerprint size is a preferred example. The determination of whether to execute the network duplicate elimination process using the dirty data size within the data block and the fingerprint size is not limited to the above comparison.

[0097] In addition to the above-mentioned dirty data size within the data block and the fingerprint size, the network duplicate elimination rate can also be used for the determination of whether to execute the network duplicate elimination process. For example, it can also be that when the fingerprint size + dirty data size within the data block × (1 - network duplicate elimination rate) < (dirty data size within the data block), the process advances to step S206, and when it is not satisfied, the process advances to step S209.

[0098] For example, the network duplicate elimination rate can also be determined based on the attributes of the file. For example, it can also be that if the attribute of the file is an encrypted file, it is determined that the possibility of duplication is low, so the network duplicate elimination rate is set to a low ratio (e.g., 10%). When the attribute of the file indicates a file of a word processing software such as Microsoft Office (registered trademark) or a spreadsheet software, it is determined that the possibility of duplication is relatively high, so the network duplicate elimination rate is set to a high ratio (e.g., 80%), and the conditions of step S205 are changed according to the attributes of the file.

[0099] S206: The data transfer program P7 sends the fingerprint of the object data block to the core storage device 200, and the core storage device 200 executes the network duplicate elimination process for determining the duplication of the object data block (refer to Figure 10 ). In addition, in this example, step S206 is executed for each data block, but it is also possible to comprehensively execute the process of step S206 for all data blocks that satisfy the condition of step S205.

[0100] S207: The data transfer program P7 determines whether there is data in the core storage device 200 that duplicates the object data block based on the result formed by the network duplicate elimination process S206. As a result, when there is data that duplicates the object data block (S207: Yes), the data transfer program P7 causes the process to proceed to step S208. On the other hand, when there is no data that duplicates the object data block (S207: No), the data transfer program P7 causes the process to proceed to step S209.

[0101] S208: The data transfer program P7 deletes the area within the object data block from the transfer part list, and adds an entry including the in-file offset of the file containing the object data block, the data block length, the file ID of the data block on the core storage device 200 that duplicates the data block, and the reference offset to the transfer duplicate data block list T50.

[0102] S209: The data transfer program P7 determines whether the object data block is the end of the data block (data block end). When the object data block is the end of the data block (S209: Yes), the data transfer program P7 causes the process to proceed to step S210. On the other hand, when the object data block is not the end of the data block (S209: No), the data transfer program P7 causes the process to proceed to step S204 to process the next data block.

[0103] S210: The data transfer program P7 obtains the data recorded in the transfer part list from the main file of the user file using the file ID.

[0104] S211: The data transfer program P7 obtains the access path of the object file to the core storage device 200 (the value of the access path C11 of the object file), and requests an update of this access path to the core storage device 200. At this time, the data transfer program P7 transfers the data obtained in step S210, the transfer part list, and the transfer duplicate data block list T50.

[0105] S212: The differential reflection program P9 of the core storage device 200 (the CPU 111 that strictly executes the differential reflection program P9) receives the update request from the edge storage device 100, and performs a differential reflection process on the access path specified by the update request (refer to Figure 11 ).

[0106] S213: The difference reflects the response of program P9 of the edge storage device 100 to the return of the update completion. In addition, the data transfer program P7 of the edge storage device 100 ends the differential transfer process by receiving this response.

[0107] Next, the network duplicate exclusion process in step S206 will be described.

[0108] Figure 10 It is a flowchart of an example of the network duplicate exclusion process of the first embodiment.

[0109] The network duplicate exclusion process is as follows: It is confirmed whether the same data block as the edge storage device 100 exists in the core storage device 200, and if it exists, the storage destination of the duplicate data block is obtained.

[0110] S301: The data transfer program P7 calculates the fingerprint of the target data block. For example, with respect to the target data block, the fingerprint can also be set as the hash value calculated by SHA-1.

[0111] S302: The data transfer program P7 sends a request (duplicate retrieval request) to the core storage device 200 to confirm whether the data corresponding to the target data block exists. The fingerprint of the target data block calculated in step S301 is included in this duplicate retrieval request.

[0112] S303: The data capacity reduction program P11 of the core storage device 200 receives the duplicate retrieval request from the edge storage device 100, refers to the duplicate determination table T40, and confirms whether there is a data block (duplicate data block) that is the same as the fingerprint (target fingerprint) contained in the duplicate retrieval request. As a result, if there is a data block that is the same as the fingerprint (S303: Yes), the data capacity reduction program P11 proceeds to step S304. On the other hand, if there is no data block that is the same as the fingerprint (S303: No), the data capacity reduction program P11 proceeds to step S305.

[0113] In addition, it is also possible to recalculate the fingerprint of the data block of the file when the file storing the data block that is the same as the fingerprint is not a duplicate data block saving file and a user file and the data block status C25 in the duplicate status management table T20 is non-duplicate, and confirm that there is no change in the fingerprint, that is, the data block is not updated. Moreover, it is also possible to perform duplicate exclusion of the target data block, save the data block in the duplicate data block saving file, and update the file ID, etc. of the entry corresponding to the data block in the duplicate determination table T40 to the file ID, etc. of the duplicate data block saving file.

[0114] S304: The data capacity reduction program P11 of the core storage device 200 returns the values of the file ID C42, offset C43, and data block length C44 of the entry in the duplicate determination table T40 that matches the object fingerprint to the edge storage device 100, and ends the network duplicate elimination process.

[0115] S305: The data capacity reduction program P11 of the core storage device 200 returns the case where there is no data identical to the object data block to the edge storage device 100, and ends the network duplicate elimination process.

[0116] Next, the differential reflection process of step S212 will be described.

[0117] Figure 11 It is a flowchart of an example of the differential reflection process of the first embodiment.

[0118] The differential reflection process is as follows: The differential reflection program P9 of the core storage device 200 receives the update request from the edge storage device 100, and reflects the update of the access path specified by the update request using the specified transfer part list and reflects the duplicate elimination result using the transfer duplicate data block list T50.

[0119] S401: The differential reflection program P9 obtains an entry (object entry) of the received transfer part list.

[0120] S402: The differential reflection program P9 reflects the update of the area (update area, differential area) shown by the object entry on the access path specified by the update request.

[0121] S403: The differential reflection program P9 refers to the duplicate status management information T20 to confirm whether there is an entry with a duplicate data block status C25 in the update area of the file corresponding to the access path specified by the update request. As a result, if there is an entry with a duplicate data block status (S403: Yes), the differential reflection program P9 proceeds to step S404. On the other hand, if there is no entry with a duplicate data block status (S403: No), the differential reflection program P9 proceeds to step S405.

[0122] S404: The differential reflection program P9 subtracts 1 from the reference count C34 in the entry of the duplicate data block management table T30 corresponding to the file ID of the duplicate data block save file ID C26 and the reference offset C27 of the reference offset in the entry with the duplicate data block status.

[0123] S405: The differential reflection program P9 changes the value of the data reduction processing flag C24 in the entry corresponding to the update area of the duplicate status management table T20 to False (false), and changes the value of the data block status C25 to non-duplicate.

[0124] S406: The differential reflection program P9 confirms whether the object entry is the end of the transfer part list. As a result, if the object entry is the end (S406: Yes), the differential reflection program P9 proceeds to step S407. On the other hand, if the object entry is not the end (S406: No), the differential reflection program P9 proceeds to step S401 and uses the next entry as the processing object for subsequent processing.

[0125] S407: The differential reflection program P9 obtains an entry of the received transfer duplicate data block list T50 as the processing object entry (object entry).

[0126] S408: The differential reflection program P9 increments the reference count C34 in the entry of the duplicate data block management table T30 corresponding to the file ID of the duplicate data block save file ID C53 and the reference offset C54 of the object entry by 1.

[0127] S409: The differential reflection program P9 refers to the duplicate status management information T20 of the file corresponding to the access path specified by the update request, and confirms whether the area corresponding to the offset of the file internal offset C51 and the data block length C52 of the object entry of the transfer duplicate data block list T50 contains a data block with a data block status C25 of duplicate (duplicate data block). As a result, if a duplicate data block is included (S409: Yes), the differential reflection program P9 proceeds to step S410. On the other hand, if no duplicate data block is included (S409: No), the differential reflection program P9 proceeds to step S411.

[0128] S410: The differential reflection program P9 decrements the reference count C34 in the entry of the duplicate data block management table T30 corresponding to the file ID of the duplicate data block save file ID C26 and the offset of the reference offset C27 in the entry of the duplicate data block of the duplicate status management information T20 by 1.

[0129] S411: The differential reflection program P9 sets, in the entry of the duplicate data block of the duplicate state management information T20, the offset of the in-file offset C51 of the object entry of the transmission duplicate data block list T50 and the data block length C52 of the data block length C23 with respect to the in-file offset C22, sets true (true) for the data reduction processing flag C24, sets duplicate for the data block state C25, and sets the duplicate data block save file ID C53 of the object entry of the transmission duplicate data block list T50 and the offset of the reference offset C54 for the file ID and the reference offset C27.

[0130] S412: The differential reflection program P9 confirms whether the object entry is the end of the transmission duplicate data block list T50. As a result, if the object entry is the end (S412: yes), the differential reflection program P9 ends the differential reflection process. On the other hand, if the object entry is not the end (S412: no), the differential reflection program P9 advances the process to step S407 and performs subsequent processing on the next entry.

[0131] According to the above file migration process, since it is determined whether to perform network deduplication processing (S206) or to send dirty data without performing network deduplication processing (S211) based on the fingerprint size and the Dirty (dirty data) size within the data block, and the determination result is executed, the data transfer volume between the edge storage device 100 and the core storage device 200 can be suppressed.

[0132] Embodiment 2

[0133] Next, the computer system of the second embodiment will be described. In addition, the same parts as those of the computer system of the first embodiment may be denoted by the same reference numerals and redundant descriptions may be omitted.

[0134] The configuration of the computer system of the second embodiment is the same as that of Figure 1 the computer system shown.

[0135] The configurations of the edge storage device 100 and the core storage device 200 of the second embodiment are the same as those of Figure 2 the edge storage device 100 and the core storage device 200 of the first embodiment shown. In addition, in the second embodiment, the data reduction program P11 performs deduplication (fixed-length deduplication) on a fixed-length data block (fixed-length data block: an example of a data division unit) as an object, which is different.

[0136] The file virtualization management information T10 of the second embodiment is the same as that of Figure 3 the file virtualization management information T10 of the first embodiment shown.

[0137] The duplicate status management table T20 of the second embodiment is the same as Figure 4 the duplicate status management table T20 of the first embodiment shown. In addition, in the second embodiment, since the data reduction program P11 performs fixed-length duplicate elimination, the field of the data block length C23 may also be absent from the duplicate status management table T20.

[0138] The duplicate data block management table T30 of the second embodiment is the same as Figure 5 the duplicate data block management table T30 of the first embodiment shown. In addition, in the second embodiment, since the data reduction program P11 performs fixed-length duplicate elimination, the field of the data block length C33 may also be absent from the duplicate data block management table T30.

[0139] The duplicate determination table T40 of the second embodiment is the same as Figure 6 the duplicate determination table T40 of the first embodiment shown. In addition, in the second embodiment, since the data reduction program P11 performs fixed-length duplicate elimination, the field of the data block length C44 may also be absent from the duplicate determination table T40.

[0140] The transfer duplicate data block list T50 of the second embodiment is the same as Figure 7 the transfer duplicate data block list T50 of the first embodiment shown. In addition, in the second embodiment, since the data reduction program P11 performs fixed-length duplicate elimination, the field of the data block length C52 may also be absent from the transfer duplicate data block list T50.

[0141] The file migration process of the second embodiment is the same as Figure 8 the file migration process of the first embodiment shown. However, in the file migration process of the second embodiment, Figure 12 the differential transfer process shown is performed as the differential transfer process in step S101.

[0142] Next, the differential transfer process will be described.

[0143] Figure 12 is a flowchart of an example of the differential transfer process of the second embodiment. In addition, the same symbols are assigned to the parts that are the same as Figure 8 the differential transfer process shown, and the repeated description is omitted.

[0144] The differential transfer process is basically the following process: The differential parts of the files created and updated in the edge storage device 100 are transferred to the core storage device 200 that performs fixed-length duplicate elimination and reflected in the core storage device 200.

[0145] Perform the following processing of steps S501 to S504 and S205 to S209 on each fixed-length data block obtained by dividing a file into fixed-length units. In addition, it is also possible to perform the processing after step S501 only on the fixed-length data blocks included in the transfer parts of the transfer part list.

[0146] S501: The data transfer program P7 uses one data block among the respective fixed-length data blocks as the data block to be processed (object data block), and acquires the management information (such as file offset within the file, data block length) of the object data block.

[0147] S502: The data transfer program P7 confirms whether a stubbed area (stub area) is included in the object data block. Specifically, referring to the partial management information T12, it is confirmed whether there is a part where the partial status C18 is Stub in the object data block. As a result, when a stub area exists in the object data block (S502: Yes), the data transfer program P7 proceeds to step S503. On the other hand, when no stub area exists in the object data block (S502: No), the data transfer program P7 proceeds to step S205.

[0148] S503: The data transfer program P7 confirms whether the size of the part with the partial status of Dirty (dirty data) in the object data block (dirty data size within the data block) is larger than the size obtained by adding the size of the Stub (stub) area within the data block (stub size within the data block) to the fingerprint size used for the network deduplication process S206 (that is, the total data size to be transferred when recalling the data in the stub area and transmitting the fingerprint size). As a result, when the dirty data size within the data block is larger (S503: Yes), since the party that recalls the data in the stub area and sends the fingerprint may be able to suppress the data volume, the data transfer program P7 proceeds to step S504. On the other hand, when the dirty data size within the data block is not large (S503: No), it means that transmitting the data of the differential part can more effectively suppress the data volume. Therefore, the data transfer program P7 proceeds to step S209.

[0149] In addition, when the probability (network deduplication rate: prediction ratio) estimated to have the same data as the data block of the object file in the core storage device 200 is 100% or a value close thereto, it is a preferred example to determine whether to execute the network deduplication process based on the size of Dirty (dirty data) within the data block, the size of the fingerprint, and the size of Stub (stub) within the data block. The determination of whether to execute the network deduplication process using the size of Dirty (dirty data) within the data block, the size of the fingerprint, and the size of Stub (stub) within the data block is not limited to the above comparison.

[0150] In addition, in addition to the size of Dirty (dirty data) within the data block, the size of the fingerprint, and the size of Stub (stub) within the data block, the network deduplication rate can also be used to determine whether to execute the network deduplication process. For example, when the fingerprint size + the size of Stub (stub) within the data block + the size of Dirty (dirty data) within the data block × (1 - network deduplication rate) < the size of Dirty (dirty data) within the data block, the process proceeds to step S504, and when not satisfied, the process proceeds to step S209.

[0151] S504: The data transfer program P7 recalls (reads) the data in the stub area within the data block from the core storage device 200, and the process proceeds to step S206. After step S206, in the second embodiment, since the data block length is fixed, the process related to the data block length can also be skipped.

[0152] According to the above second embodiment, even when the data block contains a stub area, the data transfer volume between the edge storage device 100 and the core storage device 200 can be suppressed.

[0153] Embodiment 3

[0154] Next, the computer system of the third embodiment will be described. In addition, the same parts as those of the computer system of the first embodiment may be denoted by the same reference numerals and redundant descriptions may be omitted.

[0155] The structure of the computer system of the third embodiment is the same as that of Figure 1 the computer system shown.

[0156] The structures of the edge storage device 100 and the core storage device 200 of the third embodiment are the same as those of Figure 2The edge storage device 100 and the core storage device 200 of the first embodiment shown are the same. In addition, in the third embodiment, the point where the data reduction program P11 operates in the edge storage device 100 to perform variable-length duplicate elimination is different. In addition, in the core storage device 200, the data reduction program P11 may or may not operate, and performs the same duplicate elimination as in the first embodiment.

[0157] The file virtualization management information T10 of the third embodiment is the same as Figure 3 the file virtualization management information T10 of the first embodiment shown, but is stored in the edge storage device 100.

[0158] The duplicate status management table T20 of the third embodiment is the same as Figure 4 the duplicate status management table T20 of the first embodiment shown, but is stored in the edge storage device 100.

[0159] The duplicate determination table T40 of the third embodiment is the same as Figure 6 the duplicate determination table T40 of the first embodiment shown, but is stored in the edge storage device 100.

[0160] In the third embodiment, there is no need to transfer the duplicate data block management table T50.

[0161] Next, the duplicate elimination process for stub files in the third embodiment will be described.

[0162] Figure 13 is a flowchart of an example of the duplicate elimination process for stub files in the third embodiment. In addition, the same symbols are used for the parts that are the same as the processes of Figure 8 , 9 and the like, and the repeated description is omitted.

[0163] In each edge storage device 100, the CPU 111 of the controller 110 executes the data capacity reduction program P11 to perform the duplicate elimination process for stub files. The duplicate elimination process for stub files can be performed regularly or irregularly, for example, when the user terminal 400 performs an I / O operation (such as a data write operation) on the edge storage device 100 when the specified conditions are met.

[0164] S603: The data transfer program P7 of the edge storage device 100 obtains the access path to the core storage device 200 (the value of the access path C11 of the object file) from the user file management information T11 corresponding to the object file, and requests the core storage device 200 to update the access path and calculate the fingerprint of the updated area. At this time, the data transfer program P7 transfers the data obtained in S210, the transfer part list, and the information related to the split point of the data block of the object file (split point information).

[0165] S604: The differential reflection program P9 of the core storage device 200 updates the areas (update area, differential area) shown by each entry of the transfer part list reflected on the access path specified by the update request.

[0166] S605: The differential reflection program P9 divides the updated file into variable-length data blocks through processing such as rolling hash, and calculates the division points of the data blocks.

[0167] S606: The differential reflection program P9 compares the calculated division points with the received division point information, calculates the fingerprints of the data blocks whose division points have changed and the data blocks whose data has been updated, and creates a fingerprint list including the offsets of the calculated data blocks, the data block lengths, and the entries storing the fingerprints.

[0168] S607: The differential reflection program P9 returns the information of the calculated division points (changed division point information) and the fingerprint list to the edge storage device 100.

[0169] S609: The data transfer program P7 returns the changed division point information and the fingerprint list to the data capacity reduction program P11. The data capacity reduction program P11 updates the duplicate status management information T20 based on the returned changed division point information. Here, the data capacity reduction program P11 changes the value of the data reduction process completion flag C24 to False (false) in the entry corresponding to the data block whose division point has changed, and changes the value of the data block status C25 to non-duplicate. In addition, when the data block status C25 of this entry is duplicate, the data capacity reduction program P11 subtracts 1 from the reference count C34 of the entry in the duplicate data block management table T30 corresponding to the file ID of the file storing the duplicate data block ID C26 and the offset of the reference offset C27.

[0170] S610: The data capacity reduction program P11 obtains an entry (object entry) to be processed from the fingerprint list.

[0171] S611: The data capacity reduction program P11 confirms whether there is an entry in the duplicate determination table T40 that is consistent with the fingerprint of the object entry. As a result, if there is a consistent entry, the data capacity reduction program P11 proceeds to step S612. If not, the data capacity reduction program P11 proceeds to step S613.

[0172] S612: The data capacity reduction program P11 performs duplicate exclusion on the data block corresponding to the object entry.

[0173] S613: The data capacity reduction program P11 adds the data block corresponding to the object entry to the duplicate determination table T40.

[0174] S614: The data capacity reduction program P11 determines whether the object entry is the end of the fingerprint list. As a result, when the object entry is the end, the data capacity reduction program P11 ends the duplicate exclusion process of the stub file. On the other hand, when the object entry is not the end of the fingerprint list, the data capacity reduction program P11 proceeds to step S610 and uses the next entry as the processing object for subsequent processing.

[0175] In addition, an example of variable-length duplicate exclusion performed by the edge storage device 100 is shown, but the present invention is not limited thereto. Fixed-length duplicate exclusion may also be performed by the edge storage device 100. In this case, the split point calculation is not performed in step S605, fingerprints are only calculated for the updated data blocks, and the processing of step S609 is not performed.

[0176] In the third embodiment, the file migration process may perform the same processing as Figure 8 the first embodiment shown. Additionally, in the third embodiment, the differential transfer process in step 101 may not be performed Figure 9 steps S202 to S208 in the differential transfer process in the first embodiment shown. Also, in the third embodiment, the differential reflection process in step S212 may not be performed Figure 11 steps S403 to 405, S407 to S412 in the differential reflection process in the first embodiment shown.

[0177] In the above-described third embodiment, when the file is a stub file having a stub area, since it is not necessary to send the data in the stub area of the file from the core storage device 200 to the edge storage device 100, the data transfer volume between the edge storage device 100 and the core storage device 200 can be suppressed.

[0178] Furthermore, the present invention is not limited to the above embodiments and can be appropriately modified and implemented without departing from the gist of the present invention.

[0179] For example, in the differential transmission process in the above embodiment, in step S205, based on the fingerprint size and the Dirty (dirty data) size within the data block, it is determined whether to perform network deduplication processing or send the differential part. However, the present invention is not limited thereto. It is also possible to determine whether to send the differential part to the core storage device 200 or send the fingerprint of the differential data block to the core storage device 200 based on the first total amount and the second total amount. The above first total amount is the total amount obtained by summing the billing amount in communication based on the size of the differential part of the data block and the billing amount for the network deduplication process of determining the duplication of the differential data block in the core storage device 200 with the data block in the core storage device 200 using the differential part of the data block containing the differential (differential data block). The above second total amount is the amount obtained by summing the billing amount in communication based on the size of the fingerprint of the data block and the billing amount for the process of determining the duplication of the differential data block in the core storage device 200 with the data block in the core storage device 200 using the fingerprint of the differential data block.

[0180] Specifically, it can also be determined that: when the formula of the billing amount per network transmission volume * (fingerprint size + Dirty (dirty data) size within the data block × (1 - network deduplication rate)) + the billing amount per calculation volume * the calculation volume of the network deduplication process of the core storage device 200 < the billing amount per network transmission volume * the Dirty (dirty data) size within the data block + the billing amount per calculation volume * (the calculation volume of the differential reflection process and the network deduplication process in the core storage device 200) is satisfied, perform the network deduplication process of sending the fingerprint of the differential data block to the core storage device 200. When the above formula is not satisfied, send the differential part to the core storage device 200. Accordingly, not only can the reduction of the data volume during transmission be focused on, but also the billing amount for transmission and the billing amount for processing can be increased to perform appropriate data transmission.

[0181] In addition, in the above embodiment, a part or all of the processing performed by the CPU can also be performed by a hardware circuit. In addition, the program in the above embodiment can be installed from a program source. The program source can also be a program distribution server or a storage medium (such as a portable storage medium).

[0182] Symbol Explanation

[0183] 1… computer system, 10, 10-1, 10-2… edge sites, 20… core site, 30… network, 100… edge storage device, 110… controller, 111… CPU, 112… memory, 113… cache, 114… LAN I / F, 115… WAN I / F, 116… I / F, 130… storage system, 131… CPU, 132… memory, 133… cache, 134… storage device, 136… I / F, 200… core storage device, 400… client.

Claims

1. A storage device having a processor and connected via a network to another storage device capable of performing deduplication and management on data of a specified data unit in units of specified divided data units. It is characterized in that: The processor, When transmitting data of a data unit containing a difference to the other storage device, based on the size of the difference part and the size of the fingerprint of the divided data unit containing the difference part, i.e., the differential divided data unit, determines whether to send the difference part to the other storage device or send the fingerprint of the differential divided data unit to the other storage device. Based on the determination result, sends the difference part or the fingerprint to the other storage device. The processor changes the condition according to the attribute of the data unit, and this condition is a condition for determining whether to send the difference part to the other storage device or send the fingerprint of the differential divided data unit to the other storage device, and is based on the size of the difference part and the size of the fingerprint of the differential divided data unit.

2. The storage device according to claim 1, It is characterized in that, The divided data unit is a data block of fixed length, The processor, When transmitting data of a data unit containing a difference to the other storage device, and when the differential divided data unit containing the difference part in the data unit containing the difference, i.e., the differential data block, contains a stubbed stub area, based on the size of the difference part of the differential data block, the size of the stub area, and the size of the fingerprint of the differential data block, determines whether to send the difference part to the other storage device or obtain the data of the stub area from the other storage device. When it is determined to obtain the data of the stub area from the other storage device, sends a transmission request for the data of the stub area to the other storage device, receives the data of the stub area from the other storage device, calculates the fingerprint of the differential data block using the received data of the stub area, and sends the fingerprint of the differential data block to the other storage device.

3. The storage device according to claim 1, It is characterized in that, The divided data unit is a data block of variable length, The processor, when transmitting data of a data unit containing a difference to the other storage device, and when the differential divided data unit containing the difference part in the data unit containing the difference, i.e., the differential data block, contains a stubbed stub area, sends the difference part to the other storage device.

4. The storage device according to claim 1, It is characterized in that, When transmitting data of a data unit including a difference to the other storage device, the processor determines whether to send the difference part of the data unit or the fingerprint of the difference-divided data unit to the other storage device based on the size of the difference part of the data unit divided based on the difference, the size of the fingerprint of the difference-divided data unit, and a prediction ratio indicating the possibility that the divided data unit of the data unit duplicates the divided data unit of the other storage device.

5. The storage device according to claim 4, wherein, the divided data unit is a variable-length data block, the difference-divided data unit is a difference data block, when transmitting data of a data unit including a difference to the other storage device, the processor determines to send the fingerprint of the difference data block to the other storage device when (the size of the fingerprint of the difference data block)+(the size of the difference part of the data block of the data unit)×(1 - prediction ratio) < (the size of the difference part of the data block of the data unit), and determines to send the difference part to the other storage device when the condition is not satisfied.

6. The storage device according to claim 4, wherein, the divided data unit is a fixed-length data block, the difference-divided data unit is a difference data block, when transmitting data of a data unit including a difference to the other storage device and when the difference data block including the difference part in the data unit including the difference includes a stubbed stub area, the processor determines to obtain data of the stub area from the other storage device when (the size of the fingerprint of the difference data block)+(the size of the stub area in the difference data block)+(the size of the difference part of the data block of the data unit)×(1 - prediction ratio) < (the size of the difference part of the data block of the data unit), and determines to send the difference part to the other storage device when the condition is not satisfied.

7. The storage device according to claim 1, wherein, when transmitting data of a data unit including a difference to the other storage device, the processor determines whether to send the difference part or the fingerprint of the difference-divided data unit to the other storage device based on a first total amount and a second total amount, the first total amount is an amount obtained by summing up a charging amount in communication based on the size of the difference part of the divided data unit of the data unit and a charging amount for processing to determine duplication of the difference-divided data unit in the other storage device with the divided data unit of the other storage device using the difference part of the divided data unit, The second total amount is the sum of the charging amount in communication based on the size of the fingerprint of the differential segmentation data unit and the charging amount for processing to determine the duplication of the differential segmentation data unit in the other storage device with the segmentation data unit of the other storage device using the fingerprint of the differential segmentation data unit.

Citation Information

Patent Citations

  • Efficient data storage system

    US6928526B1

  • Cache management

    CN104221016A

  • Data Deduplication Cache Comprising Solid State Drive Storage and the Like

    US20170206022A1