Deduplication storage device, file hierarchical movement method, and program

The deduplication storage device optimizes file movement to the cloud tier by using pool-level attribute information, addressing inefficiencies in backup systems by minimizing unnecessary data writing and deletion, thereby improving resource utilization.

JP7841277B2Active Publication Date: 2026-04-07NEC CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Backup systems that move files to a cloud tier based on individual file attributes in deduplication storage devices lead to inefficient data movement and resource consumption due to repeated writing and deletion of unique data, diminishing the effectiveness of freeing up capacity on the deduplication storage device.

Method used

A deduplication storage device and method that utilizes attribute information from multiple file storage pools to create a list of move candidates, considering the unique data rate and pool weights to optimize file movement to the cloud tier, reducing unnecessary data writing and deletion.

Benefits of technology

Reduces operational inefficiencies by minimizing extra unique data writing and movement to the deduplication storage device and cloud tier, enhancing write performance and resource utilization in backup systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007841277000001
    Figure 0007841277000001
  • Figure 0007841277000002
    Figure 0007841277000002
  • Figure 0007841277000003
    Figure 0007841277000003
Patent Text Reader

Abstract

To enable reduction in inefficiency on operations such as writing / deletion processing of excess unique data to / from a duplication exclusion storage device and moving excess data to a cloud layer, in a backup system for performing operations for repeatedly backing up similar data.SOLUTION: An on-premises duplication exclusion storage device of a backup system including a cloud layer includes: at least two file storage pools for storing files: and movement target file selection means for preparing a movement candidate list for selecting a file to be moved for all files written to at least two file storage pools using attribute information related to unique data on the files written to the file storage pools when moving the file from the file storage pool to a cloud layer of the backup system. The unique data is data that does not duplicate the other data on the duplication exclusion storage device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a deduplication storage device, a method for hierarchical movement of files, and a program.

Background Art

[0002] Many organizations have a dedicated backup system for backing up business data so that the business can continue even if data loss occurs due to equipment failure, incorrect operation, disaster, etc.

[0003] Patent Document 1 relates to a storage device that performs deduplication processing when detecting a section where no write request has been received for a certain period or more after receiving the final write request.

[0004] Patent Document 2 relates to a storage device that executes deduplication processing in consideration of two or more deduplication mechanisms based on the state of the storage device.

[0005] Patent Document 3 relates to a data management system that specifies storage data that matches the time of the request target included in the request from the stored data and outputs the specification result.

[0006] Patent Document 4 relates to a file control device that solves the problem that there is a possibility that free capacity cannot be secured even when a file is transferred to a lower layer in a hierarchical storage having deduplication storage in the upper layer.

[0007] Patent Document 5 relates to a network distributed deduplication file storage including a plurality of information processing devices that constructs a computer network in a Peer to Peer manner, and disperses file data to each storage included in the information processing device and stores only non-duplicate data.

Prior Art Documents

Patent Documents

[0008] [Patent Document 1] Republished Patent No. 2017 / 149592 [Patent Document 2] Republished Patent No. 2014 / 125582 [Patent Document 3] Japanese Patent Publication No. 2020-067681 [Patent Document 4] Japanese Patent Publication No. 2019-053477 [Patent Document 5] Japanese Patent Publication No. 2018-190227 [Overview of the Initiative] [Problems that the invention aims to solve]

[0009] The following analysis is provided by the present invention.

[0010] Backup systems that allow data migration to the cloud tier are configured to free up capacity on the deduplication storage device by moving files to the cloud tier, enabling the writing of new files. Therefore, the backup system checks the amount of unique data for each file on the deduplication storage device (the amount of data that does not overlap with other data on the deduplication storage device and can be freed up by moving the file), and based on this file attribute information, it moves files to the cloud tier or suggests file migration to the user.

[0011] However, due to the nature of backup operations, which involve repeatedly writing data containing a large amount of overlap, moving files from a deduplication storage device to a cloud tier based on the attribute information of individual files may result in unnecessary data movement or unnecessary writing / deletion of unique data to the deduplication storage when performing repeated backup operations.

[0012] For example, suppose a deduplication storage device uses a rule to move files to the cloud tier. This rule involves creating a list of files on the deduplication storage device sorted by their unique data percentage (the ratio of data in a file that does not overlap with other data on the deduplication storage device) in descending order of unique data percentage, and then moving the files on this list to the cloud tier in descending order of unique data percentage until a certain amount of capacity on the deduplication storage device is freed up.

[0013] Using such rules, if a file newly designated by the user for backup has a very low duplication rate with the data on the deduplication storage (i.e., a high unique data rate), then if a command to move such backup target files to the cloud tier is issued immediately after moving (writing) them to the deduplication storage device, there is a high probability that the high unique data rate files that were written to the deduplication storage device immediately before the command will be moved to the cloud tier.

[0014] However, data in files that users have added as backup targets is likely to remain backup targets for a certain period of time. In such cases, each time a backup is performed, files with a high percentage of unique data are moved to the cloud tier, and then the same unique data is written again to the deduplication storage.

[0015] Therefore, while moving the initial files to the cloud tier temporarily freed up capacity on the deduplication storage device, the significance of moving the initial files to the cloud tier diminishes because the freed capacity will be used again by the same unique data in subsequent backups.

[0016] Furthermore, writing and deleting unique data to a deduplication storage device consumes a significant amount of computing resources. Therefore, repeatedly writing and deleting the same unique data from the same file to a deduplication storage device, as in the example above, is computationally inefficient, leading to wasted computing resources and a decrease in write performance.

[0017] The present invention aims to provide a deduplication storage device, a file hierarchical movement method, and a program that contribute to reducing operational inefficiencies such as writing / deleting extra unique data to a deduplication storage device and moving extra data to a cloud tier in a backup system that repeatedly backs up similar data. [Means for solving the problem]

[0018] According to a first aspect of the present invention, an on-premises deduplication storage device for a backup system including a cloud tier, At least two file storage pools for storing files, When moving files from the file storage pool to the cloud tier of the backup system, the system includes a file selection means for selecting files to be moved, which uses attribute information relating to the unique data of the files written to the file storage pool to create a list of move candidates for all files written to at least two of the file storage pools. A data deduplication storage device can be provided, characterized in that the unique data is data that does not overlap with other data on the data deduplication storage device.

[0019] According to a second aspect of the present invention, in a deduplication storage device including at least two on-premises file storage pools of a backup system including a cloud tier, Executed by a computer having a processor and memory, The steps include storing a file in the aforementioned file storage pool, When moving a file from the file storage pool to the cloud tier of the backup system, using the attribute information regarding the unique data of the file written in the file storage pool, for all the files written in at least two of the file storage pools, a step of creating a movement candidate list for selecting the file to be moved, is included. The unique data is data that does not overlap with other data on the deduplication storage device. A method for hierarchical movement of files, characterized by this, can be provided.

[0020] According to a third aspect of the present invention, in a computer provided in a deduplication storage device including at least two file storage pools on-premises of a backup system including a cloud tier a process of storing a file in the file storage pool, and when moving a file from the file storage pool to the cloud tier of the backup system, using the attribute information regarding the unique data of the file written in the file storage pool, for all the files written in at least two of the file storage pools, a process of creating a movement candidate list for selecting the file to be moved, is executed. The unique data is data that does not overlap with other data on the deduplication storage device. A program characterized by this can be provided. Note that this program can be recorded on a computer-readable storage medium. The storage medium can be a non-transient one such as a semiconductor memory, a hard disk, a magnetic recording medium, an optical recording medium, etc. The present invention can also be embodied as a computer program product.

Advantages of the Invention

[0021] According to the present invention, in a backup system that repeatedly backs up similar data, it is possible to provide a deduplication storage device, a file hierarchical movement method, and a program that contribute to reducing operational inefficiencies such as writing / deleting extra unique data to a deduplication storage device and moving extra data to a cloud tier. [Brief explanation of the drawing]

[0022] [Figure 1] This figure shows an example configuration of a backup system and a deduplication storage device for the backup system according to one embodiment of the present invention. [Figure 2] This figure shows an example of the configuration of a backup system and a deduplication storage device for the backup system according to the first embodiment of the present invention. [Figure 3] This figure shows an example configuration of a typical backup system and a deduplication storage device for that backup system. [Figure 4] This flowchart shows an example of the operation of a deduplication storage device in a backup system according to the first embodiment of the present invention. [Figure 5] This figure shows an example of backup operation using a backup system that utilizes a deduplication storage device according to the first embodiment of the present invention. [Figure 6] This figure shows an example of backup operation using a backup system that does not use a deduplication storage device according to the first embodiment of the present invention. [Figure 7] This figure shows the configuration of the computer that constitutes the deduplication storage device of the present invention. [Modes for carrying out the invention]

[0023] First, an overview of one embodiment of the present invention will be described with reference to the drawings. The reference numerals in the drawings attached to this overview are provided for convenience as examples to aid understanding and are not intended to limit the present invention to the illustrated embodiment. Furthermore, the connecting lines between blocks in the drawings and other references in the following description include both bidirectional and unidirectional lines. Unidirectional arrows schematically represent the flow of the main signal (data) and do not exclude bidirectional flow.

[0024] Figure 1 shows an example configuration of a backup system and a deduplication storage device for the backup system according to one embodiment of the present invention. Referring to Figure 1, the backup system 30 includes an on-premises deduplication storage device 40 and a cloud tier 50 in the cloud. The on-premises deduplication storage device 40 and the cloud tier 50 in the cloud are connected via a network 72.

[0025] The deduplication storage device 40 includes at least two file storage pools 405A and 405B for storing files, and a file selection means 406 for selecting files to be moved when moving files from file storage pools 405A and 405B to the cloud tier 50 of the backup system 30. This means uses attribute information about the unique data of the files written to the file storage pools to create a list of move candidates 4061 for selecting files to be moved from all files written to file storage pools 405A and 405B. Unique data refers to data that does not overlap with other data on the deduplication storage device 40. Although Figure 1 shows two file storage pools 405A and 405B, it is not intended to limit the number of file storage pools to two.

[0026] According to the deduplication storage device 40 of one embodiment of the present invention, when moving files from the deduplication storage device 40 to the cloud tier 50, a list of move candidates 4061 for selecting files to move is created for all files written to file storage pools 405A and 405B, using attribute information relating to the unique data of the files written to file storage pools 405A and 405B. Therefore, files can be moved not only using the attribute information of individual files, but also using rules based on the respective attributes of file storage pools 405A and 405B of the deduplication storage device 40.

[0027] As a result, the deduplication storage device 40 of one embodiment of the present invention can provide a deduplication storage device 40 that contributes to reducing operational inefficiencies such as extra unique data writing / deletion processing to the deduplication storage device 40 and extra data movement to the cloud tier 50.

[0028] [First Embodiment] Next, a data deduplication storage device 40 of the first embodiment of the present invention will be described with reference to the drawings. Figure 2 is a diagram showing an example of the configuration of a backup system 30 and a data deduplication storage device 40 of the backup system 30 of the first embodiment of the present invention. In Figure 2, components with the same reference numerals as in Figure 1 represent the same components.

[0029] Referring to Figure 2, the backup system 30 includes an on-premises deduplication storage device 40 and a cloud tier 50 in the cloud. The on-premises deduplication storage device 40 and the cloud tier 50 are connected via network 72. The backup system 30 is connected to the backup server 20 via network 71, and the backup server 20 is connected to the backup target environment 10 via network 70.

[0030] The backup target environment 10 includes multiple backup target files 101-1, 101-2, and 101-3 (referred to as backup target file 101 if they are not distinguished from each other), and the backup server 20 includes a file write / read means 201.

[0031] The deduplication storage device 40 of the first embodiment of the present invention includes inter-hierarchical data movement means 401, a deduplication data storage area 402, a deduplication table 403, file splitting / deduplication means 404, file storage pools 405A and 405B (referred to as file storage pool 405 if not distinguished from each other), a file to be moved selection means 406, a unique data writing recording means 407, and a weight management table 408. Although two file storage pools 405A and 405B are shown in Figure 2, it is not intended to limit the number of file storage pools to two.

[0032] The cloud tier 50 includes a deduplicated data storage area 502, a deduplication table 503, a file splitting / deduplicating means 504, and, as an example, file storage pools 505A and 505B. Although Figure 2 shows two file storage pools 505A and 505B, it is not intended to limit the number of file storage pools to two.

[0033] [General operation of backup systems] First, the general operation of the backup system will be explained with reference to Figure 3. Figure 3 shows an example of a typical backup system and the configuration of the backup system's deduplication storage device 40. In Figure 3, components with the same reference numerals as in Figure 2 represent the same components.

[0034] In backup operations, the user reads one or more data files (one or more backup target files 101-1, 101-2, 101-3) from the backup target environment 10 using the file write / read means 201 of the backup server 20 and writes them to the backup system 30. Conversely, if file restoration is required, the backup server 20 reads the necessary files from the backup system 30 and writes them to the backup target environment 10 or another restoration destination.

[0035] Generally, backups repeatedly write data containing many duplicates to the backup system 30 at regular intervals, so the backup data (a collection of backup target files 101-1, 101-2, 101-3 written at once) often overlaps with previously written data. Such a backup system 30 stores the backup data in a deduplication storage device 40 to maximize capacity efficiency.

[0036] The deduplication storage device 40 includes one or more file storage pools 405A and 405B on which users can write files. When a user writes a file, the file is divided into deduplication-enabled blocks by the file splitting / deduplication means 404, compared with data already written to the deduplication storage device 40 from the deduplication table 403, and capacity is allocated only to unique, non-duplicate blocks for storage in the deduplication-deduplication data storage area 402. All data in file storage pools 405A and 405B of the deduplication storage device 40 constitutes a single deduplication domain, and therefore there is no duplication of unique blocks between the respective pools within file storage pools 405A and 405B.

[0037] Furthermore, in such a backup system 30, for the long-term storage of numerous backup target files 101-1, 101-2, and 101-3, the deduplication storage device 40 may be able to move the data to a data storage area built on cloud storage via the Internet, i.e., a cloud tier 50. In this case, after moving the files from the deduplication storage device 40 to the cloud tier 50, either manually or according to rules based on file attribute information, it is assumed that the backup data on the on-premises deduplication storage device 40 will be deleted to free up capacity for storing new backups.

[0038] Furthermore, the cloud tier 50 includes a deduplicated data storage area 402, a deduplicated table 403, a file splitting / deduplicating means 404, and file storage pools 405A and 405B, which correspond to the deduplicated data storage area 402, a deduplicating table 503, a file splitting / deduplicating means 404, and file storage pools 505A and 505B, respectively, in order to store data written from the deduplicating storage device 40 as deduplicated data.

[0039] [Operation of the deduplication storage device according to the first embodiment of the present invention] Next, the operation of the deduplication storage device 40 of the backup system 30 of one embodiment of the present invention shown in Figure 2 will be described with reference to Figure 4. The deduplication storage device 40 of the backup system 30 of one embodiment of the present invention shown in Figure 2 has a configuration that includes a file selection means 406, a unique data writing recording means 407, and a weight management table 408, compared to the general backup system and the deduplication storage device 40 of the backup system shown in Figure 3. In one embodiment of the present invention, attribute information held by each file storage pool 405A and 405B (and the set of files contained therein) is used to select the files to be moved when moving files from the deduplication storage device 40 to the cloud tier 50.

[0040] Furthermore, by using NFS file systems, SMB shares, etc., users can create multiple file storage pools 405A and 405B on the deduplication storage device, and file storage pools 405A and 405B refer to endpoints that can write to and read files from the network.

[0041] In many backup operations, users configure multiple file storage pools, subdivided by each backup target environment 10 and each backup operation method. As a result, data characteristics may emerge between file storage pools, and the first embodiment of the present invention utilizes this characteristic.

[0042] In other words, ideally, file selection for file movement to cloud tier 50 should be based not on individual files, but on the entire set of files written in a single backup, and it is considered effective to make decisions using attribute information from this set. However, the files that make up each backup target environment 10, as well as the backup method and settings, are information held by the backup server 20. Therefore, since the backup system 30 writes data on a file-by-file basis, it is necessary to determine which files to select based on file-by-file information. In the first embodiment of the present invention, the trend of files written to file storage pools 405A and 405B is used to indirectly select files to move to cloud tier 50 based on information held by the backup server 20, such as the backup target environment 10.

[0043] The attribute information stored in the file storage pool is, for example, information that reflects the amount of unique data written by the deduplication storage device 40 when files are written to the file storage pool within a certain period of time. In the first embodiment of the present invention, the attribute information stored in the file storage pool is retained and used for selecting files to move to the cloud tier 50.

[0044] Next, the process of selecting files to move to cloud tier 50 using the amount of unique data written to file storage pools 405A and 405B within a certain period, and the unique data rate of each file, will be explained with reference to Figure 4.

[0045] Figure 4 is a flowchart showing an example of the operation of the deduplication storage device 40 of the backup system 30 of the first embodiment of the present invention shown in Figure 2.

[0046] Next, referring to Figure 4, the operation of the file selection means 406, the unique data writing and recording means 407, and the weight management table 408 of the deduplication storage device 40 will be explained.

[0047] In the deduplication storage device 40 shown in Figure 2, the unique data writing recording means 407 constantly monitors the writing capacity (data amount) and time of unique data, which is data from files written to each file storage pool 405A and 405B that does not overlap with data on the deduplication storage device 40, and stores it in the weight management table 408.

[0048] Furthermore, the unique data writing recording means 407 calculates a weight (W) for each file storage pool 405A and 405B, which is a value proportional to the amount of unique data written to each file storage pool 405A and 405B within a predetermined period, based on the amount of unique data of the file written to the weight management table 408 and the time the writing was performed. The calculated weights for file storage pools 405A and 405B are then stored in the weight management table 408.

[0049] In step S101 of Figure 4, the user instructs the deduplication storage device 40 to move files (data) from the deduplication storage device 40 to the cloud tier 50 by specifying the amount of data to be moved.

[0050] Upon receiving a move instruction to cloud tier 50 (step S101), the file selection means 406 of the deduplication storage device 40 starts creating a move candidate list 4061 in step S102, which contains the move evaluation value (M) for each file on the deduplication storage device 40, in order to determine which files to move.

[0051] The file selection means 406 performs the processing from steps S104 to S106 between steps S103 and S107 for each file on the deduplication storage device 40 in order to create the list of candidate files to be moved 4061. The processing from steps S104 to S106 will be described below.

[0052] In step S104, the file selection means 406 obtains the amount of data that can be reduced by moving the target file (the amount of unique data) from the weight management table 408, and calculates the ratio of the amount of unique data to the total amount of data of the target file (hereinafter referred to as the unique data rate (U)).

[0053] In step S105, the file selection means 406 calculates a file movement evaluation value (M) from the weight management table 408, using the weight (W) of the file storage pool 405A or 405B in which the target file is stored and the unique data rate (U) obtained in step S104. An example of the formula for calculating the movement evaluation value (M) is shown below: M = Unique data rate of the file (U) ÷ The weight (W) of the file storage pool containing that file

[0054] Next, in step S106, the file selection means 406 adds the target file and the move evaluation value (M) to the move candidate list 4061 based on the move evaluation value (M). The file selection means 406 adds the target file to the appropriate position in the move candidate list 4061 so that the move candidate list 4061 is sorted in descending order of move evaluation value.

[0055] In step S107, the file selection means 406 of the deduplication storage device 40 completes adding all files and their move evaluation values ​​(M) to the move candidate list 4061 in descending order of their move evaluation values. In step S108, it selects files to be moved from the move candidate list 4061, starting with the files with the highest move evaluation values, until the specified amount of data to be moved is satisfied.

[0056] Finally, in step S109, the deduplication storage device 40 moves the file set containing the selected files to be moved to the cloud tier 50.

[0057] [An example of backup operation using a backup system that utilizes the deduplication storage device 40 of the first embodiment of the present invention] Next, an example of backup operation using a backup system 30 that utilizes the deduplication storage device 40 of the first embodiment of the present invention shown in Figure 2 will be described. In backup operations such as those in which similar files are written to the backup system 30 on a regular basis, if the files are moved to the cloud tier 50 immediately after being written, using the deduplication storage device 40 of the first embodiment of the present invention may reduce the amount of unique data written and the amount moved to the cloud tier 50.

[0058] Figure 5 shows an example of backup operation using a backup system that utilizes a deduplication storage device according to the first embodiment of the present invention.

[0059] In the deduplication storage device 40, there are multiple file storage pools, including file storage pool 405A (hereinafter also referred to as pool 405A) and file storage pool 405B (hereinafter also referred to as pool 405B), which store the files File A1, File A2, and File B1 as described in step 1 of Figure 5, respectively. Although Figure 5 shows two file storage pools 405A and 405B, the number of file storage pools is not limited to two.

[0060] The file is divided into blocks by the file splitting / deduplication means 404, and the unique data rate of the file refers to the proportion of the total data volume (capacity) of the file that is unique to that file and does not overlap with other files.

[0061] When a file is written to the deduplication storage device 40, the unique data writing recording means 407 records the amount of unique data and the writing time of the file in the weight management table 408. The weights (W) of file storage pools 405A and 405B are calculated to increase in proportion to the amount of unique data written to file storage pool 405A or 405B within a predetermined period. The weight management table 408 can determine the weight (W) for each file storage pool 405A and 405B using the recorded amount of unique data and the writing time of the file write. At step 1, the weight values ​​for pools 405A and 405B are set to W=0. Also, the size of all files in this example of backup operation is assumed to be uniform.

[0062] Furthermore, in an example of backup operation using a backup system that utilizes the deduplication storage device of the first embodiment of the present invention, the file selection means 406 to be moved uses the following formula when calculating the move evaluation value (M) of each file using a weight (W) (step S105 in Figure 4). M = Unique data rate of the file (U) ÷ The weight (W) of the file storage pool containing that file

[0063] In step 1 of Figure 5, when an instruction is given to the deduplication storage device 40 to move a file to the cloud tier 50, File B1 in pool 405B is selected. In step 2, when the file (data) is moved to the cloud tier 50, File B1 is deleted from pool 405B of the deduplication storage device 40.

[0064] Next, in step 3, another backup is performed, and File B1_2, which has a large portion (over 99%) of its data overlap with File B1, is written to pool 405B. As a result, the unique data contained in File B1_2 is written again to the deduplication storage device 40. At this time, because the unique data is written, the weight of pool 405B changes from W=0 to W=X (X > 1). If X is sufficiently large, the move evaluation value (M) of File B1_2 becomes smaller than the move evaluation value (M) of File A2, and File A2 in pool 405A is selected as the target for move.

[0065] In step 4, File A2, which was selected in step 3, is moved to cloud tier 50 and deleted from pool 405A.

[0066] Finally, in step 5, the backup is performed again, and File B1_3, which has a large portion (over 99%) of its data overlap with File B1_2, is written to pool 405B, but no unique data is written.

[0067] Thus, in this example of backup operation using a backup system 30 that utilizes the deduplication storage device 40 of the first embodiment of the present invention, the writing of unique data to the deduplication storage device 40 occurs only once in step 3.

[0068] In contrast to the above, an example of the same operation without using the attribute information (weight W) of file storage pools 405A and 405B will be described. Figure 6 is a diagram showing an example of backup operation using a backup system that does not use a deduplication storage device according to the first embodiment of the present invention.

[0069] Referring to Figure 6, if the deduplication storage device of the first embodiment of the present invention is not used, weights (W) are not used, so the movement evaluation value (M) = unique data rate (U).

[0070] Referring to Figure 6, in step 3, File B1_2 in file storage pool 405B is selected to determine whether or not to move it, using only the unique data rate, and in step 4, File B1_2 is moved to cloud tier 50.

[0071] Next, in step 5, File B1_3, which has a large portion (over 99%) of its data overlap with File B1_2, is written to pool 405B, and the writing of unique data occurs.

[0072] In this example of backup operation using a backup system that does not use a deduplication storage device according to the first embodiment of the present invention, unique data is written to the deduplication storage device 40 twice.

[0073] This reveals that not using the weight (W), which is attribute information for file storage pools 405A and 405B, can result in the writing of unnecessary unique data.

[0074] According to the first embodiment of the present invention, when moving files from the deduplication storage device 40 to the cloud tier 50, the files are moved using rules based not only on the attribute information of individual files but also on the attribute information of the file storage pools 405A and 405B of the deduplication storage device 40. This reduces inefficiencies in backup operations, such as writing extra unique data to the deduplication storage device 40 and moving extra files (data) to the cloud tier 50, when repeatedly backing up similar data.

[0075] Although embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above, and further modifications, substitutions, and adjustments can be made without departing from the basic technical idea of ​​the present invention. For example, the network configuration, the configuration of each element, and the message representation form shown in each drawing are examples to aid in understanding the present invention, and are not limited to the configurations shown in these drawings. Also, in the following description, "A and / or B" means at least one of A or B.

[0076] Furthermore, the procedure shown in the first embodiment described above can be implemented by a program that enables a computer (9000 in Figure 7) functioning as the deduplication storage device 40 of the present invention to perform the functions of the deduplication storage device 40. Such a computer is exemplified by a configuration comprising a CPU (Central Processing Unit) 9010, a communication interface 9020, a memory 9030, and an auxiliary storage device 9040, as shown in Figure 7. That is, the CPU 9010 in Figure 7 executes the control program for the deduplication storage device 40 and performs the update process of each calculation parameter held in the auxiliary storage device 9040, etc.

[0077] Memory 9030 refers to RAM (Random Access Memory), ROM (Read Only Memory), etc.

[0078] In other words, each part (processing means, function) of the deduplication storage device 40 shown in the first embodiment described above can be realized by a computer program that causes the computer's processor to execute each of the above-described processes using its hardware.

[0079] Finally, preferred embodiments of the present invention are summarized. [First form] (See the first perspective above regarding data deduplication storage devices.) [Second form] The deduplication storage device described in the first embodiment is: The attribute information relating to the unique data of the aforementioned file is a file movement evaluation value obtained by dividing the unique data rate of the file, which indicates the ratio of the amount of data of the unique data of the aforementioned file to the total amount of data of the aforementioned file, by the weight of the file storage pool in which the aforementioned file is stored. Preferably, the weight of the file storage pool is a value proportional to the amount of unique data of the files written to the file storage pool within a predetermined period of time. [Third form] The deduplication storage device described in the second embodiment is: Weight management table, The system further includes a unique data writing recording means that monitors the amount of unique data of the file written to the file storage pool and the time the writing was performed, stores this information in the weight management table, and calculates the weight for the file storage pool from the amount of unique data and the time the writing was performed, and stores the weight in the weight management table, Preferably, the file selection means for moving files calculates the move evaluation value of the file based on the amount of data in the file, the amount of data in the unique data of the file stored in the weight management table, and the weights for the file storage pool. [Fourth form] The deduplication storage device described in the second or third embodiment is: Preferably, the file selection means for moving files adds all files and their respective move evaluation values ​​to the move candidate list in descending order of their move evaluation values, and once the addition is complete, it selects files to be moved from the move candidate list in descending order of their move evaluation values ​​until the specified amount of move data is satisfied. [Fifth form] The deduplication storage device according to any one of the first to fourth embodiments is: It is preferable to move the file set containing the selected files to be moved to the cloud tier. [Sixth form] (See the second perspective above for how to navigate file hierarchies.) [Seventh form] The file hierarchy navigation method described in the sixth form is: The attribute information relating to the unique data of the aforementioned file is a file movement evaluation value obtained by dividing the unique data rate of the file, which indicates the ratio of the amount of data of the unique data of the aforementioned file to the total amount of data of the aforementioned file, by the weight of the file storage pool in which the aforementioned file is stored. Preferably, the weight of the file storage pool is a value proportional to the amount of unique data of the files written to the file storage pool within a predetermined period of time. [Eighth form] The file hierarchy navigation method described in the seventh form is: The steps include monitoring the amount of data of the unique data of the file written to the file storage pool and the time the writing was performed, and storing them in a weight management table. A step of calculating the weight for the file storage pool from the amount of data of the unique data and the time at which it was written, The process further includes the step of storing the weights in the weight management table, Preferably, the step further includes calculating the movement evaluation value of the file based on the amount of data in the file, the amount of data in the unique data of the file stored in the weight management table, and the weights for the file storage pool. [Ninth form] (See the program from the third perspective above) [Tenth form] The program described in the ninth form is The attribute information relating to the unique data of the aforementioned file is a file movement evaluation value obtained by dividing the unique data rate of the file, which indicates the ratio of the amount of data of the unique data of the aforementioned file to the total amount of data of the aforementioned file, by the weight of the file storage pool in which the aforementioned file is stored. Preferably, the weight of the file storage pool is a value proportional to the amount of unique data of the files written to the file storage pool within a predetermined period of time. Furthermore, the sixth form described above can be expanded from the fourth to the fifth form, similar to the first form, and the ninth form described above can be expanded from the third to the fifth form, similar to the first form.

[0080] Furthermore, each disclosure in the above-mentioned patent documents is incorporated into this book by reference. Within the framework of the full disclosure of the present invention (including the claims), further modifications and adjustments to the embodiments or examples are possible based on the fundamental technical concept. Also, within the framework of the disclosure of the present invention, various combinations or selections of various disclosed elements (including each element of each claim, each element of each embodiment or example, each element of each drawing, etc.) are possible. In other words, the present invention naturally includes various modifications and changes that a person skilled in the art could make in accordance with the full disclosure, including the claims, and the technical concept. In particular, with respect to the numerical ranges described in this book, any numerical value or sub-range included within that range should be interpreted as being specifically described, even if not otherwise stated. [Explanation of Symbols]

[0081] 10 Backup Target Environment 20 Backup Servers 30 Backup Systems 40 Deduplication Storage Devices 50 Cloud Tiers Networks 70, 71, and 72 101-1, 101-2, 101-3 Backup target files 201 File writing / reading methods 401 Inter-hierarchical data transfer method 402 Deduplication-free data storage area 403 Duplicate Removal Table 404 File splitting / deduplication methods 405, 405A, 405B File Storage Pool 406 Selection method for files to be moved 407 Unique data writing and recording means 408 Weight Management Table 502 Deduplication-free data storage area 503 Duplicate Removal Table 504 File splitting / deduplication methods 505A, 505B File Storage Pool 4061 List of potential moves 9000 Computers 9010 CPU 9020 Communication Interface 9030 memory 9040 Auxiliary storage device

Claims

1. An on-premises deduplication storage device for a backup system including a cloud tier, At least two file storage pools for storing files, When moving files from the file storage pool to the cloud tier of the backup system, the system includes a file selection means for selecting files to be moved from among all the files written to at least two of the file storage pools, using attribute information relating to the unique data of the files written to the file storage pool. The deduplication storage device is characterized in that the unique data is data that does not overlap with other data on the deduplication storage device, The attribute information relating to the unique data of the aforementioned file is a file movement evaluation value obtained by dividing the unique data rate of the file, which indicates the ratio of the amount of data of the unique data of the aforementioned file to the total amount of data of the aforementioned file, by the weight of the file storage pool in which the aforementioned file is stored. A deduplication storage device characterized in that the weight of the file storage pool is a value proportional to the amount of unique data of files written to the file storage pool within a predetermined period of time.

2. Weight management table, The system further includes a unique data writing recording means that monitors the amount of unique data of the file written to the file storage pool and the time the writing was performed, stores this information in the weight management table, and calculates the weight for the file storage pool from the amount of unique data and the time the writing was performed, and stores the weight in the weight management table, The deduplication storage device according to claim 1, characterized in that the file to be moved selection means calculates the move evaluation value of the file based on the amount of data of the file, the amount of data of the unique data of the file stored in the weight management table, and the weight for the file storage pool.

3. The deduplication storage device according to claim 1 or 2, wherein the file selection means for moving files adds the file and the move evaluation value to the move candidate list for all files in descending order of the move evaluation value of the file, and when the addition is completed, it selects the files to be moved from the move candidate list in descending order of the move evaluation value until the specified amount of data to be moved is satisfied.

4. A deduplication storage device according to any one of claims 1 to 3, characterized in that it moves a file set including the selected files to be moved to the cloud tier.

5. In a deduplication storage device that includes at least two on-premises file storage pools for a backup system including a cloud tier, Executed by a computer equipped with a processor and memory, The steps include storing a file in the aforementioned file storage pool, When moving files from the file storage pool to the cloud tier of the backup system, the process includes the step of creating a list of move candidates for selecting files to be moved from among all the files written to at least two of the file storage pools, using attribute information relating to the unique data of the files written to the file storage pool. A method for moving files hierarchically, characterized in that the unique data is data that does not overlap with other data on the deduplication storage device, The attribute information relating to the unique data of the aforementioned file is a file movement evaluation value obtained by dividing the unique data rate of the file, which indicates the ratio of the amount of data of the unique data of the aforementioned file to the total amount of data of the aforementioned file, by the weight of the file storage pool in which the aforementioned file is stored. A method for moving files hierarchically, characterized in that the weight of the file storage pool is a value proportional to the amount of data of the unique data of the files written to the file storage pool within a predetermined period of time.

6. The steps include monitoring the amount of data of the unique data of the file written to the file storage pool and the time the writing was performed, and storing them in a weight management table. A step of calculating the weight for the file storage pool from the amount of data of the unique data and the time at which it was written, The process further includes the step of storing the weights in the weight management table, A step of calculating the movement evaluation value of the file based on the amount of data in the file, the amount of data in the unique data of the file stored in the weight management table, and the weights for the file storage pool. The file hierarchy movement method according to claim 5, further comprising the above.

7. A computer located on a deduplication storage device that includes at least two on-premises file storage pools of a backup system including a cloud tier, The process of storing files in the aforementioned file storage pool, When moving files from the file storage pool to the cloud tier of the backup system, the system performs a process to create a list of move candidates for selecting files to be moved from all the files written to at least two of the file storage pools, using attribute information relating to the unique data of the files written to the file storage pool. The program is characterized in that the unique data is data that does not overlap with other data on the deduplication storage device, The attribute information relating to the unique data of the aforementioned file is a file movement evaluation value obtained by dividing the unique data rate of the file, which indicates the ratio of the amount of data of the unique data of the aforementioned file to the total amount of data of the aforementioned file, by the weight of the file storage pool in which the aforementioned file is stored. A program characterized in that the weight of the file storage pool is a value proportional to the amount of unique data of files written to the file storage pool within a predetermined period of time.

Citation Information

Patent Citations

  • Movement controller, program and storage device

    JP2013196190A

  • Block storage

    JP2017130103A

  • Network distributed duplication exclusion file storage system

    JP2018190227A

  • File control device, file control method, and program

    JP2019053477A

  • Storage system and control method

    JP2019074912A