File storage

The file storage system enhances data reduction rates by dynamically selecting a compression algorithm based on the data volume and time constraints, thereby maintaining optimal response performance during data storage.

JP7697838B2Active Publication Date: 2025-06-24HITACHI VANTARA LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021115973
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-13
Publication Date
2025-06-24
Estimated Expiration
2041-07-13

AI Technical Summary

Technical Problem

Existing file storage systems face challenges in increasing data reduction rates without compromising response performance during data storage, particularly when compressing image data on a file-by-file basis.

Method used

A file storage system that includes a processor which receives a write request for a file, writes the data to a storage device, and later compresses the data using a determined compression algorithm based on the amount of data written within a predetermined time, thereby optimizing data reduction without degrading response performance.

Benefits of technology

The proposed solution effectively increases the data reduction rate without deteriorating response performance during data storage, by selecting an appropriate compression algorithm based on the data volume and time constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697838000001
    Figure 0007697838000001
  • Figure 0007697838000002
    Figure 0007697838000002
  • Figure 0007697838000003
    Figure 0007697838000003
Patent Text Reader

Abstract

To provide file storage capable of increasing the reduction rate of data without degrading the response performance at the time when the data is stored.SOLUTION: The file storage includes a processor that is configured to receive a request for writing the same on a file from an application, to write the data of the file to a storage device, and later, to compress the data of the written file and write the same to a storage device. The processor determines the compression algorithm to be used for compression depending on the amount of data written in the predetermined time of one or more files on which the writing has been made.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a file storage having a batch compression function using a flash memory or a magnetic disk as a storage (storage medium).

Background Art

[0002] Patent Document 1 is a patent document related to an image-related compression algorithm. In recent years, with the explosive expansion of the data volume, the development of data volume reduction technologies has been actively carried out. In particular, research on image-related compression algorithms with a large data volume has been active. The feature of these compression algorithms is that data loss due to irreversible compression can be suppressed by specializing it for specific applications. For example, an image compressor can be created so that the data loss is difficult for human recognition.

[0003] The most important in the compression algorithm is the compression ratio, which is the data reduction rate, but the compression speed is also important. Generally, when trying to improve the compression ratio, the compression speed decreases. Also, the relationship between the increase and decrease of the compression ratio and the increase and decrease of the compression speed is not linear. When trying to improve the compression ratio, the compression speed decreases rapidly. Also, the decompression speed when reading data generally becomes slower as the compression ratio is higher.

[0004] Patent Document 2 discloses an example of selecting a suitable compression algorithm according to the access frequency in a storage having a plurality of compression algorithms with different compression and decompression processing times.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] Compression of image data is often performed on a file-by-file basis. The reason is that, on a file-by-file basis, the type of data, whether it is still image data, video data, or audio data, is determined. Depending on the type of data, the compression algorithm to be applied is determined. Therefore, by enabling a file storage that stores and reads data on a file-by-file basis to recognize the type of data, compression on a file-by-file basis becomes possible.

[0007] In this case, it is desirable to apply the compression algorithm with the highest compression rate, but there are restrictions on the compression speed. In particular, when data is stored in a file storage and compression processing is performed, the response performance as seen from the application may deteriorate significantly.

[0008] The present invention has been made in consideration of the above points, and intends to propose a file storage or the like that can increase the data reduction rate without deteriorating the response performance during data storage.

Means for Solving the Problems

[0009] In order to solve such problems, in the present invention, a file storage including a processor that receives a write request for a file from an application, writes the data of the file to a storage device, and later compresses and writes the data of the written file to the storage device, wherein the processor determines a compression algorithm to be used for compression according to the amount of data written in a predetermined time for one or more written files.

[0010] According to the above configuration, since the data of the written file is compressed later, for example, the data reduction rate can be increased without deteriorating the response performance during data storage.

Effects of the Invention

[0011] According to the present invention, it is possible to increase the data reduction rate without degrading the response performance during data storage.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

[0013] Hereinafter, an embodiment of the present invention will be described in detail. However, the present invention is not limited to the embodiments.

[0014] In view of the data reduction rate in the file storage, it is desirable to apply the compression algorithm with the highest compression rate, but there are restrictions on the compression speed. In particular, when storing data in the file storage, if the compression process is executed, the response performance seen from the application may deteriorate significantly.

[0015] Also, if compression is performed using a compression algorithm with a compression speed below the data generation speed over a certain period of time, the compression will not be completed in time, uncompressed data will accumulate, and capacity reduction will not be possible.

[0016] Also, when reading the compressed data, if the decompression speed is slow, the response performance seen from the application may deteriorate significantly, similar to the case of storage.

[0017] In this embodiment, the problem of deterioration of the response performance during data storage is solved by the file storage performing the compression process later in a batch process.

[0018] Also, by preparing a plurality of compression algorithms with different compression speeds, grasping the data generation amount per unit time of the file group for which the compression process is to be executed, and selecting a compression algorithm from among the compression algorithms that can complete the compression process within the allowable time, a high data reduction rate can be achieved.

[0019] In addition, to address the performance degradation of the read process, a cache area is provided in the file storage, and the extended file is stored in the cache area. When a read request is received, if the file hits in the cache area, the extended data is directly read from the cache area. This solves the problem of performance degradation in reading files with high read frequencies.

[0020] Next, embodiments of the present invention will be described with reference to the drawings. The following description and drawings are examples for explaining the present invention, and for clarity of explanation, appropriate omissions and simplifications are made. The present invention can also be implemented in various other forms. Unless otherwise limited, each component may be singular or plural.

[0021] The notations such as "first", "second", "third", etc. in this specification and the like are for identifying components and do not necessarily limit the number or order. Also, the numbers for identifying components are used for each context, and the numbers used in one context do not necessarily indicate the same configuration in other contexts. Further, a component identified by a certain number is not precluded from also having the functions of a component identified by another number.

[0022] Figure 1 shows the configuration of the information system according to the present invention. The information system is composed of one or more file storages 100, one or more servers 110, and a network 120 that connects the file storage 100 and the server 110. The server 110 is connected to the network 120 through a server port 195, and the file storage 100 is connected to the network 120 through a storage port 197. The server 110 has one or more server ports 195, and the file storage 100 has one or more storage ports 197 connected to the network 120. The server 110 is a system in which the user application 140 operates, and reads and writes necessary data to and from the file storage 100 via the network 120 according to the request of the user application 140. The protocol used in the network 120 is, for example, NFS or CIFS.

[0023] Figure 2 shows the configuration of the file storage 100. The file storage 100 is composed of one or more processors 200, a main memory 210, a shared memory 220, one or more connection devices 250 connecting these components, and a storage device 130. In this embodiment, the file storage 100 includes the storage device 130 and reads and writes data directly to and from the storage device 130. However, the present invention is also effective in a configuration where the file storage 100 does not include the storage device 130 and specifies a logical volume (such as LUN) for the block storage including the storage device 130 to read and write data. Further, the present invention is also effective in a configuration where the file storage 100 is installed on the server 110 as software and operates on the same device as the user application 140. In this case, the storage device 130 is a device connected to the server 110. The storage device 130 includes storage devices 130 such as HDD (Hard Disk Drive) and flash storage using flash memory as a storage medium. Also, there are several types of flash storage, including high-price, high-performance SLC with a large number of erasable times and, in contrast, low-price, low-performance MLC with a small number of erasable times. Further, new storage media such as phase change memory may be included. The processor 200 processes read and write requests issued from the server 110. The main memory 210 stores programs executed by the processor 200, internal information of each processor 200, etc.

[0024] The connection device 250 is a mechanism for connecting each component within the file storage 100.

[0025] The shared memory 220 is composed of volatile memory such as normal DRAM, but is made non-volatile by a battery or the like. Also, in this embodiment, for high reliability, each is assumed to be duplicated. However, the present invention is effective even if the shared memory 220 is not made non-volatile or is not duplicated. Information shared between the processors 200 is stored in the shared memory 220.

[0026] In this embodiment, it is assumed that the file storage 100 does not have a RAID (Redundancy Array Independent Device) function that can recover the data of a device even if one of the devices in the storage device 130 fails. The present invention is also effective when the file storage 100 has a RAID function.

[0027] FIG. 3 shows information related to this embodiment in the shared memory 220 of the file storage 100 in this embodiment, which is composed of file storage information 2000, file information 2100, storage device information 2200, virtual page capacity 2300, free file information pointer 2400, free page information 2500, LRU head pointer 2600, LRU tail pointer 2700, total compression amount 2800, and total decompression time 2900.

[0028] Among these, as shown in FIG. 4, the file storage information 2000 is information regarding the file storage 100, and is composed of a file storage identifier 2001, a media type 2002, the number of algorithms 2007, a compression algorithm 2003, a compression rate 2004, compression performance 2005, and decompression performance 2006. In this embodiment, when issuing a read / write request according to an instruction from the user application 140, the server 110 shall specify the identifier of the file storage 100, the identifier of the file, the relative address within the file, and the data length (the length of the data to be read / written). The identifier of the file storage 100 specified in the read / write request is the file storage identifier 2001 included in the file storage information 2000. Further, in this embodiment, it is assumed that media information and compression information of the file are specified in the read / write request. Note that the media information and compression information of the file may be notified by other means, and the present invention is still effective. The present invention targets files storing media information such as videos and images for which a high compression rate can be expected, performs compression corresponding to the media, and implements data reduction. The media type 2002 represents the type of media (such as still images, videos, etc.) for which the file storage 100 performs compression. The number of algorithms 2007 indicates the number of compression algorithms that this file storage 100 has for the corresponding media type. The compression algorithm 2003 indicates the compression algorithm that the file storage 100 has. The compression rate 2004 and the compression performance 2005 indicate the compression rate and compression performance (speed) of the corresponding compression algorithm. Also, the decompression performance 2006 indicates the decompression performance (speed). The compression algorithm 2003, the compression rate 2004, the compression performance 2005, and the decompression performance 2006 will be repeated the number of times set in the number of algorithms 2007. After that, information regarding the media indicated by the next media type 2002 is set. The file storage 100 has one or more compression algorithms corresponding to the media type 2002. The media information specified in the read / write request represents the media type of the file, and the compression information indicates whether compression is performed or, if compression is being performed, the compression algorithm being used.

[0029] The feature of this embodiment is that the file storage 100 supports a capacity virtualization function. However, the present invention is also effective even if the file storage 100 does not have a capacity virtualization function. Usually, in the capacity virtualization function, the allocation unit of the storage area is called a page. In this embodiment, it is assumed that the space of the file is divided in units of virtual pages, and the storage device 130 is divided in units of physical pages. When the capacity virtualization function is realized, when no physical page is allocated to the virtual page including the address instructed to write in the write request from the server 110 to the file storage 100, a physical page is allocated. The virtual page capacity 2300 is the capacity of the virtual page. In this embodiment, the virtual page capacity 2300 is equal to the capacity of the physical page. However, the present invention is also effective even if the physical page includes redundant data and the virtual page capacity 2300 is not equal to the physical page capacity.

[0030] FIG. 5 shows the format of the file information 2100, which is composed of a file identifier 2101, a file size 2102, a file medium 2103, initial compression information 2104, applied compression information 2105, a compressed file size 2106, a receiving start pointer 2107, a receiving end pointer 2108, a compression start pointer 2109, a compression end pointer 2110, a cache start pointer 2111, a cache end pointer 2112, a next LRU pointer 2113, a previous LRU pointer 2114, an uncompressed flag 2115, a schedule flag 2116, a cache flag 2117, a next free pointer 2118, and an access address 2119.

[0031] In this embodiment, when the file storage 100 receives a read / write request from the server 110, it recognizes the corresponding file by the specified file identifier. The present invention targets files that store media information such as videos and images that can be expected to have a high compression rate. In addition, a characteristic of such files is that data is written in order from the first address when the file is created. For this reason, areas that have already been written are usually not rewritten. In addition, when reading a file, data is usually read from the beginning of the file to the end in address order.

[0032] The file identifier 2101 is the identifier of the file. The file size 2102 is the amount of data written to the file. The file medium 2103 represents the type of medium of the file, such as the type of video. The initial compression information 2104 indicates the compression state of the data initially written from the server 110. The initial compression information 2104 indicates whether compression is applied or not, and if it is compressed, the compression algorithm applied. In the present invention, later, a compression algorithm with a higher compression ratio than the initially applied compression algorithm is applied to improve the data reduction rate. The applied compression information 2105 indicates the compression algorithm to be applied later. The compressed file size 2106 indicates the file size when the applied compression information 2105 is applied. The reception start pointer 2107 and the reception end pointer 2108 indicate the first page and the last page storing the data received first. The compression start pointer 2109 and the compression end pointer 2110 indicate the first page and the last page storing the data compressed by the file storage 100. When the file storage 100 receives a read request for the data stored in the data compressed by the file storage 100, it is necessary to convert the data into the initially written data and then pass the data to the server 110. At this time, in the present invention, in order to ensure the response performance of files with high access frequencies, the converted data is stored in the cache area provided in the storage device 130. The cache start pointer 2111 and the cache end pointer 2112 indicate the first page and the last page of the data stored in the cache area. When such control is performed, it is necessary to expel the data of files with decreased access frequencies from the cache area. In the present invention, LRU management of the files storing data in the cache area is performed to determine the files to be expelled. The next LRU pointer 2113 and the previous LRU pointer 2114 are pointers to the file information 2100 of the file with one higher access frequency and the file with one lower access frequency than the file, respectively. The uncompressed flag 2115 is a flag indicating that the file storage 100 has not yet performed compression.The schedule flag 2116 is a flag indicating that the file is to be compressed. The cache flag 2117 indicates that the file is being stored in the cache area. In the present invention, when a write request is received for the address at the head of a file, it means that a write request for a new file has been received. Therefore, at this opportunity, it is necessary to allocate file information 2100. For this reason, it is necessary to manage the file information 2100 in an empty state. The next free pointer 2118 is a pointer to the file information in an empty state next. The access address 2119 indicates the address at which to perform the next read when reading the compressed data in the file storage 100. Since the length of the compressed data is variable length, generally, the data stored where the compressed data is stored cannot be calculated from the relative address specified in the read request. However, since media data and the like are accessed in address order, the next data to be accessed is the next address even in the compressed data space. If this is memorized, the address of the compressed data to be accessed in the next request can be recognized.

[0033] Figure 6 shows the storage device information 2200. The storage device information 2200 has a storage device identifier 2201, a storage capacity 2202, and actual page information 2203. The storage device identifier 2201 is the identifier of the storage device 130. The storage capacity 2202 is the capacity of the storage device 130. The actual page information 2203 is information corresponding to the actual pages included in the storage device 130, and the number thereof is the value obtained by dividing the storage capacity by the virtual page capacity.

[0034] FIG. 7 shows the format of the actual page information 2203. The actual page information 2203 is composed of a storage identifier 3000, a relative address 3001, and a next page pointer 3002. The storage identifier 3000 indicates the identifier of the storage device 130 of the corresponding actual page. The relative address 3001 indicates the relative address within the storage device 130 of the corresponding actual page. In the present invention, the actual page takes several states. It is in an empty state (unallocated) or an allocated state. In the allocated state, there are states of storing the data written first, storing the data compressed in the file storage 100, and storing in the cache area. In total, there are four states. Since the actual pages in the same state are connected by pointers, the next page pointer 3002 is a pointer to the next actual page information 2203 in the same state.

[0035] FIG. 8 shows the file information 2100 that becomes empty and is managed by the free file information pointer 2400. This queue is called the free file information queue 800. The free file information pointer 2400 indicates the first file information 2100 in the empty state. The next free pointer 2118 in the file information 2100 then indicates the file information 2100 in the empty state.

[0036] FIG. 9 shows the actual page information 2203 in the empty state managed by the free page information 2500. This queue is called the free actual page information queue 900. The free page information 2500 indicates the first actual page information 2203 in the empty state. The next page pointer 3002 in the actual page information 2203 then indicates the actual page information 2203 in the empty state.

[0037] In the present invention, the file storage 100 periodically performs compression processing on the data of the received files. The feature of the present invention is to grasp the amount of data that needs to be compressed and select a compression algorithm that can complete the compression processing by the next period. Thereby, within the range where the compression processing can be completed in time, the compression algorithm with the highest data reduction effect can be applied. The total compression amount 2800 is the amount of data that needs to be compressed in this period. Also, in the present invention, initially, it is allowed to receive the compressed data. In this case, when attempting to apply a compression algorithm with a higher compression ratio than the initial compression algorithm, it is necessary to decompress the data once. Therefore, actually, it is necessary to complete the compression processing including this decompression time. The total decompression time 2900 is the total value of the time taken for the decompression process.

[0038] FIG. 10 shows the management state of the file information 2100 to which the cache area managed by the LRU head pointer 2600 and the LRU tail pointer 2700 is allocated. This queue is called the file information LRU queue 1000. The file information 2100 indicated by the LRU head pointer 2600 is the file information 2100 of the file that has been recently read, and the file information 2100 indicated by the LRU tail pointer 2700 is the file information 2100 of the file that has not been read for the longest period. When a new file for which a cache area is to be allocated appears, the actual page is released from the file information 2100 indicated by the LRU tail pointer 2700 and returned to the free actual page managed by the free page information 2500 shown in FIG. 9.

[0039] Figure 11 shows the structure of the actual page information 2203 managed by the reception start pointer 2107 and the reception end pointer 2108. The reception start pointer 2107 indicates the actual page information 2203 that stores the data received first, that is, the data at the address of the head of the file. In the next page pointer 3002 of the actual page information 2203, the actual page information 2203 that stores the data at the next address of the file is indicated. The address of the actual page information 2203 that stores the data received last, that is, the data at the last address, is stored in the reception end pointer 2108.

[0040] Since the structures of the actual page information 2203 managed by the compression start pointer 2109 and the compression end pointer 2110, and the structures of the actual page information 2203 managed by the cache start pointer 2111 and the cache end pointer 2112 are the same as the structures shown in Figure 11 respectively, the description is omitted.

[0041] Next, the operation of the processor 200 of the file storage 100 will be described using the management information described above. The program executed by the processor 200 of the file storage 100 is stored in the main memory 210. Figure 12 shows the program related to the present embodiment stored in the main memory 210. The program related to the present embodiment includes a write processing unit 4000, a read processing unit 4100, and a compression processing unit 4200.

[0042] Figure 13 shows the processing flow of the write processing unit 4000. The processing flow of the write processing unit 4000 is the processing flow executed when a write request is received from the server 110.

[0043] Step 50000: Check whether the specified relative address is the address of the head of the file. If it is not the head, jump to step 50004.

[0044] Step 50001: Assign the file information 2100 indicated by the free file information pointer 2400 to the file. Set the value indicated by the next free pointer 2118 of the assigned file information 2100 in the free file information pointer 2400.

[0045] Step 50002: Set the file identifier, media type, and compression information specified in the write request in the file identifier 2101, file media 2103, and initial compression information 2104.

[0046] Step 50003: Make the physical page information 2203 in the free state indicated by the free page information 2500 such that both the reception start pointer 2107 and the reception end pointer 2108 of the file information indicate it. Also, set the information indicated by the next page pointer 3002 of the assigned physical page information 2203 in the free page information 2500. After that, jump to step 50005.

[0047] Step 50004: Find the corresponding file information 2100 from the file identifier specified in the write request.

[0048] Step 50005: Check whether the data can be stored only in the currently assigned physical pages based on the relative address and data length of the received write request. If it can be stored, jump to step 50007.

[0049] Step 50006: Make the physical page information 2203 in the free state indicated by the free page information 2500 (the physical page information 2203) such that the next page pointer 3002 of the physical page information 2203 indicated by the reception end pointer 2108 indicates it. Also, make the physical page information 2203 such that the reception end pointer 2108 indicates it. In addition, set the information indicated by the next page pointer 3002 of the physical page information 2203 (the assigned physical page information 2203) in the free page information 2500.

[0050] Step 50007: Receive the write data. Calculate from the relative address and the data length which address on which page the data should be written to.

[0051] Step 50008: Issue a write request to the storage device 130.

[0052] Step 50009: Wait for completion.

[0053] Step 50010: Update the file size 2102 from the received data length.

[0054] Step 50011: Report completion to the server 110.

[0055] Figure 14 shows the processing flow of the read processing unit 4100. The processing flow of the read processing unit 4100 is the processing flow executed when the file storage 100 receives a read request from the server 110.

[0056] Step 60000: Find the corresponding file information 2100 from the specified file identifier.

[0057] Step 60001: Check if the uncompressed flag 2115 is on. If it is on, jump to step 60018.

[0058] Step 60002: Check if the cache flag 2116 is on. If it is on, jump to step 60017.

[0059] Step 60003: Check if the relative address specified in the read request is the starting address. If not, jump to step 60005.

[0060] Step 60004: In the case of the head, set the address of the head of the actual page corresponding to the compression head pointer 2109 to the access address 2119. Also, move the actual page information 2203 assigned to the file information 2100 indicated by the LRU tail pointer 2700 shown in FIG. 10, that is, the actual page information 2203 existing between the cache head pointer 2111 and the cache tail pointer 2112 of the file information 2100, to the free actual page information queue 900 indicated by the free page information 2500. Also, turn off the cache flag 2117 of the file information 2100. Further, set the address of the file information 2100 indicated by the previous LRU pointer 2114 in the file information 2100 that the LRU tail pointer 2700 has shown so far to the LRU tail pointer 2700.

[0061] Step 60005: In the page storing the compressed data, issue a read request to the storage device 130 to read data from the address indicated by the access address 2119 and wait for completion.

[0062] Step 60006: Refer to the applied compression information 2105 etc. of the file information 2100 and convert the read data into the data received from the server 110.

[0063] Step 60007: Send the converted data to the server 110 and issue a completion report.

[0064] Step 60008: Check whether the specified relative address is the address of the head of the file. If it is not the head, jump to Step 60010.

[0065] Step 60009: Make the free actual page information 2203 in the free state indicated by the free page information 2500 as indicated by both the cache head pointer 2111 and the cache tail pointer 2112 of the file information. Also, set the information indicated by the next page pointer 3002 of the assigned actual page information 2203 in the free page information 2500. Also, move the file information 2100 to the position indicated by the LRU head pointer 2600 shown in FIG. 10.

[0066] Step 60010: Check whether the data can be stored only in the currently allocated physical page based on the relative address and data length of the received read request. If it can be stored, jump to step 60012.

[0067] Step 60011: Set the physical page information 2203 (the said physical page information 2203) in the free state indicated by the free page information 2500 to be as indicated by the next page pointer 3002 of the physical page information 2203 indicated by the cache end pointer 2112. Also, set the said physical page information 2203 as indicated by the cache end pointer 2112. In addition, set the information indicated by the next page pointer 3002 of the said physical page information 2203 (the allocated physical page information 2203) in the free page information 2500.

[0068] Step 60012: Calculate from the received relative address and data length which address on which page the data should be written.

[0069] Step 60013: Issue a write request to the storage device 130.

[0070] Step 60014: Wait for completion.

[0071] Step 60015: Update the access address 2119. Check whether the writing of the entire file is completed. If not, end the process.

[0072] Step 60016: If completed, turn on the cache flag 2117 and end the process.

[0073] Step 60017: Recognize the address of the physical page storing the data to be read by referring to the received relative address, the cache start pointer 2111, and the cache end pointer 2112. Jump to step 60019.

[0074] Step 60018: Recognize the address of the actual page storing the data to be read by referring to the received relative address, the start pointer 2107 at the time of reception, and the end pointer 2108 at the time of reception.

[0075] Step 60019: Issue a read request to the storage device 130.

[0076] Step 60020: Wait for the read to complete.

[0077] Step 60021: Send the read data to the server 110 and issue an end report. After this, end the process.

[0078] FIG. 15 shows the processing flow of the compression processing unit 4200. The processing flow of the compression processing unit 4200 is periodically activated in the file storage 100.

[0079] Step 70000: Initialize the total compression amount 2800 and the total expansion time 2900.

[0080] Step 70001: Find the file information 2100 with the uncompressed flag 2115 on. If the file information 2100 with the uncompressed flag 2115 on cannot be found, jump to step 70005.

[0081] Step 70002: Turn off the uncompressed flag 2115 of the found file information 2100 and turn on the schedule flag 2116. Add the file size 2102 to the total compression amount 2800.

[0082] Step 70003: If the initial compression information 2104 is no compression, jump to step 70001.

[0083] Step 70004: If there is compression, recognize the compression algorithm 2003 used from the file media 2103 and the initial compression information 2104, and recognize the speed at which this data is decompressed according to the corresponding decompression performance 2006. Further, add the value obtained by multiplying this speed by the file size 2102 (= decompression time) to the total decompression time 2900. After that, jump to step 70001.

[0084] Step 70005: Subtract the total decompression time 2900 from the time until the next schedule. Compression processing must be completed within the subtracted time. Divide the total compression amount 2800 by the subtracted value to calculate the required compression speed.

[0085] Step 70006: For each media type 2002, determine the compression algorithm to apply by selecting the compression algorithm 2003 with the highest compression ratio among those that satisfy the compression speed from the compression algorithms 2003 held by the file storage 100.

[0086] Step 70007: Find the file information 2100 with the schedule flag 2116 on. If not found, complete the process.

[0087] Step 70008: Refer to the file media 2103 and set the compression algorithm determined in step 70006 in the applied compression information 2105.

[0088] Step 70009: Read out the data stored in the actual page corresponding to the actual page information 2203 indicated by the reception start pointer 2107 and the reception end pointer 2108. Here, take the first data as the read target and proceed to the next step.

[0089] Step 70010: Issue a read request to the storage device 130 to read the data to be read. Also, calculate the address of the data to be read next.

[0090] Step 70011: Wait for completion.

[0091] Step 70012: Refer to the initial compression information 2104. If there is no compression, jump to step 70014.

[0092] Step 70013: Recognize the compression algorithm applied to the read data with the initial compression information 2104, perform decompression processing, and return to the uncompressed state.

[0093] Step 70014: Refer to the applied compression process 2105 and compress the data using the applied compression algorithm.

[0094] Step 70015: Check whether the current address is the address at the beginning of the file. If it is not the beginning, jump to step 70017.

[0095] Step 70016: Make the physical page information 2203 in the free state indicated by the free page information 2500 as indicated by both the compression start pointer 2109 and the compression end pointer 2110 of the file information. Also, set the information indicated by the next page pointer 3002 of the allocated physical page information 2203 in the free page information 2500. Set the address for writing to the beginning of the allocated physical page.

[0096] Step 70017: Check whether the data can be stored only in the currently allocated physical page based on the length of the compressed data. If it can be stored, jump to step 70019.

[0097] Step 70018: Make the actual page information 2203 (the said actual page information 2203) indicated by the free page information 2500 in the free state be as indicated by the next page pointer 3002 of the actual page information 2203 indicated by the compression end pointer 2110. Also, make the said actual page information 2203 be as indicated by the compression end pointer 2110. In addition, set the information indicated by the next page pointer 3002 of the said actual page information 2203 (the assigned actual page information 2203) in the free page information 2500.

[0098] Step 70019: To write the compressed data to the recognized area for writing, issue a write request to the storage device 130.

[0099] Step 70020: Wait for completion.

[0100] Step 70021: Check whether all of the data of the file has been completed. If completed, jump to Step 70023.

[0101] Step 70022: Calculate the address for the next write from the length of the compressed data. After that, jump to Step 70010.

[0102] Step 70023: Return all of the actual page information 2203 pointed to by the reception start pointer 2107 to the free actual page information queue 900 indicated by the free page information 2500. After that, return to Step 70007.

[0103] According to this embodiment, in a file storage that performs compression in a batch later, by selecting a compression algorithm to be applied according to the amount of data that must be compressed, the data reduction rate can be improved. Also, for a file with a high access frequency, the response performance can be improved by caching the temporarily expanded data.

[0104] (Supplementary Note) The above-described embodiment includes, for example, the following content.

[0105] In the above-described embodiments, the case where the present invention is applied to a file storage has been described. However, the present invention is not limited thereto, and can be widely applied to various systems, apparatuses, methods, and programs.

[0106] Also, in the above-described embodiments, the case where the data in the cache area is managed in units of files has been described. However, the present invention is not limited thereto. For example, the data in the cache area may be managed in units of read requests.

[0107] Also, in the above-described embodiments, when receiving the first compression algorithm from an application, when receiving a file read request from the application, data of the file that is compressed by the second compression algorithm is read from a storage device, the read compressed data is decompressed by the second compression algorithm, and the decompressed data is compressed by the first compression algorithm and responded to the application. However, the present invention is not limited thereto. For example, when receiving the first compression algorithm from an application, when receiving a file read request from the application, data of the file that is compressed by the second compression algorithm is read from a storage device, the read compressed data is decompressed by the second compression algorithm, and the decompressed data may be compressed by a third compression algorithm different from the first compression algorithm and responded to the application.

[0108] Also, the configuration of the above-described embodiments may be, for example, the following configuration.

[0109] (1) A file storage (e.g., file storage 100, server 110) comprising a processor (e.g., processor 200) that receives a write request for a file from an application (e.g., user application 140), writes the data of the file to a storage device (e.g., storage device 130), and later compresses the data of the written file and writes it to the storage device (e.g., storage device 130). The processor may determine a compression algorithm to be used for compression according to the amount of data written (e.g., total compression amount 2800) to one or more written files within a predetermined time in step 70006. The file storage may be, for example, one in which a sensor selects a compression method according to the data generation rate. As the storage device, the data generation rate generated by the sensor corresponds to the amount of data written within a predetermined time.

[0110] For example, when the amount of written data does not exceed a threshold, the processor determines a compression algorithm with a first compression speed, and when the amount of written data exceeds the threshold, the processor determines a compression algorithm with a second compression speed greater than the first compression speed. Also, for example, the processor may determine a compression algorithm with a first compression speed during a time period when the amount of written data is small (e.g., at night), and determine a compression algorithm with a second compression speed greater than the first compression speed during a time period when the amount of written data is large (e.g., during the day).

[0111] Here, the compression algorithm is, for example, an application program (compression software). In this case, the processor may change the setting related to the compression speed (compression ratio) in the compression software and execute the compression software with the changed setting to compress the data, or may execute the determined compression software from a plurality of compression softwares with different compression speeds to compress the data.

[0112] According to the above configuration, since the data of the written file is compressed later, for example, the data reduction rate can be increased without degrading the response performance during data storage.

[0113] (2) (1) The file storage according to the above, wherein the processor may determine a compression algorithm to be used for compression according to the amount of data written in one or more files written in a predetermined time and the compression speed of each of a plurality of compression algorithms (for example, compression performance 2005) in step 70006.

[0114] For example, when 100 GB of data is written, the processor determines a compression algorithm capable of compressing 100 GB of data within a predetermined time (for example, a pre-specified time, the time from the end of the operation related to the user application 140 to the start of the operation, a periodic time such as every day).

[0115] According to the above configuration, for example, it is possible to determine the compression algorithm with the highest compression rate from among the compression algorithms with a compression speed higher than the data generation speed, so that it is possible to avoid a situation where uncompressed data accumulates.

[0116] (3) (1) The file storage according to the above, wherein the processor receives the media type of the data written to the file from the application in step 50002, and may determine a compression algorithm to be used for compression according to the amount of data written in one or more files written in a predetermined time and the received media type in step 70006.

[0117] The above-mentioned processor determines different compression algorithms for video data, still image data, and audio data, respectively. Also, if the video data, still image data, and audio data are uncompressed data, the total amount of write data is 4500 MB, and the time available for compression is 45 seconds, for example, the above-mentioned processor determines the compression algorithm with the highest compression ratio from among the compression algorithms that satisfy a compression speed of 100 MB / s for each of the video, still image, and audio. In this way, the compression algorithm may be determined at the average compression speed. However, the method for determining the compression algorithm is not limited to this.

[0118] According to the above configuration, for example, since a compression algorithm suitable for the media type can be determined, the data reduction rate can be further increased.

[0119] In addition, even if it is the same media type, if data that has not deteriorated is transmitted from the application, the processor may determine a compression algorithm that prioritizes quality (such as image quality, sound quality, etc.), and if data with a reduced size is transmitted from the application, the processor may determine a compression algorithm that does not prioritize quality.

[0120] (4) (3) The file storage described in (3), wherein the above-mentioned processor may determine the compression algorithm to be used for compression according to the amount of data written in a predetermined time of one or more files written in step 70006, the received media type, and the compression speed of each of the plurality of compression algorithms.

[0121] According to the above configuration, for example, for each media type, the compression algorithm with the highest compression ratio can be determined from among the compression algorithms with a compression speed greater than the data generation speed, so that the data reduction rate can be further increased and the situation where uncompressed data accumulates can be avoided.

[0122] (5) The file storage according to (1), wherein the processor receives from the application whether compression is to be performed on the data transmitted from the application and the compression algorithm when compression is being performed (see, for example, step 50002 in FIG. 13).

[0123] In the above configuration, for example, when the processor receives a first compression algorithm from the application, the processor can decompress the compressed data transmitted from the application using the first compression algorithm and then compress and store it using a second compression algorithm with a higher compression rate than the first compression algorithm. In the above configuration, for example, when there is a read request from the application, the processor can decompress the target data using the second compression algorithm, compress the decompressed data using the first compression algorithm, and respond to the application.

[0124] Also, for example, when the processor receives a first compression algorithm from the application, the processor can determine a second compression algorithm with similar properties to the first compression algorithm. For example, the processor can determine the second compression algorithm in consideration of whether the compression of the first compression algorithm is reversible compression or irreversible compression, so that the data received from the application can be compressed without impairing its properties.

[0125] (6) The file storage according to (5), wherein the processor may determine the compression algorithm to be used for compression according to the amount of data written to one or more files written in a predetermined time and the compression speed of each of the plurality of compression algorithms in step 70006.

[0126] According to the above configuration, for example, a situation where uncompressed data accumulates can be avoided. Further, according to the above configuration, for example, the compressed data transmitted from the application can be decompressed and compressed with a compression algorithm having a higher compression rate, so that the reduction rate of the compressed data transmitted from the application can be further increased.

[0127] (7) (5) The file storage according to the above, wherein when the data written in step 70006 is compressed data, the processor determines the compression algorithm to be used according to the time for decompressing the data (for example, total decompression time 2900), the amount of data written in a predetermined time of one or more files written, and the compression speed of each of a plurality of compression algorithms.

[0128] According to the above configuration, for example, the processor can determine the compression algorithm in consideration of the time for decompressing the compressed data transmitted from the application, so that a situation where compressed data with a low compression rate transmitted from the application accumulates can be avoided.

[0129] (8) (5) The file storage according to the above, wherein the processor receives the media type of the data to be written to the file from the application in step 50002, and determines the compression algorithm to be used according to the amount of data written in a predetermined time of one or more files written and the received media type in step 70006.

[0130] According to the above configuration, for example, a compression algorithm suitable for the media type can be determined, so that the reduction rate of the compressed data transmitted from the application can be further increased.

[0131] (9) The file storage described in (8), wherein the processor may determine a compression algorithm to be used for compression according to the amount of data written in one or more files written in a predetermined time, the received media type, and the compression speed of each of a plurality of compression algorithms in step 70006.

[0132] According to the above configuration, for example, it is possible to further increase the reduction rate of the compressed data transmitted from the application and avoid a situation where uncompressed data accumulates.

[0133] (10) The file storage described in (9), wherein when the written data is compressed data, the processor may determine a compression algorithm to be used for compression according to the time for decompressing the data, the amount of data written in one or more files written in a predetermined time, the received media type, and the compression speed of each of a plurality of compression algorithms in step 70006.

[0134] According to the above configuration, for example, it is possible to further increase the reduction rate of the compressed data transmitted from the application and avoid a situation where the compressed data with a low compression rate transmitted from the application accumulates.

[0135] (11) A file storage (e.g., file storage 100, server 110) comprising a processor (e.g., processor 200) that receives a write request for a file from an application (e.g., user application 140), writes the data of the file to a storage device (e.g., storage device 130), and later compresses the data of the written file and writes it to the storage device (e.g., storage device 130). When the processor receives a read request for a file storing compressed data from the application, in step 60006, it decompresses the compressed data, in step 60013, stores the decompressed data in a cache area, in step 60002, determines whether the data of the file for which the read request was received from the application exists in the cache area, and if it exists in the cache area, in steps 60017 and 60019, reads the data from the cache area and in step 60021, passes the read data to the application.

[0136] According to the above configuration, for example, it is possible to increase the data reduction rate without degrading the response performance during data storage and avoid a situation where the read performance of data of a file with a high read frequency deteriorates.

[0137] (12) A file storage (e.g., file storage 100, server 110) comprising a processor (e.g., processor 200) that receives a write request for a file from an application (e.g., user application 140), writes the data of the file to a storage device (e.g., storage device 130), and later compresses the data of the written file and writes it to the storage device (e.g., storage device 130). The processor, in step 50002, receives from the application the presence or absence of compression for the data transmitted from the application and the compression algorithm when compression is being performed. When receiving a read request for a file storing compressed data from the application, in step 60006, the processor decompresses the compressed data, and if it has received a compression algorithm from the application, compresses the decompressed data using the received compression algorithm. In step 60013, the compressed data is stored in a cache area. In step 60002, it is determined whether the data of the file for which a read request has been received from the application exists in the cache area. If it exists in the cache area, in steps 60017 and 60019, the data is read from the cache area, and in step 60021, the read data is passed to the application.

[0138] According to the above configuration, for example, it is possible to increase the data reduction rate without degrading the response performance during data storage and avoid a situation where the read performance of compressed data of frequently read files deteriorates.

[0139] Also, regarding the above-described configuration, within a range not exceeding the gist of the present invention, it may be appropriately changed, recombined, combined, or omitted.

Explanation of Reference Numerals

[0140] 100 File storage 110 Server 120 Network 130 Storage device 140 User application 200 Processor 210 Main memory 220 Shared memory 2000 File storage information 2100 File information 2200 Storage device information 2203 Actual page information 4000 Write processing unit 4100 Read processing unit 4200 Compression processing unit

Claims

1. A file storage comprising a processor that receives a write request for a file from an application, writes the data of the file to a storage device, and later compresses the data of the written file and writes it to the storage device, wherein the processor receives from the application the presence or absence of compression for the data transmitted from the application and the compression algorithm when compression is performed, and the processor determines a compression algorithm to be used for compression according to the amount of data written to one or more written files within a predetermined time. File storage.

2. The file storage according to claim 1, wherein the processor determines a compression algorithm to be used for compression according to the amount of data written to one or more written files within a predetermined time and the compression speed of each of a plurality of compression algorithms. File storage.

3. The file storage according to claim 1, wherein the processor receives the media type of the data to be written to the file from the application, and determines a compression algorithm to be used for compression according to the amount of data written to one or more written files within a predetermined time and the received media type. File storage.

4. The file storage according to claim 3, wherein the processor determines a compression algorithm to be used for compression according to the amount of data written to one or more written files within a predetermined time, the received media type, and the compression speed of each of a plurality of compression algorithms. File storage.

5. The file storage according to claim 1, wherein the processor determines a compression algorithm to be used for compression according to the amount of data written to one or more written files within a predetermined time and the compression speed of each of a plurality of compression algorithms. File storage.

6. The file storage according to claim 1, wherein when the written data is compressed data, the processor determines a compression algorithm to be used for compression according to the time required to decompress the data, the amount of data written to one or more written files within a predetermined time, and the compression speed of each of a plurality of compression algorithms. File storage.

7. The file storage according to claim 1, wherein the processor receives the media type of data to be written to a file from an application, and determines a compression algorithm to be used for compression according to the amount of data written in one or more files within a predetermined time and the received media type; File storage.

8. The file storage according to claim 7, wherein the processor determines a compression algorithm to be used for compression according to the amount of data written in one or more files within a predetermined time, the received media type, and the compression speed of each of a plurality of compression algorithms; File storage.

9. The file storage according to claim 8, wherein, when the data written is compressed data, the processor determines a compression algorithm to be used for compression according to the time required to decompress the data, the amount of data written in one or more files within a predetermined time, the received media type, and the compression speed of each of a plurality of compression algorithms; File storage.

10. The file storage according to claim 1, wherein the processor when receiving a read request for a file storing compressed data from an application, decompresses the compressed data, stores the decompressed data in a cache area, determines whether the data of the file for which the read request is received from the application exists in the cache area, and if it exists in the cache area, reads the data from the cache area and passes the read data to the application; File storage.

11. A file storage comprising a processor that receives a write request for a file from an application, writes the data of the file to a storage device, and later compresses and writes the data of the file that has been written to the storage device, wherein the processor receives from the application the presence or absence of compression for the data transmitted from the application and the compression algorithm when compression is performed; When receiving a read request for a file storing compressed data from an application, decompress the compressed data, and if receiving a compression algorithm from the application, compress the decompressed data using the received compression algorithm, store the compressed data in a cache area, determine whether the data of the file for which a read request is received from the application exists in the cache area, and if it exists in the cache area, read the data from the cache area and pass the read data to the application, File storage.

Citation Information

Patent Citations

  • Data file pushing method, apparatus and system

    CN105045873A

  • On-vehicle video recording device

    JP2012023556A

  • Imaging apparatus and control method and program of the same

    JP2013093816A

  • Method for Compressing Data by a Server and a Device

    JP2018503882A

  • Storage drive, compression system thereof, and data compression method thereof

    JP2019008792A