Method, apparatus, and computer program product for data writing
By determining unavailable storage segments in the storage device and generating continuous write requests, the problem of data rewriting is solved, the write performance and the reconstruction ability of backup metadata are improved, and the stability and data integrity of the storage system are achieved.
Patent Information
- Application Number
- CN202110013351.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-06
AI Technical Summary
When the prior art rewrites data in a storage device, the data is discontinuous and the unavailable storage segment cannot be processed effectively, affecting the write performance and backup metadata reconstruction.
By determining the unavailable storage segment in the storage segment, obtaining the reference compression header and generating a continuous write request, retaining the compressed header information to construct a continuous write request, supporting the reconstruction of backup metadata.
Improves the write performance and stability of the storage system, ensures the integrity of backup metadata, and avoids data loss in unavailable storage segments.
Smart Images

Figure CN114721584B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to methods, devices, and computer program products for data writing. Background Art
[0002] To save the storage space occupied by data in a storage device, before writing data into a storage area, it can be compressed at a certain compression ratio. When the storage area is initially written with data compressed at a certain compression ratio (e.g., input / output (I / O) instructions), the data is continuous within the storage area. Subsequent data can be rewritten into the same storage area to overwrite the original data.
[0003] Since a physical storage area can be divided into page-sized segments according to a paging management scheme, where the page size is the smallest allocation unit of 4KB or 8KB. If the subsequent data has a higher compression ratio than the original data, the subsequent data is usually non-continuous within the same storage area, and there is an interval (also called a "hole") between the subsequent data and the original data. In addition, some of the original data may be recycled and marked as unavailable, and become an interval affecting data writing. Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for data writing.
[0005] According to a first aspect of the present disclosure, a method for data writing is proposed. The method includes: determining unavailable storage segments among a plurality of storage segments of a storage area, each storage segment being used to store a compression header and compression data corresponding to the compression header; obtaining a reference compression header for the unavailable storage segment, the reference compression header including metadata indicating the segment size of the unavailable storage segment; and generating a continuous write request for the storage area based at least on target data to be written into the storage area and the reference compression header, so as to write the target data into available storage segments among the plurality of storage segments.
[0006] According to a second aspect of the present disclosure, an electronic device is provided. The device includes: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the device to perform actions including: determining unavailable storage segments among a plurality of storage segments of a storage area, each storage segment being for storing a compressed header and compressed data corresponding to the compressed header; obtaining a reference compressed header for the unavailable storage segment, the reference compressed header including metadata indicating the segment size of the unavailable storage segment; and generating a consecutive write request for the storage area at least based on target data to be written to the storage area and the reference compressed header to write the target data to an available storage segment among the plurality of storage segments.
[0007] In a third aspect of the present disclosure, a computer program product is provided. The computer program product is stored in a non-transitory computer storage medium and includes machine-executable instructions that, when running on a device, cause the device to perform any of the steps of the method described in the first aspect of the present disclosure.
[0008] The summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. The summary is not intended to identify key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0010] Figures 1A - 1D A schematic diagram illustrating conventional data rewriting is shown;
[0011] Figure 2 A schematic diagram illustrating the mapping between backup metadata and compressed data is shown;
[0012] Figure 3 A schematic diagram illustrating an exemplary environment in which embodiments of the present disclosure may be implemented is shown;
[0013] Figure 4 A flowchart illustrating the process of data writing according to an embodiment of the present disclosure is shown;
[0014] Figure 5 A schematic diagram illustrating a candidate compressed header according to an embodiment of the present disclosure is shown; and
[0015] Figure 6The figure illustrates a schematic block diagram of an example device that can be used to implement embodiments of the present disclosure. Detailed Description
[0016] Preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure will be more thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0017] As used herein, the term "comprising" and its variations mean open inclusion, i.e., "including but not limited to". Unless specifically stated otherwise, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included hereinafter.
[0018] Modern storage systems (e.g., all-flash array (AFA) storage devices) that employ real-time data compression (ILC) technology can significantly save disk space. As previously mentioned, the storage area is divided into multiple storage segments according to the page size. When rewriting data in a sector of the storage area where the original data was previously stored, no change in metadata will occur. If the rewritten data has the same compression ratio as the original data (i.e., the ratio of the amount of data before compression to the amount of data after compression), the storage space of the storage area can be fully utilized to sequentially write the rewritten data to the backend drive device in a continuous manner.
[0019] However, the inventors have found that if the rewritten data has a greater compression ratio than the original data, it will cause the originally continuous data to be stored in a discontinuous manner, with gaps left between the storage segments. In addition, some of the original data may be recycled, in which case the metadata will be modified to indicate that the corresponding storage segment is unavailable, and the stored original data will not be replaced. The following will specifically describe the traditional data rewriting process with reference to Figures 1A to 1D to specifically describe the traditional data rewriting process.
[0020] As Figure 1AAs shown, the storage area 110 includes four storage segments 120-1, 120-2, 120-3, and 120-4 (collectively or individually referred to as storage segment 120). Each storage segment 120 stores raw data of different sizes. For example, the size of storage segment 120-1 is 13 sectors, the size of storage segment 120-2 is 13 sectors, the size of storage segment 120-3 is 6 sectors, and the size of storage segment 120-4 is 16 sectors.
[0021] In Figure 1B , the data in storage segment 120-3 is recycled, and storage segment 120-3 is marked as unavailable. In this case, the user will not be able to access storage segment 120-3 or write data to storage segment 120-3. It should be understood that based on the ILC technology, the raw data stored in storage segment 120-3 is not deleted, but only the corresponding metadata is modified to indicate that storage segment 120-3 is unavailable.
[0022] In Figure 1C the target data 130 (also referred to as rewritten data 130) includes three items of rewritten data that need to be written to the storage area 110 to overwrite the raw data. Specifically, the first item of rewritten data 130-1 has a size of 7 sectors, the second item of rewritten data 130-2 has a size of 12 sectors, and the third item of rewritten data 130-3 has a size of 16 sectors.
[0023] In this case, the first item of rewritten data will be written to storage segment 120-1, resulting in a hole 140-1 with a size of 6 sectors; the second item of rewritten data will be written to storage segment 120-2, resulting in a hole 140-2 with a size of 1 sector.
[0024] As Figure 1C shown, in this case, three holes will be generated, namely hole 140-1, hole 140-2, and the unavailable storage segment 120-3.
[0025] According to the traditional scheme, a continuous write request can be constructed by writing padding data. For example, as Figure 1D shown, a write request can be constructed according to the storage granularity of the storage area (for example, each storage page is M sectors, and M is 8 for example).
[0026] For example, the first item of rewritten data 130-1 (with a size of 7 sectors) can be combined with padding data 150-1 with a size of 1 sector to form write data for the first 8 sectors. The data 150-3 of the first 3 sectors in the second item of rewritten data 130-2 can be combined with padding data 150-2 with a size of 5 sectors to form write data for the second 8 sectors.
[0027] However, traditional solutions cannot effectively handle the unavailable storage section 120-3. Some traditional solutions do not process the unavailable storage section 120-3 to avoid affecting the useful data in the unavailable storage section 120-3. However, this will result in the constructed write requests being discontinuous, affecting the write performance of the storage system.
[0028] In addition, some traditional solutions rewrite the unavailable storage section 120-3 by simply writing padding data (e.g., writing 0) to construct continuous write requests. However, since the unavailable storage section 120-3 includes compressed header information (Zip Header), such a compressed header can help reconstruct backup metadata (e.g., VBM files). If the compressed header in the unavailable storage section 120-3 is directly overwritten, this will cause the file system to be unable to reconstruct the backup metadata.
[0029] Figure 2 A schematic diagram 200 showing the mapping between backup metadata and compressed data is as follows Figure 2 As shown, the backup metadata 210 can maintain the metadata corresponding to different storage sections 120, and it constructs the mapping between the metadata and the data part 220 by representing the starting position of each section 120 through its length.
[0030] Exemplarily, corresponding to the example in Figure 1C the backup metadata 210 includes backup metadata 212, 214, 216, and 218 corresponding to four storage sections 120 respectively. Each item of backup metadata can maintain the corresponding length information to indicate the corresponding data part.
[0031] For example, the backup metadata 212 can correspond to the first rewritten data 130-1 (which includes a compressed header and the corresponding compressed data) and the hole 140-1; the backup metadata 214 corresponds to the second rewritten data 130-2 and the hole 140-2; the backup metadata 216 corresponds to the unavailable storage section 120-3; the backup metadata 218 corresponds to the third rewritten data 130-3.
[0032] It can be seen that if the compressed header included in the unavailable rewritten data 120-3 is overwritten, this will cause the file system to be unable to reconstruct the backup metadata based on the compressed header included in the data part 220 once the backup metadata 210 is damaged.
[0033] According to an embodiment of the present disclosure, a solution for data writing is provided. This solution enables the efficient construction of large continuous writes by rewriting a compression header corresponding to the length of an unavailable storage section. In addition, this solution preserves the useful information of the compression header, thereby enabling the reconstruction of backup metadata to be supported.
[0034] Embodiments of the present disclosure will be specifically described below with reference to the accompanying drawings. Figure 3 A schematic diagram of an example environment 300 for data writing according to an embodiment of the present disclosure is shown. As Figure 3 shown, the example environment 300 includes a host 310, a storage manager 320, and a storage device 330. It should be understood that the structure of the example environment 300 is described only for exemplary purposes, without implying any limitation on the scope of the present disclosure. For example, embodiments of the present disclosure can also be applied to environments different from the example environment 300.
[0035] The host 310 can be, for example, any physical computer, virtual machine, server, etc. that runs a user application. The host 310 can send I / O requests to the storage manager 320, such as for reading data from the storage device 330 and / or writing data to the storage device 330, etc. In response to receiving a read request from the host 310, the storage manager 320 can read data from the storage device 330 and return the read data to the host 310. In response to receiving a write request from the host 310, the storage manager 320 can write data to the storage device 330. The storage device 330 can be any currently known or future-developed non-volatile storage medium, such as a disk, a solid-state drive (SSD), or a disk array (RAID), etc.
[0036] The storage manager 320 can be deployed with a compression / decompression engine (not shown). For example, when the storage manager 320 receives a request from the host 310 to write data to the storage device 330, the storage manager 320 can use the compression / decompression engine to compress the data to be stored, and then store the compressed data in the storage device 330.
[0037] As described above, when the storage manager 320 performs data writing, the storage manager 320 can construct continuous write requests. The detailed process of data writing according to an embodiment of the present disclosure will be described below in conjunction with Figures 4 to 5 to describe.
[0038] Figure 4 A flowchart of an example process 400 for data writing according to an embodiment of the present disclosure is shown. For example, the process 400 can be performed by, such as Figure 3The storage manager 320 shown is used to execute. It should be understood that process 400 can also be executed by any other suitable device, and may include additional actions not shown and / or actions shown may be omitted. The scope of the present disclosure is not limited in this regard. For ease of description, process 400 will be described below with reference to FIGS. 1 to Figure 3 to describe process 400.
[0039] At block 402, the storage manager 320 determines an unavailable storage segment 120-3 among a plurality of storage segments 120 of the storage area 110, where each storage segment 120 is used to store a compressed header and compressed data corresponding to the compressed header.
[0040] As shown in FIG. 1, when receiving a rewrite request to write target data to the storage area 110, the storage manager 320 may indicate that the storage area 110 includes an unavailable storage segment 120-3.
[0041] In some implementations, the storage segment 120-3 may be marked as unavailable in response to being recycled. Alternatively, the storage segment 120-3 may also be marked as unavailable in response to receiving a rewrite request with a data size greater than the size of the storage segment.
[0042] In some implementations, the storage manager 320 may mark the storage segment 120-3 as unavailable by modifying the metadata corresponding to the storage segment 120-3 without deleting the compressed data stored in the storage segment 120-3. Such an unavailable storage segment 120-3 will not be accessible, thus forming the hole (or, gap) introduced above.
[0043] At block 404, the storage manager 320 obtains a reference compressed header for the unavailable storage segment 120-3, where the reference compressed header includes metadata indicating the segment size of the unavailable storage segment.
[0044] In some implementations, in order for the storage manager 320 to construct consecutive write requests without affecting the compressed headers included in the unavailable storage segment 120-3, the storage manager 320 may pre-construct a set of candidate compressed headers.
[0045] Figure 5 FIG. 500 illustrates a schematic diagram of candidate compressed headers according to an embodiment of the present disclosure. As Figure 5 shown, the storage manager 320 may allocate a buffer of a predetermined size in the memory for storing a set of candidate compressed headers 510-1, 510-2 to 510-N (collectively or individually referred to as candidate compressed headers 510).
[0046] As Figure 5As shown, each candidate compression header 510 corresponds to a different section size (ZLEN). For example, candidate compression header 510-1 corresponds to a section with a section size of 16 sectors and stores metadata indicating a section size of 16 sectors. Candidate compression header 510-2 corresponds to a section with a section size of 15 sectors and stores metadata indicating a section size of 15 sectors.
[0047] Since the file system only needs to utilize the section size information in the compression header when reconstructing backup metadata, other appropriate metadata can also be included in candidate compression header 510-1 to meet the needs of verification. It should be understood that other metadata can be initialized to any appropriate content, and the present disclosure is not intended to limit this.
[0048] After completing the construction of a set of candidate compression headers 510, the storage manager 320 can utilize this set of candidate compression headers 510 to determine the reference compression header corresponding to the unavailable storage section 120-3.
[0049] Specifically, the storage manager 320 can determine the section size of the unavailable storage section 120-3. For Figure 1C example, the storage manager 320 determines that the section size of the unavailable storage section 120-3 is 6 sectors based on metadata in the memory (e.g., the VBM file).
[0050] The storage manager 320 can determine the index for the reference compression header based on the section size. Exemplarily, depending on the organization order of a set of candidate compression headers 510, different section sizes can correspond to different indexes. For example, in Figure 5 , the candidate compression headers 510 are arranged in descending order of section size. Correspondingly, the index corresponding to the section size (6) can be determined to be 10, which means the 11th candidate compression header in this set of candidate compression headers 510.
[0051] Additionally, the storage manager 320 determines the reference compression header from a set of candidate compression headers 510 corresponding to different section sizes based on the index. For Figure 5 example, the storage manager 320 can determine the 11th candidate compression header as the reference compression header for the unavailable storage section 120-3.
[0052] Continuing to refer to Figure 2 , at block 406, the storage manager 320 generates a sequential write request for the storage area 110 based at least on the target data 130 to be written to the storage area and the reference compression header, so as to write the target data to the available storage sections among the multiple storage sections 120.
[0053] In some implementations, for the unavailable storage section 120-3, the storage manager 320 may utilize a reference compression header to generate section padding data for the unavailable storage section, where the section padding data includes the reference compression header and first padding data used to overwrite the previously stored compressed data in the unavailable storage section.
[0054] For Figure 1C example, the storage manager 320 may generate section padding data for the unavailable storage section 120-3, where the first sector may be filled with the reference compression header, and the remaining 5 sectors may be filled with a predetermined value (e.g., 0).
[0055] Additionally, the storage manager 320 may generate a sequential write request based at least on the target data 130 and the section padding data. Specifically, if only the data in a part of the available storage sections among the multiple storage sections 120 needs to be overwritten by the target data, the storage manager 320 may determine the remaining part of the available storage sections that does not need to be replaced by the target data.
[0056] For Figure 1C example, the storage section 120-1 includes a remaining part that is not covered, i.e., the hole 140-1. The storage section 120-2 includes a remaining part that is not covered, i.e., the hole 140-2.
[0057] The storage manager 320 may generate second padding data for the remaining part. Exemplarily, the storage manager 320 may generate the second padding data by writing a predetermined value (e.g., 0).
[0058] Additionally, the storage manager 320 may generate a sequential write request based on the target data, the section padding data, and the second padding data. As Figure 1C shown, according to the solution of the present disclosure, the storage manager 320 may determine that the first 7 sectors of the storage section 120-1 will be written with the first rewrite data 130-1, the last 6 sectors will be written with the padding data; the first 12 sectors of the storage section 120-2 will be written with the second rewrite data 130-2, the last 1 sector will be written with the padding data; the storage section 120-3 will be written with the padding data; the storage section 120-4 will be written with the third rewrite data 130-3. Based on such a manner, the storage manager 320 may generate a large sequential write request.
[0059] In some implementations, in order to align the written data with the storage granularity of the storage area 110, the storage manager 320 may further determine multiple data parts from the target data, the section padding data, and the second padding data according to the storage granularity associated with the storage area 110, and the size of each data part corresponds to the storage granularity.
[0060] Taking Figure 1D as an example, if the storage granularity of the storage area 110 (i.e., the size of each storage page) is 8 sectors, the storage manager 320 can further divide the continuous data determined above into multiple data parts each with a size of 8 sectors. For example, the first rewritten data 130-1 and the padding data 150-1 form the first data part; the padding data 150-2 and the data 150-3 of the first 3 sectors in the second rewritten data 130-2 form the second data part; the data 150-4 of 8 sectors in the second rewritten data 130-2 forms the third data part; the data 150-5 of the last 1 sector in the second rewritten data 130-2, the padding data 150-6 and the padding data 150-7 for the unavailable storage section 120-3 form the fourth data part; the third rewritten data 130-3 forms the fifth data part. Additionally, the storage manager can generate a continuous write request based on the multiple data parts. In this way, the write request can correspond to the storage page size of the storage area, thus ensuring data alignment.
[0061] In some implementations, in response to a request to reconstruct the backup metadata for the storage area 110, the storage manager 320 can reconstruct the backup metadata based on the section size of the unavailable storage section 120 = 3, where the backup metadata indicates the distribution of multiple storage sections. Specifically, when Figure 2 the backup metadata 210 as shown is damaged, the storage manager 320 can use the section size (e.g., 6 sectors) indicated in the reference compression header rewritten into the unavailable storage section 150-7 to reconstruct the VBM file, and such a VBM file can indicate the distribution of multiple storage sections 120.
[0062] Based on the methods discussed above, the embodiments of the present disclosure can construct large continuous write requests, thereby improving the write performance of the storage system. In addition, by effectively retaining the section length information of the unavailable storage section, the embodiments of the present disclosure can also support the reconstruction of backup metadata and improve the stability of the storage system.
[0063] Figure 6FIG. 0 shows a schematic block diagram of an example device 600 that can be used to implement embodiments of the present disclosure. For example, a storage manager 320 according to an embodiment of the present disclosure can be implemented by the device 600. As shown, the device 600 includes a central processing unit (CPU) 601, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or computer program instructions loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0064] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0065] The various processes and treatments described above, such as process 400, can be executed by the processing unit 601. For example, in some embodiments, process 400 can be implemented as a computer software program that is tangibly included in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the CPU 601, one or more actions of the process 400 described above can be executed.
[0066] The present disclosure can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0067] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0068] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0069] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0070] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0071] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0072] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0073] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.
[0074] The various embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A method for data writing, comprising: Determining unavailable storage segments among a plurality of storage segments of a storage area, each storage segment being for storing a compressed header and compressed data corresponding to the compressed header; Obtaining a reference compressed header for the unavailable storage segment, the reference compressed header including metadata indicating the segment size of the unavailable storage segment; And Generating a continuous write request for the storage area at least based on target data to be written to the storage area and the reference compressed header, to write the target data to available storage segments among the plurality of storage segments, Wherein obtaining the reference compressed header includes: Determining the segment size of the unavailable storage segment; Determining an index for the reference compressed header based on the segment size; and Determining the reference compressed header from a set of candidate compressed headers corresponding to different segment sizes based on the index.
2. The method according to claim 1, wherein generating a continuous write request for the storage area includes: Generating segment padding data for the unavailable storage segment by using the reference compressed header, the segment padding data including the reference compressed header and first padding data for covering previously stored compressed data in the unavailable storage segment; And Generating the continuous write request at least based on the target data and the segment padding data.
3. The method according to claim 2, wherein generating the continuous write request at least based on the target data and the segment padding data includes: If only data in a part of the available storage segments among the plurality of storage segments needs to be overwritten by the target data, determining the remaining part of the available storage segments that does not need to be replaced by the target data; Generating second padding data for the remaining part; And Generating the continuous write request based on the target data, the segment padding data, and the second padding data.
4. The method according to claim 3, wherein at least one of the first padding data and the second padding data is generated based on a predetermined value.
5. The method according to claim 3, wherein generating the continuous write request based on the target data, the segment padding data, and the second padding data includes: Determining a plurality of data parts from the target data, the segment padding data, and the second padding data according to a storage granularity associated with the storage area, each data part having a size corresponding to the storage granularity; And Generating the continuous write request based on the plurality of data parts.
6. The method according to claim 1, further comprising: In response to a request for reconstructing backup metadata for the storage area, reconstructing the backup metadata based on the segment size of the unavailable storage segment, the backup metadata indicating the distribution of the plurality of storage segments.
7. The method according to claim 1, wherein the unavailable storage segment is marked as unavailable in response to being recycled.
8. An electronic device, comprising: At least one processing unit; And At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which when executed by the at least one processing unit cause the device to perform actions, the actions including: Determining unavailable storage segments among a plurality of storage segments of a storage area, each storage segment for storing a compressed header and compressed data corresponding to the compressed header; Obtaining a reference compressed header for the unavailable storage segment, the reference compressed header including metadata indicating the segment size of the unavailable storage segment; and Generating a sequential write request for the storage area based at least on target data to be written to the storage area and the reference compressed header to write the target data to available storage segments among the plurality of storage segments, wherein obtaining the reference compressed header includes: Determining the segment size of the unavailable storage segment; Determining an index for the reference compressed header based on the segment size; and Determining the reference compressed header from a set of candidate compressed headers corresponding to different segment sizes based on the index.
9. The apparatus according to claim 8, wherein generating a sequential write request for the storage area includes: Generating segment padding data for the unavailable storage segment using the reference compressed header, the segment padding data including the reference compressed header and first padding data for overwriting compressed data previously stored in the unavailable storage segment; and Generating the sequential write request based at least on the target data and the segment padding data.
10. The apparatus according to claim 9, wherein generating the sequential write request based at least on the target data and the segment padding data includes: If only data in a part of the available storage segments among the plurality of storage segments needs to be overwritten by the target data, determining the remaining part of the available storage segments that does not need to be replaced by the target data; Generating second padding data for the remaining part; and Generating the sequential write request based on the target data, the segment padding data, and the second padding data.
11. The apparatus according to claim 10, wherein at least one of the first padding data and the second padding data is generated based on a predetermined value.
12. The apparatus according to claim 10, wherein generating the sequential write request based on the target data, the segment padding data, and the second padding data includes: Determining a plurality of data portions from the target data, the segment padding data, and the second padding data according to a storage granularity associated with the storage area, each data portion having a size corresponding to the storage granularity; and Generating the sequential write request based on the plurality of data portions.
13. The apparatus according to claim 8, the actions further including: In response to a request to reconstruct backup metadata for the storage area, reconstructing the backup metadata based on the segment size of the unavailable storage segment, the backup metadata indicating the distribution of the plurality of storage segments.
14. The apparatus according to claim 8, wherein the unavailable storage section is marked as unavailable in response to being reclaimed.
15. A computer program product tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions that, when executed by an apparatus, cause the apparatus to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Flexible shader derivation design in multi-computing kernel
CN108804219A
Storage zone set membership
US20170185312A1