Data writing method and apparatus
By computing and transmitting data blocks and temporary verification blocks on the client, and combining verification blocks by the server when the strip is full, the waste of memory and storage space of the client and server is solved, and the efficiency and realization of the storage system are improved.
Patent Information
- Application Number
- PCT/CN2024/095940
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2024-05-29
- Publication Date
- 2025-05-30
AI Technical Summary
In a storage system based on erasure coding technology, the client needs to cache a large number of temporary verification blocks, resulting in waste of memory space. At the same time, the computing complexity is high, and the server also needs to store a large number of verification blocks, resulting in waste of storage space.
The client calculates the verification block based on the written data. When the strip is full, the server merges the verification blocks to reduce the memory and storage space consumption of the client and the server.
It reduces the memory consumption of the client, reduces the number of temporary verification blocks stored on the server, and improves the utilization of storage space and the realization of the solution.
Smart Images

Figure CN2024095940_30052025_PF_FP_ABST
Abstract
Description
Data writing method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 22, 2023, with application number 202311572685.2, and invention name “A method, device and other equipment for data processing”, and the Chinese patent application filed with the State Intellectual Property Office on January 12, 2024, with application number 202410050635.6, and invention name “A method and device for data writing”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of cloud computing, and in particular to a data writing method and device. Background Art
[0003] With the rapid growth of data volumes and the increasing demand for data protection, storage systems are using erasure coding (EC) technology to improve reliability while maintaining high data reliability with low redundant storage overhead. In erasure coding, written data is split into multiple data blocks and additional parity blocks are created. These data blocks and parity blocks are combined to form a higher-redundancy set. If any data block is damaged or lost, the remaining data blocks and parity blocks can be used together to recover the lost data block.
[0004] In the current data writing process based on erasure coding technology, in order to improve storage utilization while reducing the storage space occupied by the check block in the incomplete stripe, the client writes the check block in the server using incremental erasure coding. That is, the client calculates a temporary check block for the data block first written to the stripe, and calculates a new temporary check block for the data block to be written to the stripe subsequently based on the temporary check block, until the data block fills the stripe.
[0005] Since in the incremental erasure code scheme, the data block newly written into the stripe needs to generate a temporary check block based on the last written data block to calculate the check block, the client needs to cache a large number of temporary check blocks, resulting in a waste of client memory space. At the same time, the calculation complexity of the check block based on the written data and the temporary check block is high, and the server also needs to store a large number of check blocks, resulting in a waste of storage space.
[0006] Summary of the Invention
[0007] Embodiments of the present application provide a data writing method in which a client calculates check blocks based on the written data. When a stripe is full, the server merges the check blocks, thereby reducing client memory consumption. Embodiments of the present application also provide a data writing apparatus, computing device, computing device cluster, computer-readable storage medium, and computer program product corresponding to the data writing method.
[0008] In a first aspect, an embodiment of the present application provides a data writing method, which can be executed by a distributed storage system, or by a component of the distributed storage system, such as a processor, chip, or chip system of the distributed storage system, or by a logic module or software that can implement all or part of the functions of the distributed storage system. In an embodiment of the present application, the distributed storage system includes a client and a server, the server is used to access a storage device, and the storage device is used to store data encoded based on erasure codes. The method provided in the first aspect includes: the client receives a first data block written by a user, and the first data block is used to write to a target stripe in the storage device. The client generates a first temporary check block based on the first data block. The client sends the first data block and the first temporary check block to the server. The server stores the first data block and the first temporary check block to the storage device. The server determines that the first data block is the tail data block of the target stripe, merges the first temporary check block and the second temporary check block by erasure coding, and obtains a merged target check block, the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block. The server stores the target check block to the storage device.
[0009] In the embodiment of the present application, the client in the distributed storage system calculates a temporary check block based on the written data, and when the stripe of the storage device is full, the server merges the temporary check blocks to obtain the target check block. Compared with the current client calculating the erasure code for the temporary check block and the newly written data, the client in the embodiment of the present application does not need to store the temporary check block, thereby reducing the client's memory consumption.
[0010] In one possible implementation, the erasure code encoding includes a coefficient multiplication operation and an exclusive-OR operation. In the process of erasure-coding the first temporary check block and the second temporary check block, the server performs an exclusive-OR operation on the first temporary check block and the second temporary check block to obtain a merged target check block.
[0011] In the embodiment of the present application, the server can perform an XOR operation based on the first temporary check block and the second temporary check block to obtain a merged target check block, thereby reducing the number of temporary check blocks stored by the server, reducing the server's consumption of storage space, and improving the feasibility of the solution.
[0012] In one possible implementation, during the process of generating the first temporary check block based on the first data block, the client determines that the size of the first data block is greater than or equal to a preset threshold, and then divides the first data block into at least one third data block having a size equal to the preset threshold. The client then performs erasure coding on the at least one third data block to generate a third temporary check block.
[0013] In the embodiment of the present application, the server can split the data block according to the size of the written data block. When the size of the written data block is greater than a preset threshold, the client can perform erasure coding on the split written data to generate a temporary check block, thereby reducing the number of temporary check blocks and saving the transmission bandwidth between the client and the server.
[0014] In one possible implementation, while the client is generating the first temporary check block based on the first data block, the client determines that the size of the first data block is less than a preset threshold, and performs a coefficient multiplication operation on the first data block to obtain a fourth temporary check block. Specifically, the client performs a coefficient multiplication operation on the first data block based on the check coefficient to obtain the fourth temporary check block.
[0015] In the embodiment of the present application, the server can use different encoding methods to generate temporary check blocks according to the size of the written data block, thereby improving the feasibility of the solution.
[0016] In one possible implementation, when the server determines that the first data block is the tail data block of the target stripe, the server determines that the first data block is the tail data block of the target stripe based on a position identifier of the first data block, where the position identifier is used to indicate the storage location of the first data block in the target stripe. The server determines that the stripe is full and generates a target check block based on the position identifier. That is, the position identifier can indicate the timing for the server to generate the target check block.
[0017] In the embodiment of the present application, the server can identify the tail data block of the target stripe based on the position identifier in the written data block, and determine the timing of merging the check block based on the position identifier in the written data block, thereby improving the storage space utilization rate of the storage device of the server.
[0018] In one possible implementation, the size of the location identifier is 2B, location identifier 00 indicates that the stripe is full, 01 indicates that the data block is the first data block of the data portion of the stripe, 10 indicates that the data block is in the middle of the data portion of the data block stripe, and 11 indicates that the data block is the tail data block of the data portion of the stripe.
[0019] In the embodiment of the present application, the server can indicate the storage position of the first data block in the target stripe based on the content of the two-byte position identifier, thereby improving the feasibility of the solution.
[0020] In one possible implementation, while the server is determining that the first data block is the tail data block of the target stripe, the client sends a merge indication message to the server. The server receives the merge indication message and determines that the first data block is the tail data block of the target stripe. In other words, the client can also control the server to proactively merge check blocks using the merge indication message.
[0021] In the embodiment of the present application, the server can also receive the merge indication information sent by the client, and determine that the written data block is the tail data block of the target stripe based on the merge indication information, thereby improving the feasibility of the solution.
[0022] In one possible implementation, the client can simultaneously generate a first temporary check block based on the first data block and send a second temporary check block to the server. In other words, the client can simultaneously generate the temporary check block and send the data block and other temporary check blocks to the server, i.e., perform encoding and sending in parallel.
[0023] In the embodiment of the present application, the client can send data blocks to the server while performing erasure coding, so that the encoding check block and the sending data block are performed in parallel, thereby improving the data writing efficiency of the client.
[0024] In second aspect, an embodiment of the present application provides a distributed storage system, which includes a client and a server, the server is used to access a storage device, the storage device is used to store data encoded based on erasure codes, the client is used to execute the method executed by the client in the above-mentioned first aspect or any possible implementation method of the first aspect, and the server is used to execute the method executed by the server in the above-mentioned first aspect or any possible implementation method of the first aspect.
[0025] In a third aspect, an embodiment of the present application provides a data writing device, which is applied to a distributed storage system. The distributed storage system includes a client and a server, the server is used to access a storage device, the storage device is used to store data encoded based on an erasure code, and the device includes a transceiver unit and a processing unit. The transceiver unit is used to receive a first data block written by a user, and the first data block is used to write to a target stripe in the storage device. The processing unit is used to generate a first temporary check block based on the first data block. The transceiver unit is also used to send the first data block and the first temporary check block to the server. The processing unit is also used to store the first data block and the first temporary check block to the storage device. The processing unit is also used to determine that the first data block is the tail data block of the target stripe, merge the first temporary check block and the second temporary check block using an erasure code, and obtain a merged target check block. The second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block. The processing unit is also used to store the target check block to the storage device.
[0026] In one possible implementation, the erasure code encoding includes a coefficient multiplication operation and an XOR operation, and the processing unit is specifically configured to perform an XOR operation on the first temporary check block and the second temporary check block to obtain a merged target check block.
[0027] In one possible implementation, the processing unit is specifically configured to determine that a size of the first data block is greater than or equal to a preset threshold, divide the first data block into at least one third data block having a size equal to the preset threshold, and perform erasure coding on the at least one third data block to generate a third temporary check block.
[0028] In a possible implementation, the processing unit is specifically configured to determine that the size of the first data block is smaller than a preset threshold, and perform a coefficient multiplication operation on the first data block to obtain a fourth temporary check block.
[0029] In a possible implementation, the processing unit is specifically configured to determine that the first data block is a tail data block of the target stripe based on a position identifier of the first data block, where the position identifier is used to indicate a storage position of the first data block in the target stripe.
[0030] In a possible implementation, the transceiver unit is specifically configured to send merge indication information to the server, receive the merge indication information, and determine that the first data block is the tail data block of the target stripe.
[0031] In a fourth aspect, an embodiment of the present application provides a computing device, comprising a processor coupled to a memory, the processor being used to store instructions. When the instructions are executed by the processor, the computing device executes the method described in the first aspect or any possible implementation of the first aspect.
[0032] In a fifth aspect, an embodiment of the present application provides a computing device cluster, which includes one or more computing devices, each of which includes a processor coupled to a memory, and the processor is used to store instructions. When the instructions are executed by the processor, the computing device cluster executes the method described in the first aspect or any possible implementation method of the first aspect.
[0033] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed, the computer executes the method described in the first aspect or any possible implementation method of the first aspect.
[0034] In a seventh aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed, the computer implements the method described in the first aspect or any possible implementation method of the first aspect.
[0035] It can be understood that the beneficial effects that can be achieved by any of the distributed storage systems, data writing devices, computing devices, computing device clusters, computer-readable media or computer program products provided above can be referred to the beneficial effects in the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a schematic diagram of the system architecture of a distributed storage system provided in an embodiment of the present application;
[0037] FIG2 is a schematic flow chart of a data writing method provided in an embodiment of the present application;
[0038] FIG3 is a schematic diagram of a data portion and a check portion of a stripe provided in an embodiment of the present application;
[0039] FIG4 is a schematic diagram of a stripe for writing data blocks according to an embodiment of the present application;
[0040] FIG5 is a schematic diagram of another stripe for writing data blocks provided by an embodiment of the present application;
[0041] FIG6 is a schematic diagram of another stripe for writing data blocks provided by an embodiment of the present application;
[0042] FIG7 is a schematic diagram of parallel encoding and writing provided by an embodiment of the present application;
[0043] FIG8 is a schematic diagram of a merged check block provided in an embodiment of the present application;
[0044] FIG9 is a schematic diagram of a data writing device provided in an embodiment of the present application;
[0045] FIG10 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0046] FIG11 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0047] FIG12 is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] Embodiments of the present application provide a data writing method and apparatus for reducing bandwidth consumption of a cloud-side server during data writing.
[0049] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0050] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0051] First, some terms involved in the embodiments of the present application are introduced to facilitate those skilled in the art to understand the technical solutions.
[0052] A stripe is a unit of data distribution. Data distribution evenly distributes contiguous data blocks across multiple disk drives, allowing the system to read and write data from multiple disks simultaneously, thereby parallelizing I / O operations. Data within a stripe is physically divided; for example, part of a file might be stored on one disk, while another part is on another. This data distribution allows data to be read or written from multiple disks simultaneously, significantly increasing data throughput.
[0053] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.
[0054] Please refer to Figure 1, which is a schematic diagram of the system architecture of a data writing system provided in an embodiment of the present application. In the system architecture shown in Figure 1, the distributed storage system 10 includes a client 100, a server 200, and a storage device 300. The storage device 300 includes one or more storage devices distributed in different data centers. The client 100 generates a check block for the written data based on an erasure code algorithm, and writes the data and check block to the server 200. The server 200 stores the data blocks and check blocks written by the client 100. The specific functions of the client 100 and the server 200 are respectively introduced below.
[0055] Client 100 is a device or application that initiates a storage request to server 200. Specifically, client 100 can divide the written data into blocks and encode the data blocks using an erasure coding algorithm to generate check blocks. Client 100 uploads both the data blocks and the check blocks to server 200, which then stores these blocks and the check blocks on multiple storage devices in data centers located in different regions, thereby increasing the data's resistance to failures.
[0056] The client 100 is also used to download data and recover faulty data blocks through the server 200. For example, if a data block is damaged or lost during transmission, the client 100 can reconstruct the complete data using the remaining data blocks and check blocks.
[0057] It is understood that client 100 can be different devices or applications in different application scenarios. For example, in a personal storage scenario, client 100 can be a terminal device such as a personal computer, smartphone, or tablet computer, or a software client on the terminal device, such as an application. Client 100 uploads data to a cloud storage server via the network.
[0058] In enterprise applications, client 100 can be a file server, database server, or application server. It encodes internal data and then sends it to a data center or private cloud storage. In IoT applications, client 100 can be a sensor or device. It generates data and sends it to a backend server for analysis and backup.
[0059] Server 200 is a device that maintains storage resources and processes storage requests from client 100. Specifically, server 200 is configured to receive data blocks and check blocks sent by client 100 and store the data blocks and check blocks in storage device 300, which may include one or more storage devices located in data centers in different regions.
[0060] The server 200 is further configured to perform erasure coding on the parity block generated by the client 100 to generate a merged parity block, and store the merged parity block in the storage device 300. The server 200 is further configured to allow data blocks to be recovered from single or multiple disk block or storage node failures.
[0061] It is understandable that the distributed storage device end 200 can be different devices in different application scenarios. In some scenarios, the distributed storage device end 300 can also be a distributed disk array in a single storage server, which is not specifically limited.
[0062] The distributed storage end 300 can also be used as a storage device in other application scenarios. For example, the distributed storage end 300 can also be a storage device in a data center application scenario, a distributed file system application scenario, a disaster recovery backup scenario, etc.
[0063] The client 100 and the server 200 in the embodiment of the present application can be deployed in different devices respectively, or can be deployed in the same device, without specific limitation.
[0064] Based on the distributed storage system 10 shown in Figure 1 , the present application further provides a data writing method. The data writing method provided by the embodiment of the present application is described below in conjunction with an embodiment.
[0065] Please refer to Figure 2, which is a flow chart of a data writing method provided in an embodiment of the present application. In the example shown in Figure 2, the method includes the following steps:
[0066] Step 201: The client receives a first data block written by a user, and generates a first temporary check block based on the first data block.
[0067] Client 100 receives a first data block, which is data to be stored in a stripe of storage device 300. Client 100 generates a first temporary check block based on the first data block. A stripe is a unit of data distribution in server 200. A stripe is physically divisible. For example, a stripe may include storage space of different disk blocks.
[0068] Specifically, in the process of the client 100 generating a check block based on the first data block, the client 100 needs to divide the write data into blocks according to the size of the first data block, and generate a first temporary check block based on the divided data block, and then store the data block and the first temporary check block to the server 200. The first data block and the first temporary check block can be stored on different storage devices respectively.
[0069] Please refer to Figure 3, which is a schematic diagram of a stripe of a data block stored in a storage device provided by an embodiment of the present application. In the example shown in Figure 3, the client 100 needs to write the write data to the stripe of the storage device 300. It can be seen from the example shown in Figure 3 that the storage device 300 includes 6 disk blocks, and the 6 disk blocks can be distributed in different storage devices. The size of each disk block is, for example, 1M. Among them, among the 6 disk blocks, 4 disks are used to store data blocks, namely D0, D1, D2 and D3, and 2 disk blocks are used to store check blocks, namely P0 and P1.
[0070] In the example shown in Figure 3, the storage space of each disk block can be divided into different units, with each unit having a granularity of, for example, 8KB. Units from different disk blocks form a stripe. For example, the units in each row of six disk blocks in Figure 3 form a stripe, meaning the size of a stripe is 48KB. In a stripe, the portion used to store data blocks is called the data portion of the stripe, and the portion used to store parity blocks is called the parity portion of the stripe. For example, the first four 8KB units in a stripe are the data portion of the stripe, and the last two 8KB units in the stripe are the parity portion of the stripe.
[0071] In the example shown in FIG3 , server 200 stores received data blocks in the data portion of a stripe, and server 200 stores received check blocks in the check portion of the stripe. It should be noted that if server 200 receives more data blocks than can fit in the data portion of a stripe, server 200 can sequentially store the data blocks in the remaining data portion of the stripe until the data portion of the stripe is full, thereby improving the storage space utilization of the data portion of the stripe. When the data portion of a stripe is full, server 200 stores the data blocks in the data portion of the next stripe.
[0072] In the embodiment of the present application, the client 100 can generate a first temporary check block in different ways based on the size of the first data block to be written. The first temporary check block includes a third temporary check block and a fourth temporary check block. The following is a detailed introduction:
[0073] In one possible implementation, when the first data block being written is larger than a first threshold, where the first threshold is, for example, the size of a unit, i.e., 8K. In this embodiment of the present application, written data larger than the first threshold is also referred to as large input / output (IO) data. The client 100 divides the first data block being written into at least one third data block having a size equal to a preset threshold based on the first threshold, and performs erasure coding on the at least one third data block to generate a third temporary check block.
[0074] Please refer to Figure 4, which is a schematic diagram of a method for segmenting written data and generating check blocks according to an embodiment of the present application. In the example shown in Figure 4, when the written data is greater than a first threshold, client 100 segments the written data into different data blocks based on the first threshold, and performs erasure coding on the different data blocks to obtain a third temporary check block.
[0075] For example, in the example shown in FIG4 , when the write data is 16K data, since 16K data is greater than the first threshold value 8K, the client 100 divides the 16K write data into blocks based on 8K to obtain two 8K data blocks. The client 100 performs two erasure code calculations on the two 8K data blocks to obtain two 8K check blocks. Specifically, in the process of the client 100 performing two erasure code calculations on the two 8K data blocks, in the first calculation, the two 8K data blocks are multiplied by the check coefficients C00 and C01 respectively, and then an exclusive OR (XOR) operation is performed to obtain the first check block. Then, in the second calculation, the two 8K data blocks are multiplied by the check coefficients C10 and C20 respectively, and then an exclusive OR (XOR) operation is performed to obtain the second check block.
[0076] In one possible implementation, when the first data block being written is less than or equal to a first threshold, where the first threshold is, for example, the size of a unit, i.e., 8K. In this embodiment of the present application, written data less than or equal to the first threshold is also referred to as small input / output (IO) data. The client 100 performs a coefficient multiplication operation on the first data block based on the check coefficient to obtain a fourth temporary check block.
[0077] Please refer to Figure 5, which is a schematic diagram of a method for segmenting written data and generating check blocks according to an embodiment of the present application. In the example shown in Figure 5, when the written data is less than or equal to a first threshold, the client 100 generates a first temporary check block based on the written data and then performs a coefficient multiplication operation on the written data based on the check coefficient to obtain a fourth temporary check block.
[0078] For example, in the example shown in FIG5 , when the write data is 4KB, since the 4KB data is less than the first threshold of 8KB, the client 100 performs two erasure coding operations on the 4KB write data based on different check coefficients to obtain two 4KB check blocks. Specifically, during the process of creating two replicas of the 4KB data block based on different check coefficients, the client 100 multiplies the 4KB data block by two different check coefficients C00 and C01, respectively, to obtain two check blocks.
[0079] Step 202: The client sends a first data block and a first temporary check block to the server.
[0080] After generating the first temporary check block based on the first data block, the client 100 sends the first data and the first temporary check block to the server 200 .
[0081] Step 203: The server stores the first data block and the first temporary check block in a storage device.
[0082] After receiving the first data block and the first temporary check block sent by the client 100, the server 200 stores the first data block and the first temporary check block in the storage device. Specifically, the server 200 stores the received data block in a disk block of the storage device used for storing data blocks, i.e., the data portion of the stripe in the storage device 300, and stores the first temporary check block in a disk block of the storage device used for storing check blocks, i.e., the check portion of the stripe in the storage device 300.
[0083] Please continue to refer to Figure 4. In the example shown in Figure 4, the client 100 sends two 8K data blocks and two 8K check blocks to the server 200, wherein the server 200 stores the two 8K data blocks in the units of disk block D0 and disk block D1, that is, the first two units of the data part of the stripe, and stores the two 8K check blocks in the units of disk block P0 and disk block P1, respectively.
[0084] Please continue to refer to Figure 5. In the example shown in Figure 5, the client 100 sends a 4K data block and two 4K check blocks to the server 200, wherein the server 200 stores the 4K data block in the unit of disk block D0, that is, the first unit of the data part of the stripe, and stores the two 4K check blocks in the units of disk block P0 and disk block P1 respectively.
[0085] It should be noted that in the append write scenario, the server 200 sequentially writes the data blocks corresponding to the received multiple write data to the data portion of the stripe in the server 200. When the data portion of the stripe is full, the server 200 continues to write data blocks to the data portion of the new stripe. At the same time, the server 200 also sequentially writes the parity blocks corresponding to the multiple write data to the disk blocks used to store the parity blocks.
[0086] Please refer to Figure 6, which is a flow chart of another data writing method provided by an embodiment of the present application. In the example shown in Figure 6, the client 100 needs to write 16K of write data in the server 200. The client 100 first divides the 16K write data into blocks to obtain two 8K data blocks, and performs two erasure code calculations on the two 8K data blocks to obtain two 8K check blocks. The client 100 sends two 8K data blocks and two 8K check blocks to the server 200. The server 200 stores the two 8K data blocks in the units of disk block D0 and disk block D1, respectively, and stores the two 8K check blocks in the units of disk block P0 and disk block P1, respectively.
[0087] Afterwards, client 100 needs to continue writing 4KB of write data to server 200. Client 100 obtains two 4KB parity blocks based on the 4KB of write data and sends the 4KB data block and the two 4KB parity blocks to server 200. Server 200 continues to store the 4KB data block in the unit of disk block D2 and continues to store the two 4KB parity blocks in the units of disk block P0 and disk block P1, respectively.
[0088] In the example shown in FIG6 , when client 100 needs to continue writing data to server 200, server 200 sequentially stores the received data blocks in the remaining storage space of the data portion of the stripe until the data portion of the stripe is full. For example, server 200 continues to store the 4K data block received for the third time in the remaining 4K storage space of the unit of disk block D2. Server 200 continues to store the 4K data blocks received for the fourth and fifth times in the storage space of the unit of disk block D3.
[0089] In an embodiment of the present application, when the client 100 writes data in the data portion of the stripe, the client 100 maintains a data integrity field (DIF) for each data block. For example, the client 100 maintains a data integrity field for a 4k data block, and the size of the data integrity field is 64B.
[0090] Among them, the data integrity field includes a position identifier, the size of which is 2B. The position identifier is used to record the position of the data block in the data part of the stripe. For example, 00 indicates that the stripe is full (full stripe), 01 indicates that the data block is the first data block of the data part of the stripe (start), 10 indicates that the data block is in the middle of the data part of the stripe (on the way), and 11 indicates that the data block is the last data block of the data part of the stripe (final).
[0091] It is understandable that the check block corresponding to the data block can also maintain the data integrity field, and the position identifier corresponding to the check block indicates the position of the data block corresponding to the check block in the data part of the stripe.
[0092] Please refer to Figure 6. In the example shown in Figure 6, each 4KB of data stored in server 200 maintains a 64-byte data integrity field. The 2-byte portion of the data integrity field is a position identifier, which records the location of the data block in the data portion of the stripe. For example, the first 4KB of data in stripe disk block D0 corresponds to position identifier 01, the first 4KB of data in stripe disk block D2 corresponds to position identifier 10, and the second 4KB of data in stripe disk block D3 corresponds to position identifier 11.
[0093] In one possible implementation, the client 100 can simultaneously generate a first temporary check block based on the first data block and send a second temporary check block to the server 200. That is, the client 100 can simultaneously generate the check block and send the data block and other check blocks to the server 200, i.e., perform encoding and sending in parallel.
[0094] Continuing with Figure 6, in the example shown in Figure 6, client 100 needs to write 16KB of data to storage device 300. Client 100 first divides the 16KB of data into two 8KB data blocks, and then performs two erasure code calculations on the two 8KB data blocks to obtain two 8KB parity blocks. While client 100 is performing erasure code calculations on the two 8KB data blocks, client 100 can send the two 8KB data blocks to server 200.
[0095] In the example shown in Figure 6, when the client 100 needs to continue writing 4K of write data on the server 200, the client 100 generates two 4K check blocks based on the 4K write data. While generating the 4K check block, the client 100 can also send the 8K check block generated by the last calculation to the server 200.
[0096] Please refer to Figure 7, which is a schematic diagram of a client generating a check block and sending a data block in parallel according to an embodiment of the present application. In the example shown in Figure 7, the client 100 can send a data block to the server 200 while generating a check block based on the erasure code calculation of the data block. For example, while the client 100 is performing an erasure code calculation based on data block 1 and data block 2 to obtain check block 1, the client 100 can concurrently write data block 1 to disk block D0 of the server 200 and write data block 1 to disk block D1 of the server 200.
[0097] In the example shown in Figure 7, after the client 100 performs erasure code calculation based on data block 1 and data block 2 to obtain check block 1, the check block 1 is written to the disk block P0 of the server 200. At this time, the client 100 can perform erasure code calculation based on data block 1 and data block 2 in parallel to obtain a new check block.
[0098] Step 204. The server determines that the first data block is the tail data block of the target stripe, and performs erasure coding on the first temporary check block and the second temporary check block to obtain a merged target check block. The second temporary check block includes one or more temporary check blocks generated based on other data blocks in the target stripe except the first data block. After the server 200 receives the first temporary check block sent by the client 100, when the first temporary check block generates a check block for the tail data block of the stripe, the server 200 performs erasure coding on the first temporary check block and the second temporary check block to obtain a target check block. The target check block may also be referred to as a merged check block in this application. The second temporary check block includes one or more check blocks generated based on other data blocks in the stripe except the tail data block.
[0099] That is, when the data portion of a stripe in server 200 is fully written, server 200 merges the parity blocks corresponding to the data blocks in the stripe to obtain a target parity block. Specifically, server 200 identifies the timing for merging parity blocks based on the position identifier of the temporary parity block. When the position identifier indicates that the parity block is the parity block corresponding to the last data block in the data portion of the stripe, server 200 merges all parity blocks corresponding to the data blocks in the data portion of the stripe to obtain the target parity block.
[0100] Please continue to refer to Figure 6. In the example shown in Figure 6, the client 100 writes data blocks in the data part of the stripe of the storage device 300. For example, 16k of data is written for the first time and stored in disk blocks D0 and D1 of the data part of the stripe respectively. 4k of data is written for the second time and stored in disk block D2 of the data part of the stripe. 4k of data is written for the third time and also stored in disk block D2 of the data part of the stripe. 4k of data is written for the fourth time and stored in disk block D3 of the data part of the stripe. 4k of data is written for the fifth time and stored in disk block D3 of the data part of the stripe. The stripe is full after the fifth write.
[0101] Accordingly, in the example shown in FIG6 , the two 8K parity blocks corresponding to the first 16K data block written by client 100 are stored in disk blocks P0 and P1. Since the 16K data is the header data block of the stripe, the position of the parity block is 01. The two 4K parity blocks corresponding to the second 4K data block written by client 100 are also stored in disk blocks P0 and P1. The two 4K parity blocks corresponding to the third 4K data block written by client 100 are also stored in disk blocks P0 and P1. The two 4K parity blocks corresponding to the fourth 4K data block written by client 100 are also stored in disk blocks P0 and P1. Since the data written from the second to the fourth time are intermediate data blocks of the stripe, the positions of the three parity blocks from the second to the fourth time are 10. The two 4K parity blocks corresponding to the fifth 4K data block written by client 100 are stored in disk blocks P0 and P1. Since the 16K data is the header data block of the stripe, the position of the parity block is 11.
[0102] In the example shown in FIG6 , when the server 200 identifies the parity block position identifier as 11, the server 200 performs parity block merging. The server 200 merges all parity blocks corresponding to the data blocks in the stripe data portion. That is, the server 200 merges the 8K parity block written first, the 4K parity block written second, the 4K parity block written third, the 4K parity block written fourth, and the 4K parity block written fifth to obtain a merged 8K parity block. After obtaining the merged parity block, the server 200 rewrites the merged parity block into the disk blocks in the stripe parity portion.
[0103] In one possible implementation, in addition to merging temporary parity blocks after the data portion of a stripe is fully written, the server 200 may also generate a target parity block by controlling the server 200 to proactively merge parity blocks. The specific implementation is not limited thereto. For example, the client 100 may send a merge indication message to the server 200. The server 200 receives the merge indication message and, based on the merge indication message, determines that the first data block is the tail data block of the target stripe.
[0104] Please refer to Figure 8, which is a schematic diagram of a merging check block provided in an embodiment of the present application. In the example shown in Figure 8, the server 200 merges five check blocks, where the position of the first 4k check block is identified as 01, i.e., the check block corresponding to the head data block of the stripe, and the position of the fifth 4k check block is identified as 11, i.e., the check block corresponding to the tail data block of the stripe.
[0105] In the example shown in FIG8 , when the server 200 merges five check blocks, the server 200 performs the merge check at a granularity of 8K units, that is, the server 200 combines two 4K check blocks into one 8K check block, performs an XOR operation on the three 8K check blocks, and obtains a merged 8K check block. The position of the merged 8K check block is identified as 00, which indicates the check block corresponding to the full stripe.
[0106] Step 205: The server 200 stores the target check block in a storage device.
[0107] The server 200 performs erasure coding on the first temporary check block and the second temporary check block to merge the first temporary check block and obtain the merged target check block. Then, the server 200 stores the target check block in the storage device 300 .
[0108] It can be seen from the above embodiments that in the distributed storage system in the embodiments of the present application, the client calculates a temporary check block based on the written data, and when the stripe of the storage device is full, the server merges the temporary check blocks to obtain the target check block, so that the client does not need to store the temporary check blocks, thereby reducing the client's memory consumption during the data writing process.
[0109] Based on the above method embodiment, the embodiment of the present application further provides a data writing device. The data writing device provided by the embodiment of the present application is described in detail below.
[0110] Please refer to Figure 9, which is a schematic diagram of the structure of a data writing device provided in an embodiment of the present application. In the example shown in Figure 9, the data writing device 900 is used to implement the various steps performed by the distributed storage system in the above embodiments. The data writing device 900 includes a transceiver unit 901 and a processing unit 902.
[0111] Among them, the transceiver unit 901 is used to receive the first data block written by the user, and the first data block is used to write the target stripe in the storage device. The processing unit 902 is used to generate a first temporary check block based on the first data block. The transceiver unit 901 is also used to send the first data block and the first temporary check block to the server. The processing unit 902 is also used to store the first data block and the first temporary check block to the storage device. The processing unit 902 is also used to determine that the first data block is the tail data block of the target stripe, and merge the first temporary check block and the second temporary check block with error correction code to obtain a merged target check block, and the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block. The processing unit 902 is also used to store the target check block to the storage device.
[0112] In a possible implementation, the erasure code encoding includes a coefficient multiplication operation and an XOR operation, and the processing unit 902 is specifically configured to perform an XOR operation on the first temporary check block and the second temporary check block to obtain a merged target check block.
[0113] In one possible implementation, the processing unit 902 is specifically configured to determine whether the size of the first data block is greater than or equal to a preset threshold, divide the first data block into at least one third data block having a size equal to the preset threshold, and perform erasure coding on the at least one third data block to generate a third temporary check block.
[0114] In a possible implementation, the processing unit 902 is specifically configured to determine that the size of the first data block is smaller than a preset threshold, and perform a coefficient multiplication operation on the first data block to obtain a fourth temporary check block.
[0115] In a possible implementation, the processing unit 902 is specifically configured to determine that the first data block is the tail data block of the target stripe based on a position identifier of the first data block, where the position identifier is used to indicate a storage position of the first data block in the target stripe.
[0116] In a possible implementation, the transceiver unit 901 is specifically configured to send merge indication information to the server, receive the merge indication information, and determine that the first data block is the tail data block of the target stripe.
[0117] It should be understood that the division of units in the above device is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. Moreover, the units in the device can all be implemented in the form of software called through processing elements; or they can all be implemented in the form of hardware; or some units can be implemented in the form of software called through processing elements, and some units can be implemented in the form of hardware. For example, each unit can be a separately established processing element, or it can be integrated into a certain chip of the device. In addition, it can also be stored in the memory in the form of a program, called by a certain processing element of the device and perform the function of the unit. In addition, all or part of these units can be integrated together, or they can be implemented independently. The processing element described here can also be a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each unit above can be implemented by the integrated logic circuit of the hardware in the processor element or in the form of software called through the processing element.
[0118] It is worth noting that, for the sake of simplicity of description, the above method embodiments are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited to the order of the actions described. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required for this application.
[0119] Other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0120] Please refer to Figure 10, which is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in Figure 10, the computing device 1000 includes: a processor 1001, a memory 1002, a communication interface 1003, and a bus 1004. The processor 1001, the memory 1002, and the communication interface 1003 are coupled via a bus (not labeled in the figure). The memory 1002 stores instructions. When the execution instructions in the memory 1002 are executed, the computing device 1000 performs the method performed by the computing device in the above method embodiment.
[0121] The computing device 1000 may be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. For example, when a unit in the device can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For example, these units can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0122] The processor 1001 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0123] The memory 1002 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0124] The memory 1002 stores executable program codes, and the processor 1001 executes the executable program codes to respectively implement the functions of the aforementioned units or modules, thereby implementing the aforementioned data writing method. That is, the memory 1002 stores instructions for executing the aforementioned data writing method.
[0125] The communication interface 1003 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0126] In addition to the data bus, bus 1004 may also include a power bus, a control bus, and a status signal bus. The bus may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), or a Cache Coherent Interconnect for Accelerators (CCIX). Buses can be categorized as address buses, data buses, and control buses.
[0127] Please refer to Figure 11, which is a schematic diagram of a computing device cluster provided in an embodiment of the present application. As shown in Figure 11, the computing device cluster 1100 includes at least one computing device 1000.
[0128] As shown in Figure 11, the computing device cluster 1100 includes at least one computing device 1000. The memory 1002 in one or more computing devices 1000 in the computing device cluster 1100 may store the same instructions for executing the above-mentioned data writing method.
[0129] In some possible implementations, the memory 1002 of one or more computing devices 1000 in the computing device cluster 1100 may also store some instructions for executing the above-described data writing method. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the above-described data writing method.
[0130] It should be noted that the memory 1002 in different computing devices 1000 in the computing device cluster 1100 can store different instructions, each for executing part of the functions of the above-mentioned data writing device. In other words, the instructions stored in the memory 1002 in different computing devices 1000 can implement the functions of one or more modules in the transceiver unit and the processing unit.
[0131] In some possible implementations, one or more computing devices 1000 in the computing device cluster 1100 may be connected via a network, which may be a wide area network or a local area network.
[0132] Please refer to Figure 12, which is a schematic diagram of computing devices in a computing cluster provided by an embodiment of the present application connected via a network. As shown in Figure 12, two computing devices 1000A and 1000B are connected via a network. Specifically, the connection to the network is through a communication interface in each computing device.
[0133] In one possible implementation, the memory of the computing device 1000A stores instructions for executing the functions of the transceiver unit, while the memory of the computing device 1000B stores instructions for executing the functions of the processing unit and the display unit.
[0134] It should be understood that the functions of the computing device 1000A shown in Figure 12 may also be completed by multiple computing devices. Similarly, the functions of the computing device 1000B may also be completed by multiple computing devices.
[0135] In another embodiment of the present application, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the computing device in the above method embodiment.
[0136] In another embodiment of the present application, a computer program product is provided, the computer program product including computer-executable instructions stored in a computer-readable storage medium. When a processor of a device executes the computer-executable instructions, the device performs the method performed by the computing device in the above method embodiment.
[0137] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0138] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0139] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0140] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
Claims
1. A data writing method, characterized in that: The method is applied to a distributed storage system, the distributed storage system includes a client and a server, the server is used to access a storage device, the storage device is used to store data encoded based on erasure coding, and the method includes: The client receives a first data block written by a user, where the first data block is used to be written into a target stripe in the storage device; The client generates a first temporary check block based on the first data block; The client sends the first data block and the first temporary check block to the server; The server stores the first data block and the first temporary check block in the storage device; The server determines that the first data block is a tail data block of the target stripe, performs erasure coding on the first temporary check block and the second temporary check block to obtain a merged target check block, wherein the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block; The server stores the target check block in the storage device.
2. The method according to claim 1, characterized in that The erasure coding includes a coefficient multiplication operation and an exclusive-OR operation, and the erasure coding merging of the first temporary check block and the second temporary check block includes: The server performs the XOR operation on the first temporary check block and the second temporary check block to obtain a combined target check block.
3. The method according to claim 1 or 2, characterized in that: The client generates a first temporary check block based on the first data block, including: The client determines that the size of the first data block is greater than or equal to a preset threshold, and divides the first data block into at least one third data block having a size equal to the preset threshold; The client performs the erasure coding on the at least one third data block to generate a third temporary check block.
4. The method according to claim 1 or 2, characterized in that: The client generates a first temporary check block based on the first data block, including: The client determines that the size of the first data block is smaller than the preset threshold, and performs the coefficient multiplication operation on the first data block to obtain a fourth temporary check block.
5. The method according to any one of claims 1 to 4, characterized in that: The server determines that the first data block is a tail data block of the target stripe, including: The server determines that the first data block is a tail data block of the target stripe based on a position identifier of the first data block, where the position identifier is used to indicate a storage position of the first data block in the target stripe.
6. The method according to any one of claims 1 to 4, characterized in that: The server determines that the first data block is a tail data block of the target stripe, including: The client sends merge indication information to the server; The server receives the merge indication information and determines that the first data block is the tail data block of the target stripe.
7. A distributed storage system, characterized in that: The distributed storage system includes a client and a server, wherein the server is used to access a storage device, the storage device is used to store data encoded based on erasure codes, the client is used to execute the method executed by the client in any one of claims 1 to 6, and the server is used to execute the method executed by the server in any one of claims 1 to 6.
8. A data writing device, characterized in that: The device is applied to a distributed storage system, the distributed storage system includes a client and a server, the server is used to access a storage device, the storage device is used to store data encoded based on erasure codes, and the device includes: A transceiver unit, configured to receive a first data block written by a user, wherein the first data block is used to be written into a target stripe in the storage device; A processing unit, configured to generate a first temporary check block based on the first data block; The transceiver unit is further used to send the first data block and the first temporary check block to the server; The processing unit is further configured to store the first data block and the first temporary check block in the storage device; The processing unit is further configured to determine that the first data block is a tail data block of the target stripe, and to Performing erasure code merging on the second temporary check block to obtain a merged target check block, wherein the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block; The processing unit is further configured to store the target check block in the storage device.
9. The device according to claim 8, characterized in that The erasure code encoding includes a coefficient multiplication operation and an XOR operation, and the processing unit is specifically used for: The first temporary check block and the second temporary check block are subjected to the XOR operation to obtain a combined target check block.
10. The device according to claim 8 or 9, characterized in that The processing unit is specifically used for: Determine that the size of the first data block is greater than or equal to a preset threshold, and divide the first data block into at least one third data block having a size equal to the preset threshold; The erasure code is performed on the at least one third data block to generate a third temporary check block.
11. The device according to claim 8 or 9, characterized in that The processing unit is specifically used for: It is determined that the size of the first data block is smaller than the preset threshold, and the coefficient multiplication operation is performed on the first data block to obtain a fourth temporary check block.
12. The device according to any one of claims 8 to 11, characterized in that The processing unit is specifically used for: Based on the position identifier of the first data block, it is determined that the first data block is a tail data block of the target stripe, and the position identifier is used to indicate a storage position of the first data block in the target stripe.
13. The device according to any one of claims 8 to 11, characterized in that The transceiver unit is specifically used for: Sending merge indication information to the server; The merge indication information is received, and the first data block is determined to be a tail data block of the target stripe.
14. A computing device, characterized in that The device comprises a processor coupled to a memory, wherein the processor is used to store instructions. When the instructions are executed by the processor, the computing device performs the method according to any one of claims 1 to 6.
15. A computing device cluster, characterized in that: The system comprises at least one computing device, wherein the computing device comprises a processor, wherein the processor is coupled to a memory, and the processor is used to store instructions. When the instructions are executed by the processor, the computing device cluster executes the method according to any one of claims 1 to 6.
16. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed, the computer is caused to execute the method according to any one of claims 1 to 6.
17. A computer program product, comprising instructions, characterized in that: When the instructions are executed, the computer implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data writing method and device
CN120029817A
Request response method and device, equipment and storage medium
CN111475523A
Erasure code data processing method, device and system, storage medium and processor
CN115016979A
Storage system having raid stripe metadata
US20220229730A1
CN202311572685A