Data writing method and device
After computing and transmitting data blocks on the client, the server merges the verification blocks, which solves the problems of client memory and computing complexity, and reduces the storage needs of the server, improving the efficiency and space utilization of the storage system.
Patent Information
- Application Number
- CN202410050635.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-01-12
- Publication Date
- 2025-05-23
AI Technical Summary
In a storage system based on erasure coding technology, the client needs to cache a large number of temporary verification blocks, resulting in waste of memory space. At the same time, the computing complexity is high, and the server also needs to store a large number of verification blocks, resulting in waste of storage space.
The client calculates the verification block based on the written data. When the strip is full, the server merges the verification blocks to reduce the memory consumption of the client.
By handing over the task of merging the verification blocks to the server, the client does not need to store temporary verification blocks, which reduces the client's memory consumption, and reduces the number of temporary verification blocks stored on the server, improving the utilization of storage space.
Smart Images

Figure CN120029817A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 22, 2023, with application number 202311572685.2 and invention name “A method, device and other equipment for data processing”, the entire contents of which are incorporated by reference in this application. Technical Field
[0002] The embodiments of the present application relate to the field of cloud computing, and in particular to a data writing method and device. Background Art
[0003] With the rapid growth of data scale and the increasing demand for data protection, in order to improve the reliability of storage systems, storage systems use erasure coding (EC) technology to ensure high data reliability with low redundant storage overhead. In erasure coding technology, written data is divided into multiple data blocks and additional check blocks are created. These data blocks and check blocks are combined to form a higher redundancy set. If any data block is damaged or lost, the remaining data blocks and check blocks can be used together to recover the lost data block.
[0004] In the current data writing process based on erasure coding technology, in order to improve storage utilization while reducing the storage space occupied by the check block, the client writes the check block in the server using incremental erasure coding. That is, the client calculates a temporary check block for the data block first written into the stripe, and calculates a new temporary check block for the data block to be written into the stripe subsequently based on the temporary check block, until the data block fills the stripe.
[0005] In the incremental erasure code scheme, the data block newly written into the stripe needs to generate a temporary check block based on the last written data block to calculate the check block. Therefore, the client needs to cache a large number of temporary check blocks, resulting in a waste of client memory space. At the same time, the calculation complexity of the check block based on the written data and the temporary check block is high. At the same time, the server also needs to store a large number of check blocks, resulting in a waste of storage space. Summary of the invention
[0006] The embodiment of the present application provides a data writing method, wherein the client calculates a check block based on the written data, and when the stripe is full, the server merges the check blocks, thereby reducing the client's memory consumption. The embodiment of the present application also provides a data writing device, a computing device, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the data writing method.
[0007] In the first aspect, an embodiment of the present application provides a data writing method, which can be executed by a distributed storage system, or by a component of the distributed storage system, such as a processor, chip or chip system of the distributed storage system, or by a logic module or software that can realize all or part of the functions of the distributed storage system. In the embodiment of the present application, the distributed storage system includes a client and a server, the server is used to access a storage device, and the storage device is used to store data encoded based on erasure codes. The method provided in the first aspect includes: the client receives a first data block written by a user, and the first data block is used to write a target stripe in the storage device. The client generates a first temporary check block based on the first data block. The client sends the first data block and the first temporary check block to the server. The server stores the first data block and the first temporary check block to the storage device. The server determines that the first data block is the tail data block of the target stripe, and performs erasure code merging on the first temporary check block and the second temporary check block to obtain a merged target check block, and the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block. The server stores the target check block to the storage device.
[0008] In the embodiment of the present application, the client in the distributed storage system calculates a temporary check block based on the written data, and when the stripe of the storage device is full, the server merges the temporary check blocks to obtain the target check block. Compared with the current client calculating the erasure code for the temporary check block and the newly written data, the client in the embodiment of the present application does not need to store the temporary check block, thereby reducing the client's memory consumption.
[0009] In one possible implementation, the erasure code encoding includes a coefficient multiplication operation and an XOR operation. In the process of erasure coding the first temporary check block and the second temporary check block, the server performs an XOR operation on the first temporary check block and the second temporary check block to obtain a merged target check block.
[0010] In the embodiment of the present application, the server can perform an XOR operation based on the first temporary check block and the second temporary check block to obtain a merged target check block, thereby reducing the number of temporary check blocks stored by the server, reducing the server's consumption of storage space, and improving the feasibility of the solution.
[0011] In a possible implementation, during the process of the client generating the first temporary check block based on the first data block, the client determines that the size of the first data block is greater than or equal to a preset threshold, divides the first data block into at least one third data block whose size is the preset threshold, and performs erasure coding on the at least one third data block to generate a third temporary check block.
[0012] In the embodiment of the present application, the server can split the data block according to the size of the write data block. When the size of the write data block is greater than a preset threshold, the client can perform erasure coding on the split write data to generate a temporary check block, thereby reducing the number of temporary check blocks and saving the transmission bandwidth between the client and the server.
[0013] In one possible implementation, during the process of the client generating the first temporary check block based on the first data block, the client determines that the size of the first data block is less than a preset threshold, and performs a coefficient multiplication operation on the first data block to obtain a fourth temporary check block. Specifically, the client performs a coefficient multiplication operation on the first data block based on the check coefficient to obtain the fourth temporary check block.
[0014] In the embodiment of the present application, the server can generate a temporary check block using different encoding methods according to the size of the written data block, thereby improving the feasibility of the solution.
[0015] In a possible implementation, when the server determines that the first data block is the tail data block of the target stripe, the server determines that the first data block is the tail data block of the target stripe based on the position identifier of the first data block, and the position identifier is used to indicate the storage position of the first data block in the target stripe. The server determines that the stripe is full and generates a target check block based on the position identifier, that is, the position identifier can indicate the timing for the server to generate the target check block.
[0016] In the embodiment of the present application, the server can identify the tail data block of the target stripe according to the position identifier in the written data block, and determine the timing of merging the check block according to the position identifier in the written data block, thereby improving the storage space utilization rate of the storage device of the server.
[0017] In one possible implementation, the size of the location identifier is 2B, location identifier 00 indicates that the stripe is full, 01 indicates that the data block is the first data block of the data portion of the stripe, 10 indicates that the data block is in the middle of the data portion of the data block stripe, and 11 indicates that the data block is the tail data block of the data portion of the stripe.
[0018] In the embodiment of the present application, the server can indicate the storage position of the first data block in the target stripe according to the content based on the two-byte position identifier, thereby improving the feasibility of the solution.
[0019] In a possible implementation, when the server determines that the first data block is the tail data block of the target stripe, the client sends a merge indication message to the server. The server receives the merge indication message and determines that the first data block is the tail data block of the target stripe. That is, the client can also control the server to actively merge the check blocks through the merge indication message.
[0020] In the embodiment of the present application, the server may also receive the merge indication information sent by the client, and determine that the write data block is the tail data block of the target stripe according to the merge indication information, thereby improving the feasibility of the solution.
[0021] In one possible implementation, the client can send a second temporary check block to the server while generating a first temporary check block based on the first data block. In other words, the client can generate a temporary check block while sending the data block and other temporary check blocks to the server, i.e., encoding and sending are performed in parallel.
[0022] In the embodiment of the present application, the client can send data blocks to the server while performing erasure coding, so that the encoding check block and the sending data block are performed in parallel, thereby improving the data writing efficiency of the client.
[0023] In a second aspect, an embodiment of the present application provides a distributed storage system, which includes a client and a server, the server is used to access a storage device, the storage device is used to store data encoded based on erasure codes, the client is used to execute the method executed by the client in the above-mentioned first aspect or any possible implementation scheme of the first aspect, and the server is used to execute the method executed by the server in the above-mentioned first aspect or any possible implementation scheme of the first aspect.
[0024] In a third aspect, an embodiment of the present application provides a data writing device, which is applied to a distributed storage system, wherein the distributed storage system includes a client and a server, wherein the server is used to access a storage device, and the storage device is used to store data encoded based on an erasure code, and the device includes a transceiver unit and a processing unit. Among them, the transceiver unit is used to receive a first data block written by a user, and the first data block is used to write a target stripe in the storage device. The processing unit is used to generate a first temporary check block based on the first data block. The transceiver unit is also used to send the first data block and the first temporary check block to the server. The processing unit is also used to store the first data block and the first temporary check block to the storage device. The processing unit is also used to determine that the first data block is the tail data block of the target stripe, and merge the first temporary check block and the second temporary check block by erasure code to obtain a merged target check block, and the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block. The processing unit is also used to store the target check block to the storage device.
[0025] In a possible implementation, the erasure code encoding includes a coefficient multiplication operation and an XOR operation, and the processing unit is specifically configured to perform an XOR operation on the first temporary check block and the second temporary check block to obtain a merged target check block.
[0026] In a possible implementation, the processing unit is specifically configured to determine that the size of the first data block is greater than or equal to a preset threshold, divide the first data block into at least one third data block whose size is the preset threshold, and perform erasure coding on the at least one third data block to generate a third temporary check block.
[0027] In a possible implementation manner, the processing unit is specifically configured to determine that the size of the first data block is smaller than a preset threshold, and perform a coefficient multiplication operation on the first data block to obtain a fourth temporary check block.
[0028] In a possible implementation, the processing unit is specifically configured to determine that the first data block is a tail data block of the target stripe based on a position identifier of the first data block, where the position identifier is used to indicate a storage position of the first data block in the target stripe.
[0029] In a possible implementation manner, the transceiver unit is specifically configured to send merge indication information to the server, receive the merge indication information, and determine that the first data block is a tail data block of the target stripe.
[0030] In a fourth aspect, an embodiment of the present application provides a computing device, comprising a processor, the processor being coupled to a memory, the processor being used to store instructions, and when the instructions are executed by the processor, the computing device executes the method described in the first aspect or any possible implementation manner of the first aspect.
[0031] In a fifth aspect, an embodiment of the present application provides a computing device cluster, the computing device cluster includes one or more computing devices, the computing device includes a processor, the processor is coupled to a memory, the processor is used to store instructions, when the instructions are executed by the processor, the computing device cluster executes the method described in the first aspect or any possible implementation method of the first aspect.
[0032] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed, the computer executes the method described in the first aspect or any possible implementation manner of the first aspect.
[0033] In a seventh aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed, the computer implements the method described in the first aspect or any possible implementation method of the first aspect.
[0034] It can be understood that the beneficial effects that can be achieved by any of the distributed storage systems, data writing devices, computing devices, computing device clusters, computer-readable media or computer program products provided above can be referred to the beneficial effects in the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 A schematic diagram of the system architecture of a distributed storage system provided in an embodiment of the present application;
[0036] Figure 2 A schematic diagram of a data writing method provided in an embodiment of the present application;
[0037] Figure 3 A schematic diagram of a data portion and a check portion of a stripe provided in an embodiment of the present application;
[0038] Figure 4 A schematic diagram of a stripe for writing a data block provided in an embodiment of the present application;
[0039] Figure 5 A schematic diagram of another stripe for writing data blocks provided in an embodiment of the present application;
[0040] Figure 6 A schematic diagram of another stripe for writing data blocks provided in an embodiment of the present application;
[0041] Figure 7 A schematic diagram of encoding and writing in parallel provided in an embodiment of the present application;
[0042] Figure 8 A schematic diagram of a merged check block provided in an embodiment of the present application;
[0043] Fig. 9 A schematic diagram of a data writing device provided in an embodiment of the present application;
[0044] Fig.10 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0045] Fig.11 A schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0046] Fig.12 A schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] The embodiments of the present application provide a data writing method and device for reducing the bandwidth consumption of a cloud-side server during data writing.
[0048] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0049] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0050] First, some terms involved in the embodiments of the present application are introduced to facilitate technical personnel in this field to understand the technical solution.
[0051] A stripe is a unit of data distribution. Data distribution evenly distributes continuous data blocks to multiple disk drives and allows the system to read and write from multiple disks at the same time, thereby processing IO operations in parallel. The data in a stripe is physically divided. For example, part of a file may be stored on one disk and the other part on another disk. This data distribution allows data to be read or written from multiple disks at the same time, greatly increasing data throughput.
[0052] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.
[0053] See also Figure 1 , Figure 1 A schematic diagram of the system architecture of a data writing system provided in an embodiment of the present application. Figure 1 In the system architecture shown, the distributed storage system 10 includes a client 100, a server 200 and a storage device 300, wherein the storage device 300 includes one or more storage devices distributed in different data centers, wherein the client 100 generates a check block for the written data based on an erasure code algorithm, and writes the data and the check block on the server 200, and the server 200 stores the data block and the check block written by the client 100. The specific functions of the client 100 and the server 200 are respectively introduced below.
[0054] The client 100 refers to a device or application that initiates a storage request to the server 200. Specifically, the client 100 can divide the write data into blocks, and encode the data blocks based on the erasure code algorithm to generate check blocks. The client 100 uploads both the data blocks and the check blocks to the server 200, and the server 200 stores these data blocks and the check blocks in multiple storage devices distributed in data centers in different regions, thereby increasing the data's resistance to failure.
[0055] The client 100 is also used to download data and recover faulty data blocks through the server 200. For example, if a data block is damaged or lost during the transmission process, the client 100 can reconstruct the complete data through the remaining data blocks and check blocks.
[0056] It is understandable that the client 100 can be different devices or applications in different application scenarios. For example, in a personal storage scenario, the client 100 can be a terminal device, such as a personal computer, a smart phone, or a tablet computer, or a software client on the terminal device, such as an application. The client 100 uploads data to a cloud storage server via a network.
[0057] In an enterprise application scenario, the client 100 may be a file server, a database server, or an application server. The client 100 encodes the internal data and then sends it to a data center or private cloud storage. In an IoT application scenario, the client 100 may be a sensor or device. The client 100 generates data and sends it to a backend server for analysis and backup.
[0058] The server 200 is a device that maintains storage resources and processes storage requests from the client 100. Specifically, the server 200 is used to receive data blocks and check blocks sent by the client 100, and store the data blocks and check blocks in the storage device 300, which includes one or more storage devices distributed in data centers in different regions.
[0059] The server 200 is also used to perform erasure coding again on the check block generated by the client 100 to generate a merged check block, and store the merged check block to the storage device 300. The server 200 is also used to allow data blocks to be recovered from single or multiple disk blocks or storage node failures.
[0060] It is understandable that the distributed storage device end 200 can be different devices in different application scenarios. In some scenarios, the distributed storage device end 300 can also be a distributed disk array in a single storage server, which is not specifically limited.
[0061] The distributed storage end 300 may also be applied as a storage device in other application scenarios. For example, the distributed storage end 300 may also be a storage device in a data center application scenario, a distributed file system application scenario, a disaster recovery backup scenario, and the like.
[0062] The client 100 and the server 200 in the embodiment of the present application can be deployed in different devices respectively, or can be deployed in the same device, without specific limitation.
[0063] based on Figure 1 The distributed storage system 10 shown in the figure, the present application also provides a data writing method. The data writing method provided by the embodiment of the present application is introduced below in combination with the embodiment.
[0064] See also Figure 2 , Figure 2 A flow chart of a data writing method provided in an embodiment of the present application. Figure 2 In the example shown, the method includes the following steps:
[0065] Step 201: The client receives a first data block written by a user, and generates a first temporary check block based on the first data block.
[0066] The client 100 receives a first data block, which is data to be stored in a stripe of the storage device 300. The client 100 generates a first temporary check block based on the first data block, wherein a stripe refers to a unit of data distribution in the server 200, and a stripe can be physically divided, for example, a stripe can include storage space of different disk blocks.
[0067] Specifically, in the process of the client 100 generating a check block based on the first data block, the client 100 needs to divide the write data into blocks according to the size of the first data block, and generate a first temporary check block based on the divided data block, and then store the data block and the first temporary check block to the server 200. The first data block and the first temporary check block can be stored on different storage devices respectively.
[0068] See also Figure 3 , Figure 3 A schematic diagram of a stripe of a data block stored in a storage device provided in an embodiment of the present application. Figure 3 In the example shown, the client 100 needs to write the write data to the stripe of the storage device 300. Figure 3It can be seen from the example shown that the storage device 300 includes 6 disk blocks, which can be distributed in different storage devices, and the size of each disk block is, for example, 1M. Among the 6 disk blocks, 4 disks are used to store data blocks, namely D0, D1, D2 and D3, and 2 disk blocks are used to store parity blocks, namely P0 and P1.
[0069] exist Figure 3 In the example shown, the storage space of each disk block can be divided into different units, and the division granularity of each unit is, for example, 8K, where the units of different disk blocks form a stripe. Figure 3 The units in each row of 6 disk blocks form a stripe, that is, the size of a stripe is 48K. In a stripe, the part used to store data blocks is called the data part of the stripe, and the part used to store parity blocks is called the parity part of the stripe. For example, the first 4 8K units in the stripe are the data part of the stripe, and the last 2 8K units of the stripe are the parity part of the stripe.
[0070] exist Figure 3 In the example shown, the server 200 stores the received data block in the data part of the stripe, and the server 200 stores the received check block in the check part of the stripe. It should be noted that when the server 200 receives a data block that cannot fill the data part of the stripe, the server 200 can store the data block in turn in the remaining data part of the stripe until the data part of the stripe is filled, thereby improving the storage space utilization of the data part of the stripe. When the data part of the stripe is full, the server 200 stores the data block in the data part of the next stripe.
[0071] In the embodiment of the present application, the client 100 can generate the first temporary check block in different ways based on the size of the first data block written, and the first temporary check block includes the third temporary check block and the fourth temporary check block. The following is a detailed introduction:
[0072] In a possible implementation, when the first data block to be written is larger than a first threshold, wherein the first threshold is, for example, the size of a unit, i.e., 8K, the written data larger than the first threshold in the embodiment of the present application is also referred to as large input and output (IO) data. The client 100 divides the first data block to be written into at least one third data block of a size equal to a preset threshold based on the first threshold, and performs erasure coding on the at least one third data block to generate a third temporary check block.
[0073] See also Figure 4 , Figure 4 A schematic diagram of dividing written data into blocks and generating check blocks provided in an embodiment of the present application. Figure 4In the example shown, when the write data is greater than the first threshold, the client 100 divides the write data into different data blocks based on the first threshold, and performs erasure code calculation on the different data blocks to obtain a third temporary check block.
[0074] For example, in Figure 4 In the example shown, when the write data is 16K data, since the 16K data is greater than the first threshold value 8K, the client 100 divides the 16K write data into blocks based on 8K to obtain two 8K data blocks. The client 100 performs two erasure code calculations on the two 8K data blocks to obtain two 8K check blocks. Specifically, in the process of the client 100 performing two erasure code calculations on the two 8K data blocks, in the first calculation, the two 8K data blocks are multiplied by the check coefficients C00 and C01 respectively, and then an exclusive OR (XOR) operation is performed to obtain the first check block, and then in the second calculation, the two 8K data blocks are multiplied by the check coefficients C10 and C20 respectively, and then an exclusive OR (XOR) operation is performed to obtain the second check block.
[0075] In one possible implementation, when the first data block to be written is less than or equal to a first threshold, where the first threshold is, for example, the size of a unit, i.e., 8K, the written data less than or equal to the first threshold in the embodiment of the present application is also referred to as small input and output (IO) data. The client 100 performs a coefficient multiplication operation on the first data block based on the check coefficient to obtain a fourth temporary check block.
[0076] See also Figure 5 , Figure 5 A schematic diagram of dividing written data into blocks and generating check blocks provided in an embodiment of the present application. Figure 5 In the example shown, when the written data is less than or equal to the first threshold, the client 100 generates a first temporary check block based on the written data, and multiplies the written data based on the check coefficient to obtain a fourth temporary check block.
[0077] For example, in Figure 5 In the example shown, when the written data is 4K data, since the 4K data is less than the first threshold 8K, the client 100 performs two erasure codes on the 4K written data based on different check coefficients to obtain two 4K check blocks. Specifically, during the process of creating two copies of the 4K data block based on different check coefficients, the client 100 multiplies the 4K data block by two different check coefficients C00 and C01 respectively to obtain two check blocks.
[0078] Step 202: The client sends a first data block and a first temporary check block to the server.
[0079] After generating the first temporary check block based on the first data block, the client 100 sends the first data and the first temporary check block to the server 200 .
[0080] Step 203: The server stores the first data block and the first temporary check block in a storage device.
[0081] After receiving the first data block and the first temporary check block sent by the client 100, the server 200 stores the first data block and the first temporary check block in the storage device. Specifically, the server 200 stores the received data block in a disk block in the storage device for storing data blocks, that is, the data portion of the stripe in the storage device 300, and stores the first temporary check block in a disk block in the storage device for storing check blocks, that is, the check portion of the stripe in the storage device 300.
[0082] Please continue reading Figure 4 ,exist Figure 4 In the example shown, the client 100 sends two 8K data blocks and two 8K check blocks to the server 200, wherein the server 200 stores the two 8K data blocks in the units of disk block D0 and disk block D1, i.e., the first two units of the data portion of the stripe, and stores the two 8K check blocks in the units of disk block P0 and disk block P1, respectively.
[0083] Please continue reading Figure 5 ,exist Figure 5 In the example shown, the client 100 sends a 4K data block and two 4K check blocks to the server 200, wherein the server 200 stores the 4K data block in the unit of disk block D0, i.e., the first unit of the data part of the stripe, and stores the two 4K check blocks in the units of disk block P0 and disk block P1, respectively.
[0084] It should be noted that, in the scenario of appending, the server 200 sequentially writes the data blocks corresponding to the received multiple write data into the data part of the stripe in the server 200. When the server 200 fills the data part of the stripe, the server 200 continues to write data blocks into the data part of the new stripe. At the same time, the server 200 also sequentially writes the check blocks corresponding to the multiple write data into the disk blocks used to store the check blocks.
[0085] See also Figure 6 , Figure 6 A schematic diagram of another data writing method provided in an embodiment of the present application. Figure 6In the example shown, the client 100 needs to write 16K of write data to the server 200. The client 100 first divides the 16K write data into blocks to obtain two 8K data blocks, and performs two erasure code calculations on the two 8K data blocks to obtain two 8K check blocks. The client 100 sends two 8K data blocks and two 8K check blocks to the server 200. The server 200 stores the two 8K data blocks in the units of disk block D0 and disk block D1, respectively, and stores the two 8K check blocks in the units of disk block P0 and disk block P1, respectively.
[0086] After that, the client 100 needs to continue to write 4K write data to the server 200. The client 100 obtains two 4K check blocks based on the 4K write data, and sends the 4K data block and the two 4K check blocks to the server 200. The server 200 continues to store the 4K data block in the unit of the disk block D2, and continues to store the two 4K check blocks in the units of the disk block P0 and the disk block P1 respectively.
[0087] exist Figure 6 In the example shown, when the client 100 needs to continue writing data in the server 200, the server 200 stores the received data blocks in the remaining storage space of the data portion of the stripe in sequence until the data portion of the stripe is full. For example, the server 200 continues to store the 4K data blocks received for the third time in the remaining 4K storage space of the unit of disk block D2. The server 200 continues to store the 4K data blocks received for the fourth and fifth times in the storage space of the unit of disk block D3.
[0088] In an embodiment of the present application, when the client 100 writes data in the data part of the stripe, the client 100 maintains a data integrity field (DIF) for each data block. For example, the client 100 maintains a data integrity field for a 4k data block, and the size of the data integrity field is 64B.
[0089] Among them, the data integrity field includes a position identifier, the size of which is 2B, and the position identifier is used to record the position of the data block in the data part of the stripe. For example, 00 indicates that the stripe is full (full stripe), 01 indicates that the data block is the first data block of the data part of the stripe (start), 10 indicates that the data block is in the middle of the data part of the stripe (on the way), and 11 indicates that the data block is the last data block of the data part of the stripe (final).
[0090] It can be understood that the check block corresponding to the data block can also maintain the data integrity field, and the position identifier corresponding to the check block indicates the position of the data block corresponding to the check block in the data part of the stripe.
[0091] Please continue to see 6. Figure 6 In the example shown, each 4k data stored in the server 200 maintains a 64B data integrity field, wherein 2B in the data integrity field is a position identifier, which is used to record the position of the data block in the data portion of the stripe. For example, the position identifier corresponding to the first 4k data of the disk block D0 of the stripe is 01, the position identifier corresponding to the first 4k data of the disk block D2 of the stripe is 10, and the position identifier of the second 4k data of the disk block D3 in the stripe is 11.
[0092] In one possible implementation, the client 100 can send a second temporary check block to the server 200 while generating a first temporary check block based on the first data block. That is, the client 100 can send the data block and other check blocks to the server 200 at the same time as generating the check block, that is, encoding and sending are performed in parallel.
[0093] Please continue reading Figure 6 ,exist Figure 6 In the example shown, the client 100 needs to write 16K write data in the storage device 300. The client 100 first divides the 16K write data into blocks to obtain two 8K data blocks, and performs two erasure code calculations on the two 8K data blocks to obtain two 8K check blocks. While the client 100 performs erasure code calculations on the two 8K data blocks, the client 100 can send the two 8K data blocks to the server 200.
[0094] exist Figure 6 In the example shown, when the client 100 needs to continue writing 4K of write data on the server 200, the client 100 generates two 4K check blocks based on the 4K write data. While generating the 4K check block, the client 100 can also send the 8K check block generated by the last calculation to the server 200.
[0095] See also Figure 7 , Figure 7 A schematic diagram of a client generating a check block and sending a data block in parallel provided in an embodiment of the present application. Figure 7In the example shown, the client 100 can send data blocks to the server 200 while performing erasure code calculation based on the data blocks to generate the check blocks. For example, in the process of performing erasure code calculation based on data blocks 1 and 2 to obtain the check block 1, the client 100 can write the data block 1 to the disk block D0 of the server 200 and write the data block 1 to the disk block D1 of the server 200 in parallel.
[0096] exist Figure 7 In the example shown, after the client 100 performs erasure code calculation based on data block 1 and data block 2 to obtain check block 1, the check block 1 is written to the disk block P0 of the server 200. At this time, the client 100 can perform erasure code calculation based on data block 1 and data block 2 in parallel to obtain a new check block.
[0097] Step 204. The server determines that the first data block is the tail data block of the target stripe, and performs erasure coding on the first temporary check block and the second temporary check block to obtain a merged target check block, and the second temporary check block includes one or more temporary check blocks generated from other data blocks in the target stripe except the first data block. After the server 200 receives the first temporary check block sent by the client 100, when the first temporary check block generates a check block for the tail data block of the stripe, the server 200 performs erasure coding calculation on the first temporary check block and the second temporary check block to obtain a target check block, which may also be referred to as a merged check block in this application, and the second temporary check block includes one or more check blocks generated based on other data blocks in the stripe except the tail data block.
[0098] That is to say, when the data portion of the stripe in the server 200 is fully written, the server 200 merges the check blocks corresponding to the data blocks in the stripe once to obtain the target check block. Specifically, the server 200 identifies the timing of merging the check blocks based on the position identifier of the temporary check block. When the position identifier indicates that the check block is the check block corresponding to the tail data block of the data portion of the stripe, the server 200 merges all the check blocks corresponding to the data blocks of the data portion of the stripe once to obtain the target check block.
[0099] Please continue reading Figure 6 ,exist Figure 6In the example shown, the client 100 writes data blocks in the data portion of the stripe of the storage device 300. For example, 16k data is written for the first time and stored in disk blocks D0 and D1 of the data portion of the stripe, respectively; 4k data is written for the second time and stored in disk block D2 of the data portion of the stripe; 4k data is written for the third time and also stored in disk block D2 of the data portion of the stripe; 4k data is written for the fourth time and stored in disk block D3 of the data portion of the stripe; and 4k data is written for the fifth time and stored in disk block D3 of the data portion of the stripe, wherein the stripe is full after the fifth write.
[0100] Accordingly, in Figure 6 In the example shown, the two 8K check blocks corresponding to the 16k data block written by the client 100 for the first time are stored in the disk blocks P0 and P1. Since the 16k data is the header data block of the stripe, the position of the check block is identified as 01. The two 4K check blocks corresponding to the 4k data block written by the client 100 for the second time are also stored in the disk blocks P0 and P1. The two 4K check blocks corresponding to the 4k data block written by the client 100 for the third time are also stored in the disk blocks P0 and P1. The two 4K check blocks corresponding to the 4k data block written by the client 100 for the fourth time are stored in the disk blocks P0 and P1. Since the data written from the second to the fourth time is the middle data block of the stripe, the position of the three check blocks from the second to the fourth time is identified as 10. The two 4K check blocks corresponding to the 4k data block written by the client 100 for the fifth time are stored in the disk blocks P0 and P1. Since the 16k data is the header data block of the stripe, the position of the check block is identified as 11.
[0101] exist Figure 6 In the example shown, when the server 200 identifies that the position identifier of the check block is 11, the server 200 merges the check blocks. The server 200 merges all the check blocks corresponding to the data blocks of the stripe data part, that is, the server 200 merges the 8K check block written for the first time, the 4K check block written for the second time, the 4K check block written for the third time, the 4K check block written for the fourth time, and the 4K check block written for the fifth time to obtain a merged 8K check block. After the server 200 obtains the merged check block, it rewrites the merged check block into the disk block of the stripe check part.
[0102] In a possible implementation, the server 200 generates a target check block when the data portion of the stripe is filled and then merges the temporary check block. The client 100 can also control the server 200 to actively merge the check blocks, which is not specifically limited. For example, the client 100 sends a merge indication message to the server 200. The server 200 receives the merge indication message and determines that the first data block is the tail data block of the target stripe according to the merge indication message.
[0103] See also Figure 8 , Figure 8 A schematic diagram of a merged check block provided in an embodiment of the present application. Figure 8 In the example shown, the server 200 merges 5 check blocks, where the position of the first 4k check block is identified as 01, i.e., the check block corresponding to the head data block of the stripe, and the position of the fifth 4k check block is identified as 11, i.e., the check block corresponding to the tail data block of the stripe.
[0104] exist Figure 8 In the example shown, during the process of merging 5 check blocks by the server 200, the server 200 performs the merge check at the granularity of 8K units, that is, the server 200 combines 2 4K check blocks into 1 8K check block, performs an XOR operation on the 3 8K check blocks, and obtains a merged 8K check block. The position of the merged 8K check block is identified as 00, which indicates the check block corresponding to the full stripe.
[0105] Step 205. The server 200 stores the target check block in a storage device.
[0106] The server 200 performs erasure coding on the first temporary check block and the second temporary check block to obtain a merged target check block. Then, the server 200 stores the target check block in the storage device 300 .
[0107] It can be seen from the above embodiments that in the distributed storage system in the embodiments of the present application, the client obtains a temporary check block based on the written data calculation, and when the stripe of the storage device is full, the server merges the temporary check blocks to obtain the target check block, so that the client does not need to store the temporary check blocks, thereby reducing the client's memory consumption during the data writing process.
[0108] Based on the above method embodiment, the embodiment of the present application also provides a data writing device. The data writing device provided by the embodiment of the present application is described in detail below.
[0109] See also Fig. 9 , Fig. 9 A schematic diagram of the structure of a data writing device provided in an embodiment of the present application. Fig. 9 In the example shown, the data writing device 900 is used to implement the various steps performed by the distributed storage system in the above embodiments. The data writing device 900 includes a transceiver unit 901 and a processing unit 902 .
[0110] Among them, the transceiver unit 901 is used to receive the first data block written by the user, and the first data block is used to write the target stripe in the storage device. The processing unit 902 is used to generate a first temporary check block based on the first data block. The transceiver unit 901 is also used to send the first data block and the first temporary check block to the server. The processing unit 902 is also used to store the first data block and the first temporary check block to the storage device. The processing unit 902 is also used to determine that the first data block is the tail data block of the target stripe, and merge the first temporary check block and the second temporary check block with error correction code to obtain a merged target check block, and the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block. The processing unit 902 is also used to store the target check block to the storage device.
[0111] In a possible implementation, the erasure code encoding includes a coefficient multiplication operation and an XOR operation, and the processing unit 902 is specifically configured to perform an XOR operation on the first temporary check block and the second temporary check block to obtain a merged target check block.
[0112] In a possible implementation, the processing unit 902 is specifically configured to determine that the size of the first data block is greater than or equal to a preset threshold, divide the first data block into at least one third data block whose size is the preset threshold, and perform erasure coding on the at least one third data block to generate a third temporary check block.
[0113] In a possible implementation manner, the processing unit 902 is specifically configured to determine that the size of the first data block is smaller than a preset threshold, and perform a coefficient multiplication operation on the first data block to obtain a fourth temporary check block.
[0114] In a possible implementation, the processing unit 902 is specifically configured to determine that the first data block is a tail data block of the target stripe based on a position identifier of the first data block, where the position identifier is used to indicate a storage position of the first data block in the target stripe.
[0115] In a possible implementation manner, the transceiver unit 901 is specifically configured to send merge indication information to the server, receive the merge indication information, and determine that the first data block is the tail data block of the target stripe.
[0116] It should be understood that the division of the units in the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And the units in the device can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some units can also be implemented in the form of software calling through processing elements, and some units can be implemented in the form of hardware. For example, each unit can be a separately established processing element, or it can be integrated in a certain chip of the device. In addition, it can also be stored in the memory in the form of a program, and called and executed by a certain processing element of the device. The function of the unit. In addition, all or part of these units can be integrated together, or they can be implemented independently. The processing element described here can also be a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each unit above can be implemented by an integrated logic circuit of hardware in the processor element or in the form of software calling through a processing element.
[0117] It is worth noting that, for the above method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the described order of actions. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required for the present application.
[0118] Other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with the fact that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0119] See also Fig.10 , Fig.10 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. Fig.10 As shown, the computing device 1000 includes: a processor 1001, a memory 1002, a communication interface 1003 and a bus 1004. The processor 1001, the memory 1002 and the communication interface 1003 are coupled via a bus (not marked in the figure). The memory 1002 stores instructions. When the execution instructions in the memory 1002 are executed, the computing device 1000 executes the method executed by the computing device in the above method embodiment.
[0120] The computing device 1000 may be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASIC), or one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA), or a combination of at least two of these integrated circuit forms. For another example, when a unit in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For another example, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0121] The processor 1001 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.
[0122] The memory 1002 may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. Among them, the nonvolatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0123] The memory 1002 stores executable program codes, and the processor 1001 executes the executable program codes to respectively implement the functions of the aforementioned units or modules, thereby implementing the aforementioned data writing method. That is, the memory 1002 stores instructions for executing the aforementioned data writing method.
[0124] The communication interface 1003 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0125] In addition to the data bus, the bus 1004 may also include a power bus, a control bus, a status signal bus, etc. The bus may be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus may be divided into an address bus, a data bus, a control bus, etc.
[0126] See also Fig.11 , Fig.11 A schematic diagram of a computing device cluster provided in an embodiment of the present application. Fig.11 As shown, the computing device cluster 1100 includes at least one computing device 1000 .
[0127] like Fig.11 As shown, the computing device cluster 1100 includes at least one computing device 1000. The memory 1002 in one or more computing devices 1000 in the computing device cluster 1100 may store the same instructions for executing the above data writing method.
[0128] In some possible implementations, the memory 1002 of one or more computing devices 1000 in the computing device cluster 1100 may also store some instructions for executing the above data writing method. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the above data writing method.
[0129] It should be noted that the memory 1002 in different computing devices 1000 in the computing device cluster 1100 can store different instructions, which are respectively used to execute part of the functions of the above-mentioned data writing device. That is, the instructions stored in the memory 1002 in different computing devices 1000 can implement the functions of one or more modules in the transceiver unit and the processing unit.
[0130] In some possible implementations, one or more computing devices 1000 in the computing device cluster 1100 may be connected via a network, which may be a wide area network or a local area network.
[0131] See also Fig.12 , Fig.12A schematic diagram of a computing device in a computing cluster connected via a network provided in an embodiment of the present application. Fig.12 As shown, two computing devices 1000A and 1000B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device.
[0132] In a possible implementation, the memory in the computing device 1000A stores instructions for executing the functions of the transceiver unit, and the memory in the computing device 1000B stores instructions for executing the functions of the processing unit and the display unit.
[0133] It should be understood that Fig.12 The functions of the computing device 1000A shown in FIG. 1000A may also be completed by multiple computing devices. Similarly, the functions of the computing device 1000B may also be completed by multiple computing devices.
[0134] In another embodiment of the present application, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the computing device in the above method embodiment.
[0135] In another embodiment of the present application, a computer program product is provided, the computer program product includes computer execution instructions, the computer execution instructions are stored in a computer readable storage medium. When the processor of the device executes the computer execution instructions, the device executes the method executed by the computing device in the above method embodiment.
[0136] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0137] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0138] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0140] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
Claims
1. A data writing method, characterized in that: The method is applied to a distributed storage system, the distributed storage system includes a client and a server, the server is used to access a storage device, the storage device is used to store data encoded based on erasure coding, and the method includes: The client receives a first data block written by a user, where the first data block is used to be written into a target stripe in the storage device; The client generates a first temporary check block based on the first data block; The client sends the first data block and the first temporary check block to the server; The server stores the first data block and the first temporary check block in the storage device; The server determines that the first data block is a tail data block of the target stripe, performs erasure coding on the first temporary check block and the second temporary check block to obtain a merged target check block, wherein the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block; The server stores the target check block in the storage device.
2. The method according to claim 1, characterized in that The erasure coding includes a coefficient multiplication operation and an exclusive-OR operation, and the erasure coding merging of the first temporary check block and the second temporary check block includes: The server performs the XOR operation on the first temporary check block and the second temporary check block to obtain a combined target check block.
3. The method according to claim 1 or 2, characterized in that: The client generates a first temporary check block based on the first data block, including: The client determines that the size of the first data block is greater than or equal to a preset threshold, and divides the first data block into at least one third data block having a size equal to the preset threshold; The client performs the erasure coding on the at least one third data block to generate a third temporary check block.
4. The method according to claim 1 or 2, characterized in that: The client generates a first temporary check block based on the first data block, including: The client determines that the size of the first data block is smaller than the preset threshold, and performs the coefficient multiplication operation on the first data block to obtain a fourth temporary check block.
5. The method according to any one of claims 1 to 4, characterized in that: The server determines that the first data block is a tail data block of the target stripe, including: The server determines that the first data block is a tail data block of the target stripe based on a position identifier of the first data block, where the position identifier is used to indicate a storage position of the first data block in the target stripe.
6. The method according to any one of claims 1 to 4, characterized in that: The server determines that the first data block is a tail data block of the target stripe, including: The client sends merge indication information to the server; The server receives the merge indication information and determines that the first data block is the tail data block of the target stripe.
7. A distributed storage system, characterized in that: The distributed storage system includes a client and a server, wherein the server is used to access a storage device, the storage device is used to store data encoded based on erasure codes, the client is used to execute the method executed by the client in any one of claims 1 to 6, and the server is used to execute the method executed by the server in any one of claims 1 to 6.
8. A data writing device, characterized in that: The device is applied to a distributed storage system, the distributed storage system includes a client and a server, the server is used to access a storage device, the storage device is used to store data encoded based on erasure codes, and the device includes: A transceiver unit, configured to receive a first data block written by a user, wherein the first data block is used to be written into a target stripe in the storage device; A processing unit, configured to generate a first temporary check block based on the first data block; The transceiver unit is further used to send the first data block and the first temporary check block to the server; The processing unit is further configured to store the first data block and the first temporary check block in the storage device; The processing unit is further configured to determine that the first data block is a tail data block of the target stripe, and perform erasure coding on the first temporary check block and the second temporary check block to obtain a merged target check block, wherein the second temporary check block includes one or more temporary check blocks generated by other data blocks in the target stripe except the first data block; The processing unit is further configured to store the target check block in the storage device.
9. The device according to claim 8, characterized in that The erasure code encoding includes a coefficient multiplication operation and an XOR operation, and the processing unit is specifically used for: The first temporary check block and the second temporary check block are subjected to the XOR operation to obtain a combined target check block.
10. The device according to claim 8 or 9, characterized in that The processing unit is specifically used for: Determine that the size of the first data block is greater than or equal to a preset threshold, and divide the first data block into at least one third data block having a size equal to the preset threshold; The erasure code is performed on the at least one third data block to generate a third temporary check block.
11. The device according to claim 8 or 9, characterized in that The processing unit is specifically used for: It is determined that the size of the first data block is smaller than the preset threshold, and the coefficient multiplication operation is performed on the first data block to obtain a fourth temporary check block.
12. The device according to any one of claims 8 to 11, characterized in that The processing unit is specifically used for: Based on the position identifier of the first data block, it is determined that the first data block is a tail data block of the target stripe, and the position identifier is used to indicate a storage position of the first data block in the target stripe.
13. The device according to any one of claims 8 to 11, characterized in that The transceiver unit is specifically used for: Sending merge indication information to the server; The merge indication information is received, and the first data block is determined to be a tail data block of the target stripe.
14. A computing device, characterized in that The device comprises a processor coupled to a memory, wherein the processor is used to store instructions. When the instructions are executed by the processor, the computing device performs the method according to any one of claims 1 to 6.
15. A computing device cluster, characterized in that: The system comprises at least one computing device, wherein the computing device comprises a processor, wherein the processor is coupled to a memory, and the processor is used to store instructions. When the instructions are executed by the processor, the computing device cluster executes the method according to any one of claims 1 to 6.
16. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed, the computer is caused to perform the method according to any one of claims 1 to 6.
17. A computer program product, comprising instructions, characterized in that: When the instructions are executed, the computer implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Data writing method and apparatus
EP4807562A1
Data writing method and apparatus
WO2025107561A1