Data processing method and data storage system

By uniformly formatting the IO data into a data format with DIF on the client and DIF checksum storage on the hardware offload device, the problem of degradation of IO performance when the storage system processes data in different data formats is solved, and more efficient IO processing is achieved.

CN120010747APending Publication Date: 2025-05-16HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311524253.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the storage system processes data with different data formats, it requires two sets of code processing, resulting in a degradation of IO performance.

Method used

The client uniformly formats the IO data into a data format with DIF, simplifies the processing logic of microcode, and performs DIF checksum storage on the hardware offload device.

Benefits of technology

Through the unified data format, the delay of the IO hardware through path is reduced, the consumption of CPU resources is reduced, and the IO performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010747A_ABST
    Figure CN120010747A_ABST
Patent Text Reader

Abstract

The invention provides a data block processing method which is applied to a storage system, the storage system comprises a client side and a server side, the client side and the server side conduct hardware unloading through hardware unloading equipment, IO hardware is directly communicated with a hard disk, the method comprises the steps that the client side receives a first write request sent by an upper layer application, and the first write request carries to-be-written data; the client processes data to be written to obtain at least one first data block, each first data block comprises a data block to be written and a DIF data block, and the DIF data block comprises public DIF information; the client sends a second write request to the server; and the server responds to the second write request, carries out DIF verification on the at least one data block, and stores the at least one first data block in a hard disk if the DIF verification is passed. According to the method, the client unifies the IO data into the data format with the DIF before the IO straight-through path, and the microcode processing logic is simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud storage technology, and in particular to a data processing method and a data storage system. Background Art

[0002] Customized hardware and software function hardware offloading has become an important competitive advantage for cloud services. Customized chips such as AWS Nitro and Google OneRMA have been commercialized on a large scale. Huawei also has perfect customized chip capabilities. Huawei has developed its own software and hardware direct-connect technology in the backend of cloud service storage. Through technologies such as numerical control separation, persistence log (Persistence log, Plog) semantic offloading, and data integrity field (data integrity field, DIF) calculation offloading, it realizes Plog input / output (Input / Output, IO) hardware direct-connection to the hard disk. Reduce IO latency, save CPU resources, and reduce costs.

[0003] However, storage systems often receive data with inconsistent data formats, such as data with DIF format and data without DIF format. As a result, different data formats need to be processed on the Plog IO hardware pass-through path, resulting in two sets of codes for the entire IO path and a large number of branches that reduce IO performance. Summary of the invention

[0004] The embodiments of the present application provide a data processing method and a data storage system. The client unifies the IO data into a data format with DIF before the IO pass-through path, thereby simplifying the processing logic of the microcode.

[0005] In a first aspect, the present application provides a data processing method, which is applied to a storage system, wherein the storage system includes a client (Client) and a server (Server), wherein the client and the server perform hardware unloading through a hardware unloading device to realize IO hardware direct access to a hard disk, and the method includes a client receiving a first write request sent by an upper-layer application, wherein the first write request carries data to be written; the client processes the data to be written to obtain at least one first data block, wherein each first data block includes a data block to be written and a DIF data block, and the DIF data block includes common DIF information; the client sends a second write request to the server; the server responds to the second write request, performs a DIF check on at least one data block, and if the DIF check passes, stores at least one first data block in a hard disk (disk).

[0006] The data block processing method provided by the present application inserts the received data without DIF information into DIF at the client, unifies the data into a data format with DIF, allows the IO direct path to be processed according to the unified DIF format, simplifies the processing logic of the microcode, and reduces the latency of the IO hardware direct path.

[0007] In a possible implementation, a specific implementation of the client processing the data to be written to obtain at least one first data block is: the client divides the data to be written to obtain at least one data block to be written; the client determines the common DIF information of each data block to be written in the at least one data block to be written, and the common DIF information of each data block to be written includes a cyclic redundancy check code (Cyclic Redundancy Check, CRC) of each data block to be written; the client inserts the common DIF information into the DIF data block in each first data block.

[0008] In another possible implementation, the client's processing of the data to be written also includes the client inserting private DIF information into the DIF data block in each first data block, and the private DIF information includes metadata of the data block to be written in each first data block and a cyclic redundancy check code of the metadata.

[0009] That is to say, the insertion of the complete DIF information of the data to be written is completed on the client, and in the subsequent IO hardware direct path, it is only necessary to process and verify the DIF according to the determined data format.

[0010] Optionally, the client device offloads the calculation of DataCRC in the public DIF and the calculation of metadata (Meta) CRC in the private DIF in the DIF data on the client to the hardware offload device through a hardware offload device (e.g., a hardware offload card), thereby offloading CPU-consuming computational operations to the hardware and reducing CPU overhead.

[0011] It can be understood that the meaning of DataCRC is the CRC calculated for the business data, and the meaning of MetaCRC is the CRC calculated for the metadata of the business data.

[0012] In another possible implementation, the client processes the data to be written to obtain at least one first data block, and the process also includes: the client determines that the first write request is a non-4KB aligned request; and the client performs 4KB alignment processing on the first write request. In other words, the client also performs 4KB alignment processing on the non-4KB aligned data to be written before inserting the DIF.

[0013] Hardware microcode is suitable for processing computing operations, but is not suitable for processing non-4KB aligned requests. This application further reduces the IO latency of IO hardware passing through the hard disk by completing 4KB alignment processing of non-4KB data on the client.

[0014] In another possible implementation, a specific implementation of the client performing 4KB alignment processing on the first write request is that if the client determines that the header of the first write request is not 4KB aligned, the data to be written is padded forward using page cache data, where the page cache data includes the last page data of the last written data; if the client determines that the tail of the first write request is not 4KB aligned, zero padding is performed on the tail of the data to be written to align the tail of the data to be written with 4KB.

[0015] In another possible implementation, the client performs 4KB alignment processing on the first write request, and also updates Page cache data based on the tail data of the data to be written.

[0016] Optionally, the size of the data block to be written is 4KB, and the size of the DIF data block is 64 bytes. Exemplarily, after the data to be written is aligned to 4KB, the data to be written is processed in blocks and divided into data blocks of 4KB in size. Each 4KB data block is then unified into a 4KB data block + a 64-byte DIF data block, and public DIF information and private DIF information are inserted into the 64-byte DIF data block to obtain a 4KB + 64-byte data block with DIF information in a unified format. In the subsequent IO hardware pass-through path, it is only necessary to process and verify the DIF according to the determined data format, thereby simplifying the microcode processing logic.

[0017] In another possible implementation, the hardware offloading device includes a neural processing unit (NPU); the server responds to the second write request, performs a DIF check on each of the at least one first data block, and stores the at least one first data block in the hard disk if the DIF check passes. The specific implementation is: the NPU responds to the second write request, performs a DIF check on each of the at least one first data block, and stores the at least one first data block in the hard disk if the DIF check passes. In other words, the functions of the storage service end of the server device are offloaded to the hardware offloading device for implementation, such as DIF check of data and disk entry operations.

[0018] In another possible implementation, the data processing method provided by the present application further includes the client responding to the first read request of the upper layer application and sending a second read request to the server; the server responding to the second read request sent by the client and reading at least one target data block from the hard disk, the at least one target data block including a business data block and a DIF data block; sending the at least one target data block to the client; the client sending the business data block in each data block in the at least one target data block to the upper layer application. That is, the client converts the read data in the 4KB+64 byte format into data in the business format (i.e. only business data) to facilitate the use of the upper layer application.

[0019] In another possible implementation, the client responds to the first read request of the upper-layer application and performs the following steps before sending the second read request to the server: the client determines that the first read request is a non-4KB aligned read request; the client converts the non-4KB aligned read request into a 4KB aligned read request to obtain the second read request.

[0020] For example, if the offset address of the read request of the upper-layer application is 7KB and the length (size) is 11KB, it is judged as a non-4KB read request, and then a 4KB alignment operation is performed on it, that is, its offset is changed to 4KB to achieve 4KB alignment, so that it can read data of four 4KB data blocks starting from the offset of 4KB from the disk.

[0021] In the second aspect, the present application provides a data storage system, including a client and a server, wherein the client and the server perform hardware unloading through a hardware unloading device to realize IO hardware direct access to a hard disk; the client includes a receiving module, a processing module and a sending module, wherein the receiving module is used to receive a first write request sent by an upper-layer application, and the first write request carries data to be written; the processing module is used to process the data to be written to obtain at least one first data block, wherein each first data block includes a data block to be written and a DIF data block, and the DIF data block includes common DIF information; the sending module is used to send a second write request to the server, and the second write request is used to request that at least one first data block be stored; the server includes a first response module, and the first response module is used to respond to the second write request, perform a DIF check on at least one first data block, and if the DIF check passes, store at least one first data block in the hard disk.

[0022] In one possible implementation, the processing module is specifically used to divide the data to be written to obtain at least one data block to be written; determine the common DIF information of each data block to be written in at least one data block to be written, and the common DIF information of each data block to be written includes the CRC of each data block to be written; insert the common DIF information into the DIF data block in each first data block.

[0023] In another possible implementation, the processing module is further configured to insert private DIF information into the DIF data block in each first data block, where the private DIF information includes Meta of the data block to be written in each first data block and CRC of Meta.

[0024] In another possible implementation, the client further includes a 4KB alignment module, and the 4KB alignment module is used to determine that the first write request is a non-4KB aligned request; and perform 4KB alignment processing on the first write request.

[0025] In another possible implementation, the 4KB alignment module is specifically used to determine that the header of the first write request is not 4KB aligned, and then use the Page cache data to pad the data to be written forward, and the Page cache data includes the last page data of the last written data; if it is determined that the tail of the first write request is not 4KB aligned, then fill the tail of the data to be written with zeros to align the tail of the data to be written with 4KB.

[0026] In another possible implementation, the 4KB alignment module is further used to update the Page cache data based on the tail data of the data to be written.

[0027] Optionally, the size of the data block to be written is 4 KB, and the size of the DIF data block is 64 bytes.

[0028] In another possible implementation, the hardware offloading device includes an NPU; the first response module is deployed in the NPU, and is used to respond to the second write request, perform a DIF check on each first data block in at least one first data block, and store at least one first data block in the hard disk if the DIF check passes.

[0029] In another possible implementation, the client also includes a second response module, which is used to respond to the first read request of the upper-layer application and send a second read request to the server; the first response module is used to respond to the second read request sent by the client, read at least one target data block from the hard disk, and the at least one target data block includes a business data block and a DIF data block; send the at least one target data block to the client; the sending module is used to send the business data block in each data block in the at least one target data block to the upper-layer application.

[0030] In another possible implementation, the 4KB alignment module is further used to determine that the first read request is a non-4KB aligned read request; and convert the non-4KB aligned read request into a 4KB alignable read request to obtain a second read request.

[0031] In the third aspect, an embodiment of the present application provides a network unloading device, which is applied to a server device in a storage system, and is used to perform hardware unloading of the storage service on the server device; the network unloading device includes a processor, and the processor can execute the following steps: receiving a write request sent by a client, the write request carrying at least one data block, each of the at least one data block including a data block to be written and a DIF data block; responding to the write request, performing a DIF check on at least one data block, and if the DIF check passes, storing at least one data block in the hard disk.

[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0033] In a fifth aspect, an embodiment of the present application further provides a computer program or a computer program product, wherein the computer program or the computer program product comprises instructions, and when the instructions are executed, the computer is caused to execute the method described in the first aspect.

[0034] In a sixth aspect, an embodiment of the present application further provides a chip, comprising at least one processor and a communication interface, wherein the processor is used to execute the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the system architecture in the related technology 1;

[0036] Figure 2 A schematic diagram of the structure of a storage system to which the data processing method provided in the embodiment of the present application can be applied;

[0037] Figure 3 A schematic diagram of an implementation architecture of an IO hardware pass-through hard disk provided in an embodiment of the present application;

[0038] Figure 4 A flowchart of a data processing method provided in an embodiment of the present application;

[0039] Figure 5 A schematic diagram showing four common types of DIF data formats;

[0040] Figure 6 A schematic diagram of a 4KB alignment processing process for a non-4KB aligned write request according to an embodiment of the present application is shown;

[0041] Figure 7 An architectural diagram for implementing a data processing method provided in an embodiment of the present application is shown;

[0042] Figure 8A schematic diagram showing the implementation of a data processing method for reading and writing processing provided by an embodiment of the present application is shown;

[0043] Fig. 9 An architectural diagram of a data processing device provided in an embodiment of the present application;

[0044] Fig.10 An architectural diagram of a data processing device provided in an embodiment of the present application;

[0045] Fig.11 It is a structural diagram of a cloud service system according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] The term "and / or" mentioned in this article is a kind of relationship that describes the association of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.

[0047] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, first memory chain data and second memory chain data are used to distinguish different memory chain data rather than to describe a specific order of the memory chain data.

[0048] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0049] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.

[0050] To facilitate understanding of the solutions of the embodiments of the present application, the technical terms involved in this document are first explained below.

[0051] CRC is a process for detecting errors in data transmission. The CRC check generates a number based on the transmitted data through complex calculations. The sending device performs this calculation before sending the data, and then sends the result to the receiving device. After receiving the data, the receiving device repeats the same calculation. If the calculation results of the two devices are the same, the transmission is considered to be correct. This process is called redundant checking because each transmission contains not only data but also additional (redundant) error checking values.

[0052] DIF includes the CRC checksum of user data, logical block address (LBA) information, and other custom information.

[0053] Plog: An append-only ROW writing mechanism.

[0054] Hardware offloading technology: Through dedicated hardware, specific network, computing and other functions are implemented. Hardware offloading is to separate the corresponding functions from the CPU and carry them by the hardware, reducing the CPU occupancy and improving the efficiency of the entire system. The dedicated hardware for offloading is programmable, and the corresponding hardware programming language is microcode.

[0055] Related technology 1: During the storage process, the storage backend adopts different data processing methods for different data format types. For example, when the storage backend receives data without DIF information, it generates private CRC information at the entry point to verify the data on the IO path, and fills the original data with 4K+64 before downloading it to the disk, inserts the complete DIF information, and finally downloads it to the disk (see Figure 1 , in the left branch).

[0056] The storage backend receives data with DIF information. On the IO path, it uses the DIF CRC information in the built-in DIF information to check, supplements the private DIF information before downloading, and finally downloads it to the disk (see Figure 1 , in the right branch).

[0057] This solution has the problem of needing to process two types of data types, requiring two sets of codes for the entire IO path, and causing a large number of branches that reduce IO performance.

[0058] Related technology 2 adopts the solution of storage backend IO hardware passing through the hard disk. For example, IO requests go through the NPU offload path, and management requests go through the service processing unit (SPU). The server side implements IO hardware passing through the hard disk through Plog semantics, EC, DIF technology hardware offload and other technologies.

[0059] This solution offloads CPU-consuming computational operations to hardware, reduces CPU overhead, offloads the main IO path based on the hard disk, simplifies the software system architecture, and improves the scalability and cost-effectiveness of the storage system. However, this solution still has the problem that for received data without DIF format, the hardware needs to perform the operation of inserting DIF information, but the hardware microcode is not suitable for implementing the logic of inserting DIF information. For non-4KB aligned requests, the hardware also needs to perform 4KB alignment operations, but the hardware microcode is suitable for processing computational operations, not 4KB aligned operations.

[0060] In order to solve the above-mentioned technical problems, an embodiment of the present application provides a data processing method, which unifies the data into a data format with DIF on the client, provides a consistent data source for the microcode directly passed through the IO hardware, and simplifies the processing logic of the microcode.

[0061] The specific implementation of the data processing method provided in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0062] Figure 2 Schematic diagram of a storage system to which the data processing method provided in the embodiment of the present application can be applied. Figure 2 As shown, the storage system includes a host 100 and a plurality of storage nodes 200 , and a communication connection can be established between the host 100 and the plurality of storage nodes 200 via a wired or wireless network.

[0063] For example, a network card 102 is provided on the host 100, and a network card 202 is provided on each storage node 200, and the host 100 and the storage node 200 are connected to each other through the network card.

[0064] Each storage node 200 includes multiple hard disks 203, which can be solid state disks (SSD), mechanical hard disks (HDD), shingled magneting recording (SMR) hard disks, etc. All or part of the hard disks 203 in multiple storage nodes can form a storage resource pool, and the storage resources of the storage resource pool can be virtualized into multiple storage units, each storage unit is the smallest unit for storing data in the storage resource pool, and the storage space of each storage unit can come from one or more storage nodes 200, that is, each storage unit can physically span multiple storage nodes 200. Among them, each storage unit can also be called a persistence log (Plog).

[0065] A client 101 is running in the host 100, and the client 101 can provide a data access interface, and can implement functions such as data sharding, routing, online erasure code (EC), online EC writing across availability zones (AZ), and fault handling. For example, the client 101 can receive a write request sent by an application (not shown in the figure) in the host 100, and after processing the write request, send it to the storage node 200 through the network card 102, so that the storage node 200 can store the data to be written of the write request. Among them, the availability zone refers to a geographical area, and there is a certain physical distance between different availability zones. Multiple storage nodes in the storage system can be distributed in one or more availability zones.

[0066] Each storage node 200 in the storage system also runs a server 201, which can manage multiple storage units and can also implement data background EC within a single AZ, data background EC across AZs, and data heat storage.

[0067] In one example, a manager may also run in the host 100. The management segment stores metadata and tasks and is responsible for managing the topology information and task information of the entire storage system, such as processing requests sent by clients, and supervising the status of clients and storage nodes.

[0068] In the embodiments of the present application, the hardware of the client, server, and management side can be unloaded through the hardware unloading device, for example, the hardware unloading can be unloaded to the network card for execution. Of course, some functions of the client, server, and management side can also be unloaded to the network card for execution, for example, the computing type operation hardware on the client and / or server can be unloaded to the network card, such as calculating the CRC of business data, calculating the CRC of metadata, and other computing operations.

[0069] The hardware offloading device may be any feasible hardware offloading device, such as a smart network interface card (smart NIC), an application specific integrated circuit (ASIC), and a data processing unit (DPU).

[0070] For example, you can plug in a smart network card (that is, Figure 2 The network card 102 in the figure is a smart network card), and a smart network card is plugged into the server where the storage node is located (that is, Figure 2The network card 202 in the example is a smart network card), which offloads the IO hardware from the client and the server, and enables the IO hardware to pass directly through the hard disk, thereby reducing the load on the host and storage nodes and increasing the IO speed.

[0071] Figure 2 The storage system shown may be an append only distributed storage system, and in particular may be a cloud storage system.

[0072] It should be pointed out that Figure 2 The storage system shown is only an example and does not constitute a limitation on the embodiments of the present application. It may include more or fewer components. For example, the host 100 also includes a central processing unit (CPU) and memory.

[0073] Figure 2 The storage system shown uses a hardware offload device to perform hardware offload on the client and server through the hardware offload device, realizing IO hardware pass-through to the hard disk. The hardware offload device includes, for example, an NPU and an SPU. The implementation architecture of IO hardware pass-through to the hard disk is as follows: Figure 3 shown.

[0074] like Figure 3 As shown in the figure, IO requests go through the NPU offload path, and management requests go through the SPU to achieve CNC separation. The server implements IO pass-through to the hard disk through hardware offload such as Plog semantics, EC, and DIF calculation.

[0075] Figure 4 A flow chart of a data processing method provided in an embodiment of the present application. The data processing method can be applied to Figure 3 The storage architecture shown in the figure realizes unifying IO data into a data format with DIF before the IO direct path, thus simplifying the processing logic of the microcode. Figure 4 As shown, the data processing method provided in the embodiment of the present application at least includes steps S401 to S404.

[0076] In step S401, the client receives a first write request sent by an upper layer application.

[0077] When the upper layer application needs to write data to the hard disk, it sends a write request to the client of the storage system. The write request carries the business data to be written. For the convenience of description, the business data to be written can be referred to as data to be written.

[0078] In step S402, the client processes the data to be written to obtain at least one first data block, wherein each first data block includes a data block to be written and a DIF data block, and the DIF data block includes public DIF information.

[0079] In order to achieve end-to-end data protection of IO data and solve the problem of silent data destruction in the storage system, DIF data is usually added to the IO data.

[0080] DIF data includes multiple types according to the different sizes of external data blocks (512B or 4KB) of the storage system, as well as different protection information (PI) and Meta organizations. Figure 5 Four common types of DIF data formats are shown, such as Figure 5 As shown, the DIF data format of type 1 is 512B of business data block data + 8B of PI data, which can be referred to as 512+8 type DIF, the DIF data format of type 2 is 4096B of business data block data + 8B of PI data, which can be referred to as 4KB+8 type DIF, the DIF data format of type 3 is 4096B of business data block data + 56B of Meta + 8B of PI data, which can be referred to as 4KB+64 type DIF, and the DIF data format of type 4 is 4096B of business data block data + 8B of PI data + 56B of Meta data, which can be referred to as 4KB+64 type DIF. The embodiment of the present application takes the DIF data format of type 4, 4KB+64 type DIF, as an example to introduce.

[0081] The format of the data to be written carried in the write request received by the client is not uniform. Some of the data to be written is data with DIF data, and some is data without DIF data. This will cause two types of data to be written to be processed in the IO hardware pass-through path, requiring two sets of codes, reducing IO performance. At the same time, the hardware microcode is not suitable for processing the logic of inserting private DIF information.

[0082] In the embodiment of the present application, the format of the data to be written carried in the write request of the client is unified, that is, it is unified into a format with DIF data, for example, a data format of 4KB+64, that is, 4KB is the business data block to be written + 64B of DIF data.

[0083] Exemplarily, for the data to be written in the data format without DIF data in the write request, the client divides the data to be written into data blocks of 4KB in size, and then inserts 64B of DIF data into each 4KB data block to obtain the data to be written in the format of 4KB+64 with DIF data. Before the IO direct path, the IO data is unified into the data format with DIF, providing a consistent data source for the microcode of the IO hardware direct path, simplifying the processing logic of the microcode.

[0084] DIF data includes public DIF information and private DIF information. Public DIF information includes PI information (e.g. Figure 5 8B PI of type 4 data in the private DIF information includes Meta information (e.g. Figure 5 56B of Meta of type 4 data in ).

[0085] Public DIF information includes CRC of 4K service data block Data and version information of DIF information, etc. Private DIF information includes Meta information and CRC of Meta information.

[0086] Meta, also known as intermediary data or relay data, is data about data. It mainly describes the properties of data and is used to support functions such as indicating storage location, historical data, resource search, and file records. Meta is information about the organization of data, data domains, and their relationships. In short, Meta is data about data.

[0087] In order to further increase the IO performance of IO hardware direct pass-through to the hard disk, in an example, the CYC calculation of the business data and the CRC calculation of Meta in the client's public DIF can be hardware offloaded to the hardware offload device for implementation. For example, the CYC calculation of the business data and the CRC calculation of Meta are implemented through the NPU. In this way, the DIF CRC calculation that consumes a lot of CPU resources is offloaded to the hardware processing, reducing CPU usage and further reducing IO latency.

[0088] The write request may be a non-4KB aligned request. For example, if the offset address of the write request is 10KB and the length is 21KB, it can be known that the write request is a non-4KB aligned write request and the head and tail are not 4KB aligned.

[0089] As can be seen from the above description, the microcode is not suitable for processing non-4KB aligned requests. Therefore, in order to further increase the IO performance, the client's processing of write requests also includes performing 4KB alignment processing on non-4KB aligned requests.

[0090] After receiving the write request from the upper-layer application, the client first determines whether the request is a 4KB alignment request. If it is not a 4KB alignment request, it first performs 4KB alignment processing. After 4KB alignment, it determines whether the data to be written carried by the request is data with DIF information. If it does not carry DIF information, it is unified into 4KB+64 format data to be written with DIF information, and the public DIF information and private DIF information are inserted to obtain 4KB+64 format data to be written with complete DIF information.

[0091] The embodiment of the present application uses Plog semantics to perform 4KB alignment processing on non-4KB aligned write requests.

[0092] The semantics of Plog can be understood as an append-only ROW writing mechanism, which has the feature of append-only. That is, after a Plog is full, no new data can be written or modified. It can only be deleted and read.

[0093] For example, when the client receives a non-4KB aligned request, when the header of the data to be written in the request (also called the first page) is not 4KB aligned, the Page cache data is used to fill the header; when the tail of the data to be written in the request (also called the tail page) is not 4KB aligned, the tail is filled with zeros to align to 4KB.

[0094] Figure 6 FIG. 4 shows a schematic diagram of a 4KB alignment processing process for a non-4KB aligned write request according to an embodiment of the present application. Figure 6 As shown, the write request Append req received by the client is offset=7KB, size=11KB, that is, the offset address of the data to be written is 7KB, and the length is 11KB. It can be seen that the beginning and the end of the write request are not 4KB aligned, and both need to be aligned to 4KB.

[0095] First, the page cache data is used to pad the data to be written forward, and then the tail of the data to be written is padded with zeros to 4KB alignment (for example, Figure 6 If the last page of the data to be written has only 2KB, the remaining 2KB will be filled with zeros to achieve the tail 4KB alignment), and finally the pagecache data will be updated based on the tail data of the data to be written, that is, the 2KB data of the last page.

[0096] The design of page cache is based on plog semantics. It will cache the last non-4KB aligned page data of each Plog, organize data using hash tables, and reserve memory according to the maximum number of unseal Plogs. Since page cache caches the last non-4KB aligned page data of each Plog, that is, the data of the last tail page of the last data to be written that is not 4KB aligned, and due to the append write feature of Plog, the last tail page unaligned data to be written will inevitably be spliced ​​with the 4KB aligned data of the first page of this time into a 4KB data. For example, the offset of this write request is 7KB, and the last tail page unaligned data is bound to be 3KB, that is, the size of the current page cache data is 3KB, and the first page unaligned data of this write request is 1KB. The page cache data and the first page unaligned data of the write request can just fill in 4KB of data, so the page cache data can just be used to perform 4KB alignment processing on the header of the write request.

[0097] In this way, the client performs 4KB alignment processing on the write request and inserts DIF data processing to obtain 4KB aligned data to be written in DIF format. In the subsequent IO hardware pass-through path, only the unified format needs to be processed and DIF checked, thereby increasing the IO performance of the IO hardware pass-through path.

[0098] In step S403, the client sends a second write request to the server, where the second write request is used to request to store at least one first data block.

[0099] The client sends a write request to the server. Since the server is offloaded to the smart network card by hardware, the client sends the write request to the smart network card. It can also be understood that there is a storage service client on the smart network card, and the client sends a write request to the storage service client on the smart network card. For example, Figure 3 As shown, after completing the processing of the IO request (such as a write request), the client sends the IO request to the NPU on the smart network card, requesting the NPU to write the data to be written into the hard disk.

[0100] In step S404, the server responds to the second write request and performs a DIF check on the at least one first data block. If the DIF check passes, the at least one first data block is stored in the hard disk.

[0101] After receiving the write request, the server responds to the write request, performs DIF check on the 4KB alignment and the data to be written with DIF information, and writes the data to be written to the hard disk if the DIF check passes.

[0102] Figure 7FIG. 1 shows an architecture diagram for implementing a data processing method provided in an embodiment of the present application. Figure 7 As shown, on the client, 4KB alignment processing is performed on the data to be written of the write request, and DIF information insertion processing (including public DIF information insertion and private DIF information insertion) is performed on the data after 4KB alignment processing to unify it into a 4KB+64 data format with DIF information, and then the data in 4KB+64 format is sent to the server. The server performs DIF verification on the data to be written through the hardware NPU. After the DIF verification passes, the data to be written is stored on the disk (i.e., written to the hard disk). In this way, on the client, the original non-4KB aligned data to be written is aligned to 4KB based on Plog, and then the data format is unified into a 4KB+64 data format with DIF information, and the private DIF information is inserted before the hardware unloading path. The hardware unloading logic only needs to be processed according to the determined format, which simplifies the hardware microcode development logic and reduces the latency of the IO hardware pass-through path.

[0103] Continue to see Figure 7 The client offloads the CPU-consuming DIF CRC calculation hardware to the NPU for processing, reducing CPU usage and further reducing IO latency.

[0104] The data processing method provided in the embodiment of the present application also provides a processing method for an IO request that is a read request. When the client receives an IO request from an upper-layer application that is a read request, it first determines whether the read request is a 4KB-aligned request. If it is not a 4KB-aligned request, the original read request is subjected to 4KB alignment processing. For example, if the offset of the read request is 11KB and the read request is not a 4KB-aligned request, it is subjected to 4KB alignment processing, forward-aligned, and the offset of the read request is updated to 8KB to achieve 4KB alignment. The 4KB-aligned read request is sent to the server after hardware unloading, i.e., the NPU on the smart network card. The NPU reads 4KB+64 data in DIF format from the hard disk. The NPU returns the read 4KB+64 data to the client. The client performs a DIF check on the returned data to confirm the integrity of the data. If the check passes, the data in the 4KB+64 format is converted into the business format, i.e., the 4KB business data in the 4KB+64 format data is read, and the business data required by the upper-layer application is returned to the upper-layer application (see Figure 7 The reading process on the right side of the diagram).

[0105] Optionally, the client offloads the DIF check hardware to the NPU of the network card for processing. In this way, the client offloads the calculation and verification of DIFCRC that consumes CPU resources to the NPU through hardware, further reducing CPU usage and IO latency.

[0106] Figure 8 A schematic diagram of implementing read and write processing of a data processing method provided by an embodiment of the present application is shown. Figure 8 As shown, the client processes operations that are not suitable for hardware microcode processing in the CPU, such as 4KB alignment processing and DIF information insertion processing. For IO requests that are read requests, non-4KB read requests are converted into 4KB aligned read requests, and then sent to the server that is offloaded to the smart network card NPU by the hardware to request the NPU to read the required data from the hard disk. For IO requests that are write requests, 4KB alignment processing is performed on non-4KB aligned write requests (for example, the page cache data with non-4KB aligned data in the last page of the previous Plog is used to fill the first page, and the tail is aligned with zeros), and then the format without DIF information is unified into 4KB+64 format with DIF information data, and then the 4KB aligned 4KB+64 data to be written is sent to the NPU, and the NPU performs DIF verification on the data. After the verification passes, the data is stored on the disk.

[0107] Fig. 9 The present invention provides an architecture diagram of a data processing device according to an embodiment of the present invention. The data processing device 900 is deployed on the network card of the host and is used to implement the client functions described above. The host is connected to the storage node through the network card. The data processing device 900 includes at least a receiving module 901, a processing module 902 and a sending module 903, wherein the receiving module 901 is used to receive a first write request sent by an upper-layer application, and the first write request carries the data to be written; the processing module 902 is used to process the data to be written to obtain at least one first data block, wherein each first data block includes a data block to be written and a DIF data block, and the DIF data block includes common DIF information; the sending module 903 is used to send a second write request to the server, and the second write request is used to request to store at least one first data block.

[0108] In one possible implementation, the processing module 902 is specifically used to divide the data to be written to obtain at least one data block to be written; determine the common DIF information of each data block to be written in at least one data block to be written, and the common DIF information of each data block to be written includes the CRC of each data block to be written; insert the common DIF information into the DIF data block in each first data block.

[0109] In another possible implementation, the processing module 902 is further configured to insert private DIF information into the DIF data block in each first data block, where the private DIF information includes the Meta of the data block to be written in each first data block and the CRC of the Meta.

[0110] In another possible implementation, the data processing device 900 further includes a 4KB alignment module 904, and the 4KB alignment module 904 is used to determine that the first write request is a non-4KB aligned request; and perform 4KB alignment processing on the first write request.

[0111] In another possible implementation, the 4KB alignment module 904 is specifically used to determine that the header of the first write request is not 4KB aligned, and then use the Page cache data to padded forward the data to be written, and the Page cache data includes the last page data of the last written data; if it is determined that the tail of the first write request is not 4KB aligned, then fill the tail of the data to be written with zeros to align the tail of the data to be written with 4KB.

[0112] In another possible implementation, the 4KB alignment module 904 is further configured to update the page cache data based on the tail data of the data to be written.

[0113] Optionally, the size of the data block to be written is 4 KB, and the size of the DIF data block is 64 bytes.

[0114] In another possible implementation, the data processing device 900 also includes a second response module 905, which is used to respond to the first read request of the upper-layer application and send a second read request to the server; after receiving at least one target data block returned by the server, the sending module 903 is used to send the business data block in each data block of the at least one target data block to the upper-layer application.

[0115] In another possible implementation, the 4KB alignment module 904 is further configured to determine that the first read request is a non-4KB aligned read request; and convert the non-4KB aligned read request into a 4KB aligned read request to obtain a second read request.

[0116] Fig.10 The present invention provides an architecture diagram of a data processing device according to an embodiment of the present invention. The data processing device 1000 is deployed on a network card of a storage node, and is used to implement the server functions described above. The storage node is connected to the host through the network card. The data processing device 1000 includes at least a first response module 1001, which is used to respond to a write request sent by a client, perform a DIF check on at least one first data block, and store at least one first data block in a hard disk if the DIF check passes, wherein each data block in the at least one data block includes a data block to be written and a DIF data block.

[0117] The network card includes an NPU, and a first response module is deployed in the NPU, for responding to a second write request, performing a DIF check on at least one data block, and storing at least one first data block in a hard disk if the DIF check passes.

[0118] In one possible implementation, the first response module 1001 is used to respond to a second read request sent by the client, read at least one target data block from the hard disk, the at least one target data block including a business data block and a DIF data block; send the at least one target data block to the client; and enable the client to send the business data block in each data block in the at least one target data block to an upper-layer application.

[0119] It should be noted that the host and storage node mentioned above can be a physical server or a cloud server (such as a virtual server). Fig.11 FIG. 1 is a schematic diagram of the structure of a cloud service system according to an embodiment of the present application. Fig.11 , the cloud service system 1100 includes a host and a storage node. The host includes a hardware layer and a virtual machine monitor (virtual machine monitor, VMM) running on the hardware layer and running on the hardware layer, and multiple virtual machines (virtualmachine, VM). Any virtual machine can serve as a virtual host of the cloud service system 1100. The storage node, similar to the host, includes a hardware layer and a virtual monitor running on the hardware layer and multiple virtual machines, and any virtual machine can serve as a virtual storage node of the cloud service system 1100. The composition of the host will be described in detail as an example below.

[0120] A virtual machine is a virtual computer (server) simulated on public hardware resources by virtualization software. An operating system and applications can be installed on a virtual machine, and the virtual machine can also access network resources. For applications running in a virtual machine, the virtual machine is like working in a real computer.

[0121] The hardware layer is the hardware platform on which the virtualized environment runs, and can be abstracted from the hardware resources of one or more physical hosts. Among them, the hardware layer can include a variety of hardware, for example, the hardware layer includes a processor (such as CPU / NPU / SPU, etc.), a memory and a network card, and can also include high-speed / low-speed IO devices and other devices with specific processing functions. Among them, the memory can be a volatile memory (volatile memory), such as a random access memory (random-access memory, RAM), a dynamic random access memory (dynamic random-access memory, DRAM); the memory can also be a non-volatile memory (non-volatile memory), such as a read-only memory (read-only memory, ROM), a flash memory (flash memory), a hard disk drive (hard disk drive, HDD), a solid-state drive (solid-state drive, SSD), a storage class memory (storage class memory, SCM), etc.; the memory can also include a combination of the above types of memory. The virtual machine runs an executable program based on the hardware resources provided by the VMM and the hardware layer to execute the method steps performed by the client in the above embodiment, which will not be repeated here for the sake of brevity.

[0122] Through the data processing method provided in the embodiment of the present application, the non-4KB aligned IO request data is aligned to 4KB on the client, and then the data format is unified into a 4KB+64 data format with DIF information, and the private DIF information is inserted before the hardware unloading path. The hardware unloading logic only needs to be processed according to the determined format, which simplifies the hardware microcode development logic and reduces the latency of the IO hardware pass-through path.

[0123] An embodiment of the present application also provides a network unloading device, which is applied to a server device in a storage system, and is used to perform hardware unloading of the storage service on the server device; the network unloading device includes a processor (such as an NPU), and the processor can execute the following steps: receiving a write request sent by a client, the write request carrying at least one data block, each of the at least one data block including a data block to be written and a DIF data block; responding to the write request, performing a DIF check on at least one data block, and if the DIF check passes, storing at least one data block in a hard disk.

[0124] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer instructions are executed by a processor, the above-mentioned method is implemented.

[0125] An embodiment of the present application provides a chip, which includes at least one processor and an interface, wherein the at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the method mentioned above.

[0126] An embodiment of the present application provides a computer program or a computer program product, wherein the computer program or the computer program product comprises instructions, and when the instructions are executed, the computer is caused to execute the above-mentioned method.

[0127] Those of ordinary skill in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0128] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented by hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0129] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A data processing method, characterized in that: Applied to a storage system, the storage system includes a client and a server, the client and the server perform hardware unloading through a hardware unloading device to implement IO hardware direct access to a hard disk, the method includes: The client receives a first write request sent by an upper-layer application, where the first write request carries data to be written; The client processes the data to be written to obtain at least one first data block, wherein each first data block includes a data block to be written and a DIF data block, and the DIF data block includes public DIF information; The client sends a second write request to the server, where the second write request is used to request to store the at least one first data block; The server performs a DIF check on the at least one first data block in response to the second write request, and stores the at least one first data block in the hard disk if the DIF check passes.

2. The method according to claim 1, characterized in that The client processes the data to be written to obtain at least one first data block, including: The client divides the data to be written into at least one data block to be written; The client determines the common DIF information of each data block to be written in the at least one data block to be written, wherein the common DIF information of each data block to be written includes a cyclic redundancy check code of each data block to be written; The client inserts the common DIF information into the DIF data block in each of the first data blocks.

3. The method according to claim 2, characterized in that The client processes the data to be written to obtain at least one first data block, further comprising: The client inserts private DIF information into the DIF data block in each of the first data blocks, where the private DIF information includes metadata of the data block to be written in each of the first data blocks and a cyclic redundancy check code of the metadata.

4. The method according to any one of claims 1 to 3, characterized in that: The client processes the data to be written to obtain at least one first data block, and the method also includes: The client determines that the first write request is a non-4KB aligned request; The client performs 4KB alignment processing on the first write request.

5. The method according to claim 4, characterized in that The client performs 4KB alignment processing on the first write request, including: The client determines that the first write request header is not aligned with 4KB, and then uses page cache data to padded forward the data to be written, wherein the page cache data includes the last page data of the last written data; The client determines that the tail of the first write request is not aligned with 4KB, and then performs a zero padding operation on the tail of the data to be written to align the tail of the data to be written with 4KB.

6. The method according to claim 5, characterized in that The client performs 4KB alignment processing on the first write request, further comprising: The page cache data is updated based on the tail data of the data to be written.

7. The method according to any one of claims 1 to 6, characterized in that: The size of the data block to be written is 4 KB, and the size of the DIF data block is 64 bytes.

8. The method according to any one of claims 1 to 7, characterized in that: The hardware offloading device includes a neural network processing unit; The server performs a DIF check on the at least one first data block in response to the second write request, and stores the at least one first data block in the hard disk if the DIF check passes, including: The neural network processing unit responds to the second write request and performs a DIF check on the at least one first data block, and stores the at least one first data block in the hard disk if the DIF check passes.

9. The method according to any one of claims 1 to 8, characterized in that: Also includes: The client sends a second read request to the server in response to the first read request of the upper layer application; The server reads at least one target data block from the hard disk in response to the second read request sent by the client, wherein the at least one target data block includes a service data block and a DIF data block; sending the at least one target data block to the client; The client sends the service data block in each data block in the at least one target data block to the upper layer application.

10. The method according to claim 9, characterized in that The client responds to the first read request of the upper layer application and sends a second read request to the server, which includes: The client determines that the first read request is a non-4KB aligned read request; The client converts the non-4KB aligned read request into a 4KB aligned read request to obtain the second read request.

11. A data storage system, characterized in that: It includes a client and a server, wherein the client and the server perform hardware unloading through a hardware unloading device to realize IO hardware direct access to a hard disk; The client comprises: A receiving module, configured to receive a first write request sent by an upper layer application, wherein the first write request carries data to be written; A processing module, used for processing the data to be written to obtain at least one first data block, wherein each first data block includes a data block to be written and a DIF data block, and the DIF data block includes common DIF information; A sending module, used for sending a second write request to the server, wherein the second write request is used for requesting to store the at least one first data block; The server includes: The first response module is used to respond to the second write request, perform a DIF check on the at least one first data block, and store the at least one first data block in the hard disk if the DIF check passes.

12. The system according to claim 11, characterized in that The processing module is specifically used for: Dividing the data to be written to obtain at least one data block to be written; Determine common DIF information of each data block to be written in the at least one data block to be written, wherein the common DIF information of each data block to be written includes a cyclic redundancy check code of each data block to be written; The common DIF information is inserted into the DIF data block in each of the first data blocks.

13. The system according to claim 12, characterized in that The processing module is also used for: Insert private DIF information into the DIF data block in each of the first data blocks, the private DIF information including metadata of the data block to be written in each of the first data blocks and a cyclic redundancy check code of the metadata.

14. The system according to any one of claims 11 to 13, characterized in that: The client also includes a 4KB alignment module; The 4KB alignment module is used to: Determining that the first write request is a non-4KB aligned request; Perform 4KB alignment processing on the first write request.

15. The system according to claim 14, characterized in that The 4KB alignment module is specifically used for: Determining that the first write request header is not 4KB aligned, then using page cache data to pad the data to be written forward, the page cache data including the last page data of the last written data; If it is determined that the tail of the first write request is not 4KB aligned, a zero padding operation is performed on the tail of the data to be written so that the tail of the data to be written is 4KB aligned.

16. The system according to claim 15, characterized in that The 4KB alignment module is also used to: The page cache data is updated based on the tail data of the data to be written.

17. The system according to any one of claims 11 to 16, characterized in that: The size of the data block to be written is 4 KB, and the size of the DIF data block is 64 bytes.

18. The system according to any one of claims 11 to 17, characterized in that: The hardware offloading device includes a neural network processing unit; The first response module is deployed in the neural network processing unit, and is used to respond to the second write request, perform a DIF check on the at least one first data block, and store the at least one first data block in the hard disk if the DIF check passes.

19. The system according to any one of claims 11 to 18, characterized in that: The client also includes: A second response module, configured to respond to the first read request of the upper layer application and send a second read request to the server; The first response module is used to respond to the second read request sent by the client and read at least one target data block from the hard disk, wherein the at least one target data block includes a business data block and a DIF data block; sending the at least one target data block to the client; The sending module is used to send the service data block in each data block in the at least one target data block to the upper layer application.

20. The system according to claim 19, characterized in that The 4KB alignment module is also used to: Determining that the first read request is a non-4KB aligned read request; The non-4KB aligned read request is converted into a 4KB aligned read request to obtain the second read request.

21. A network offloading device, characterized in that: The network offloading device is applied to a server device in a storage system, and is used to perform hardware offloading of the storage service client on the server device; The network offloading device comprises a processor, wherein the processor is configured to: Receive a write request sent by a client, wherein the write request carries at least one data block, and each data block in the at least one data block includes a data block to be written and a DIF data block; In response to the write request, a DIF check is performed on the at least one data block, and if the DIF check passes, the at least one data block is stored in the hard disk.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Distributed storage management method, electronic equipment, storage medium and program product

    CN120296065A

  • Data storage method and device, electronic equipment and program product

    CN120447839A