Data processing method and device, distributed storage system, equipment and program product

By aggregating operation requests on the client and batch processing on the storage side, the problem of low network communication efficiency in distributed storage systems is solved, data transmission and processing efficiency is improved, and more efficient data storage and read and write operations are achieved.

CN120353395APending Publication Date: 2025-07-22JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510472134.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In a distributed storage system, the network communication between the client and the storage node is not efficient, resulting in low data reading and writing efficiency, making it difficult to effectively utilize the advantages of the distributed storage system.

Method used

The client aggregates multiple operation requests into a single aggregate request and sends them to the storage node for batch processing. By creating data placement groups on the storage side, it can realize the unique correspondence of data objects and the sending of aggregate requests, improving transmission and processing efficiency.

Benefits of technology

Through batch processing of aggregated requests, the transmission efficiency and processing efficiency between the client and the storage node are improved, and the overall performance of the distributed storage system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353395A_ABST
    Figure CN120353395A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, a distributed storage system, equipment and a program product, relates to the technical field of storage, and can segment a file in a client into a data object and set an object identifier for the data object so as to convert a read-write operation on the file into a read-write operation on the data object. Then, when the data object is operated, the storage node uniquely corresponding to the data object can be determined in the multiple storage nodes of the storage end according to the object identifier of the data object, and the data object is distributed to the storage nodes. When the number of the data objects distributed to the same storage node reaches a preset number, the operation requests corresponding to the data objects in the storage node can be aggregated to obtain an aggregation request, and the aggregation request is sent to the storage node, so that the storage node processes the operation requests in the aggregation request in batches; in this way, a plurality of operation requests can be aggregated into a single aggregation request and issued to the storage node, so that the transmission efficiency and the processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage technology, and in particular to a data processing method, apparatus, distributed storage system, device, and program product. Background Art

[0002] In a distributed storage system, a client can write or read data through a network with storage nodes in a storage side. However, the efficiency of network communication is not high, which in turn leads to low data reading and writing efficiency between the client and the storage nodes, and it is difficult to effectively utilize the advantages of the distributed storage system. Summary of the Invention

[0003] The present application provides a data processing method, apparatus, distributed storage system, device, and program product, which can improve the transmission efficiency and processing efficiency between the client and the storage nodes by aggregating multiple operation requests in the client into a single aggregated request and sending the aggregated request to the storage nodes for batch processing.

[0004] The present application provides a data processing method applied to a client. The method includes:

[0005] Determine a data object to be operated on in a file to be operated on in the client, and determine an object identifier of the data object; wherein, the data object is obtained by splitting the file to be operated on, and the operation corresponding to the data object is a read operation or a write operation;

[0006] Determine a storage node uniquely corresponding to the data object among at least two storage nodes included in the storage side according to the object identifier, and allocate the data object to the storage node;

[0007] When the number of data objects allocated to the same storage node reaches a preset number, aggregate operation requests corresponding to the respective data objects in the storage node to obtain an aggregated request, and send the aggregated request to the storage node so that the storage node batch processes the operation requests in the aggregated request.

[0008] Optionally, determining a storage node uniquely corresponding to the data object among at least two storage nodes included in the storage side according to the object identifier, and allocating the data object to the storage node includes:

[0009] Determine a data placement group uniquely corresponding to the data object among at least two data placement groups according to the object identifier, and add the data object to the data placement group; wherein, one data placement group uniquely corresponds to one storage node;

[0010] When the number of data objects allocated to the same storage node reaches a preset number, aggregating operation requests corresponding to the respective data objects in the storage node to obtain an aggregated request, and sending the aggregated request to the storage node includes:

[0011] When the number of data objects in the same data placement group reaches a preset number, aggregate the operation requests corresponding to each data object in the data placement group to obtain an aggregation request;

[0012] Send the aggregation request to the storage node corresponding to the data placement group.

[0013] Optionally, the operation request is a write operation request, and the aggregation request includes aggregated write data and write control information;

[0014] Aggregating the operation requests corresponding to each data object in the data placement group to obtain an aggregation request includes:

[0015] Aggregate the data blocks corresponding to each data object in the data placement group in the client memory to obtain aggregated write data;

[0016] Generate a write operation request for the data object according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated write data, and the memory address of the data block corresponding to the data object;

[0017] Aggregate the write operation requests corresponding to each data object in the data placement group into write control information;

[0018] Sending the aggregation request to the storage node corresponding to the data placement group includes:

[0019] Send the write control information to the storage node through the control plane, and send the aggregated write data to the storage node through the data plane, so that the storage node performs batch writing on each data block in the aggregated write data according to the write operation requests in the write control information.

[0020] Optionally, the file to be operated on is the file to be stored;

[0021] According to the file to be operated on in the client, determine the data object to be operated on in the file to be operated on, and determine the object identifier of the data object, including:

[0022] Receive the file to be stored written by the upper-layer application to the client memory;

[0023] Split the file to be stored according to the preset data object size to obtain data objects, and set object identifiers for the data objects.

[0024] Optionally, the operation request is a read operation request, and the aggregation request includes read control information;

[0025] Aggregating the operation requests corresponding to each data object in the data placement group to obtain an aggregation request includes:

[0026] Generate a read operation request for a data object based on the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object;

[0027] Aggregate the read operation requests corresponding to the respective data objects in the data placement group into read control information;

[0028] Send the aggregated request to the storage node corresponding to the data placement group, including:

[0029] Send the read control information to the storage node through the control plane, so that the storage node reads the data blocks of the respective data objects according to the read operation requests in the read control information and sequentially aggregates them into aggregated read data;

[0030] After sending the read control information to the storage node through the control plane, it further includes:

[0031] Receive the aggregated read data returned by the storage node through the data plane;

[0032] Extract the data blocks of the respective data objects from the aggregated read data according to the read operation requests in the read control information.

[0033] Optionally, the file to be read is the file to be read;

[0034] Determine the data objects to be operated on in the file to be operated on according to the file to be operated on in the client, and determine the object identifier of the data object, including:

[0035] Receive an asynchronous read request issued by the upper-layer application and determine the file to be read corresponding to the asynchronous read request;

[0036] Determine the data objects included in the file to be read and determine the object identifier of the data object;

[0037] After extracting the data blocks of the respective data objects from the aggregated read data, it further includes:

[0038] Form the data blocks of the respective data objects into the file to be read and return the file to be read to the upper-layer application.

[0039] Optionally, before determining the data objects to be operated on in the file to be operated on according to the file to be operated on in the client and determining the object identifier of the data object, it further includes:

[0040] Create data placement groups according to the number of storage nodes in the storage side and set data placement group identifiers for each data placement group;

[0041] Determine the corresponding relationship between each data placement group and each storage node according to the data placement group identifier.

[0042] The present application also provides a data processing method, which is applied to a storage node in a storage side. The storage side includes at least two storage nodes. The method includes:

[0043] Receiving an aggregation request sent by a client; wherein, the aggregation request includes a preset number of operation requests, and the operation requests are read operation requests or write operation requests for data objects, and the data objects are obtained by splitting a file to be operated in the client;

[0044] Creating a unified transaction for the aggregation request and sequentially locking the operation requests in the aggregation request;

[0045] When the sequential locking of the operation requests is completed, batch-processing the operation requests in the aggregation request.

[0046] Optionally, the aggregation request includes aggregated write data and write control information. The aggregated write data includes each data block of each data object, and the write control information includes each write operation request corresponding to each data object. The write operation request includes the data object name of the data object, the position of the data block corresponding to the data object in the aggregated write data, and the memory address of the data block corresponding to the data object;

[0047] Receiving the aggregation request sent by the client includes:

[0048] Receiving the write control information sent by the client through the control plane and receiving the aggregated write data sent by the client through the data plane;

[0049] Batch-processing the operation requests in the aggregation request includes:

[0050] Batch-writing each data block in the aggregated write data according to each write operation request in the write control information.

[0051] Optionally, the aggregation request includes read control information. The write control information includes each read operation request corresponding to each data object. The read operation request includes the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object;

[0052] Receiving the aggregation request sent by the client includes:

[0053] Receiving the read control information sent by the client through the control plane;

[0054] Batch-processing the operation requests in the aggregation request includes:

[0055] Reading the data blocks of each data object according to each read operation request in the read control information and sequentially aggregating them into aggregated read data;

[0056] Return the aggregated read data to the client through the data plane, so that the client extracts the data blocks of each data object from the aggregated read data according to the read control information.

[0057] Optionally, the storage node includes a replica pool and an erasure code pool;

[0058] Batch process the operation requests in the aggregation request, including:

[0059] Write the data objects corresponding to the write operation requests in the aggregation request into the replica pool;

[0060] Determine whether the size of the data objects written in the replica pool has reached the preset erasure code stripe size;

[0061] If so, convert the data objects written in the replica pool into erasure code stripes and write the erasure code stripes into the erasure code pool;

[0062] If not, retain the data objects written in the replica pool in the replica pool.

[0063] Optionally, batch process the operation requests in the aggregation request, including:

[0064] Determine the storage locations of the data objects corresponding to the read operation requests in the aggregation request;

[0065] If it is determined according to the storage locations that all data objects are in the replica pool, batch read the data blocks of the data objects from the replica pool and return them to the client;

[0066] If it is determined according to the storage locations that all data objects are in the erasure code pool, batch read the data blocks of the data objects from the erasure code pool and return them to the client;

[0067] If it is determined according to the storage locations that the data objects are in both the replica pool and the erasure code pool, merge the data blocks of the data objects in the replica pool with the data blocks of the data objects in the erasure code pool, and return the merged data blocks of the data objects to the client.

[0068] Optionally, after returning the merged data blocks of the data objects to the client, it further includes:

[0069] Use the merged data blocks of the data objects to update the corresponding erasure code stripes in the erasure code pool.

[0070] Optionally, it further includes:

[0071] Statistically analyze the access frequencies of each data object and determine the hot data objects according to the access frequencies;

[0072] Copy the hot data objects from the erasure code pool to the replica pool.

[0073] The present application also provides a data processing device, which is applied to a client. The device includes:

[0074] A data object determination module, configured to determine a data object to be operated in a file to be operated in the client, and determine an object identifier of the data object; wherein, the data object is obtained by splitting the file to be operated, and the operation corresponding to the data object is a read operation or a write operation;

[0075] An allocation module, configured to determine, according to the object identifier, a storage node uniquely corresponding to the data object among at least two storage nodes included in the storage end, and allocate the data object to the storage node;

[0076] A request aggregation and distribution module, configured to, when the number of data objects allocated to the same storage node reaches a preset number, aggregate the operation requests corresponding to the respective data objects in the storage node to obtain an aggregation request, and send the aggregation request to the storage node, so that the storage node batch processes the operation requests in the aggregation request.

[0077] The present application also provides a data processing device, which is applied to a storage node in the storage end. The storage end includes at least two storage nodes. The device includes:

[0078] A request receiving module, configured to receive an aggregation request sent by the client; wherein, the aggregation request includes a preset number of operation requests, and the operation request is a read operation request or a write operation request for a data object, and the data object is obtained by splitting a file to be operated in the client;

[0079] A locking module, configured to create a unified transaction for the aggregation request, and sequentially lock the operation requests in the aggregation request;

[0080] A request batch processing module, configured to batch process the operation requests in the aggregation request when the sequential locking of the operation requests is completed.

[0081] The present application also provides a distributed storage system, including: a client and a storage end, and the storage end includes at least two storage nodes;

[0082] The client is configured to execute the above-mentioned data processing method applied to the client;

[0083] The storage node is configured to execute the above-mentioned data processing method applied to the storage node.

[0084] The present application also provides an electronic device, including:

[0085] A memory, configured to store a computer program;

[0086] A processor, configured to implement the above-mentioned data processing method when executing the computer program.

[0087] The present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the above-mentioned data processing method.

[0088] The present application also provides a non-volatile computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the above-mentioned data processing method is implemented.

[0089] Through the present application, a file in a client can be segmented into data objects, and an object identifier can be set for the data objects, so as to convert the read and write operations on the file into the read and write operations on the data objects. Subsequently, when operating on the data objects, the storage node uniquely corresponding to the data objects can be determined from multiple storage nodes at the storage end according to the object identifiers of the data objects, and the data objects can be allocated to the storage nodes. When the number of data objects allocated to the same storage node reaches a preset number, the operation requests corresponding to the respective data objects in the storage node can be aggregated to obtain an aggregated request, and the aggregated request can be sent to the storage node, so that the storage node batch processes the operation requests in the aggregated request; in this way, multiple operation requests can be aggregated into a single aggregated request and sent to the storage node, thereby improving the transmission efficiency and processing efficiency. The present application also provides a data processing device, a distributed storage system, an electronic device, a computer program product, and a computer-readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0091] Figure 1 It is a structural block diagram of a distributed storage system provided by an embodiment of the present application;

[0092] Figure 2 It is a flowchart of a data processing method provided by an embodiment of the present application;

[0093] Figure 3 It is a schematic diagram of request aggregation provided by an embodiment of the present application;

[0094] Figure 4 It is a schematic diagram of aggregated write data provided by an embodiment of the present application;

[0095] Figure 5 It is a schematic diagram of control information provided by an embodiment of the present application;

[0096] Figure 6 It is a flowchart of another data processing method provided by an embodiment of the present application;

[0097] Figure 7 It is a structural block diagram of a data processing device provided by an embodiment of the present application;

[0098] Figure 8 It is a structural block diagram of another data processing device provided by an embodiment of the present application;

[0099] Figure 9 It is a structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0100] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0101] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0102] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0103] In a distributed storage system, a client can write or read data from storage nodes in the storage end through a network. For example, data can be transmitted between the client and the storage nodes through Remote Direct Memory Access (RDMA). Remote Direct Memory Access is a network-based direct memory access technology that can directly transfer data from the memory of one computer to another computer. However, the efficiency of network communication is not high. For example, network communication often involves connection operations between communication parties and encapsulation and decapsulation operations of network data packets, and these operations are likely to affect the efficiency of data transmission and data processing, thereby reducing the transmission efficiency and processing efficiency between the client and the storage nodes and making it difficult to effectively exert the advantages of the distributed storage system.

[0104] In view of this, in response to the technical problem of how to improve the transmission efficiency and processing efficiency between the client and the storage node, the present application can provide a data processing method, which can aggregate multiple read operation requests or multiple write operation requests into a single aggregated request at the client, and can send the aggregated request to the storage side, so that the storage node can batch process the multiple operation requests in the aggregated request, thereby improving the transmission efficiency and processing efficiency between the client and the storage node.

[0105] For ease of understanding, the distributed storage system applicable to the present application will be introduced first below. Please refer to Figure 1 , Figure 1 FIG. 9 is a structural block diagram of a distributed storage system provided by an embodiment of the present application. This system may include a client 1 and a storage side 2. The storage side 2 includes at least two storage nodes 20. The client 1 can communicate with the storage nodes 20 through a network. Of course, to prevent the storage nodes 20 from being directly exposed to the outside world, the storage side 2 may also include a master node (not shown). The master node is arranged between the client 1 and each storage node 20, and it can communicate with the client 1 through a network and can also communicate with each storage node 20. The master node is mainly responsible for forwarding data between the client 1 and the storage nodes 20, such as receiving a request sent by the client 1 and forwarding the request to the corresponding storage node 20 for processing; and receiving the processing result of the storage node 20 and returning the processing result to the client 1. It should be noted that the present application does not limit the specific number of the storage nodes 20, which can be set according to actual application requirements. In addition, the present application does not limit the specific number of the clients 1 either. It can be one or multiple, that is, the storage side 2 can serve multiple clients 1 at the same time.

[0106] Based on the above introduction of the system structure, the data processing method provided by the present application will be introduced below. First, the implementation of the method on the client side will be introduced. Please refer to Figure 2 , Figure 2 FIG. 16 is a flowchart of a data processing method provided by an embodiment of the present application. This method can be applied to the client and specifically includes:

[0107] S101. Determine the data object to be operated on in the file to be operated on in the client, and determine the object identifier of the data object; wherein, the data object is obtained by splitting the file to be operated on, and the operation corresponding to the data object is a read operation or a write operation.

[0108] In this embodiment, the file to be operated on can be a file to be stored written by an upper-layer application in the client memory to the client memory, or a file to be read requested by the upper-layer application from the client memory. In the file storage scenario, when the client receives the file to be stored, it needs to store the file to be stored from the client memory to the storage end. In the file reading scenario, the client needs to extract the file to be read from the storage node in the storage end to the client memory and return it to the upper-layer application. Since the upper-layer application has strong randomness in reading and writing the client memory, it is easy to cause frequent transfer of small pieces of data between the client and the storage end, thereby reducing the transmission efficiency and processing efficiency between the two ends.

[0109] For this reason, the client in this embodiment can reorganize and aggregate the file to be operated on to increase the amount of requests and data transferred each time between the client and the storage end, thereby improving the transmission efficiency and processing efficiency. Specifically, the client can pre-chunk the file to be operated on to obtain data objects and set an object identifier for each data object. Among them, the data object is the object for which read and write operations are performed in this embodiment, and can be simply understood as a data block in the file to be operated on. In other words, a file to be operated on can be composed of one or more data objects. The object identifier is the unique identifier of the data object, used to index the data object and map the data object to the corresponding storage node.

[0110] In this way, this embodiment can convert the read or write operation on the file to be operated on into a read or write operation on the data object, and can flexibly perform data reorganization and request aggregation based on the data object, thereby increasing the amount of requests and data transferred each time between the client and the storage end, and thus improving the transmission efficiency and processing efficiency.

[0111] It should be noted that the "data object to be operated on" in this step is not necessarily all the data objects in the file to be operated on, but can be some of the data objects in the file to be operated on. For example, when the upper-layer application modifies a file in the client memory, it may only modify some of the data of the file. At this time, the client only needs to re-transmit the data objects corresponding to the modified part of the data to the storage end for storage.

[0112] Furthermore, it is worth pointing out that when the upper-layer application successfully writes data to the client memory, the client memory can directly feedback to the upper-layer application that the write is successful without waiting for the data to be further written to the storage end. Therefore, for the upper-layer application, the write efficiency can be improved and the write latency can be reduced. When the upper-layer application reads data from the client memory, since the client memory may need to read data from the storage end, the upper-layer application can initiate a read request to the client memory in an asynchronous manner, and the client can wake up the upper-layer application and return the data when it receives the data returned from the storage end.

[0113] S102. Determine the storage node uniquely corresponding to the data object among at least two storage nodes included in the storage side according to the object identifier, and allocate the data object to the storage node.

[0114] In this step, the data object to be operated will be allocated to each storage node according to the object identifier of the data object, so as to perform data reorganization and request aggregation on the data objects allocated to each storage node, and issue the aggregated request to the corresponding storage node for batch processing. Specifically, to ensure data consistency, this embodiment needs to determine the storage node uniquely corresponding to the data object according to the object identifier, so as to operate the data object only on this storage node. For example, the object identifier can be subjected to consistent hashing processing to obtain the storage node uniquely corresponding to the data object.

[0115] Furthermore, for the convenience of data reorganization and request aggregation, this embodiment can also provide a data placement group (DPG, Data Packaging Group) to perform data packaging based on the data placement group. Among them, the data placement group is an intermediate layer for consistent hashing calculation of business objects. Specifically, at least two data placement groups can be set in the client, and the corresponding relationship between each data placement group and each storage node can be established in advance to ensure that one data placement group uniquely corresponds to one storage node. Subsequently, according to the object identifier of the data object, determine the data placement group uniquely corresponding to the data object, and add the data object to the data placement group. For example, the object identifier can be subjected to hashing calculation to determine the data placement group number corresponding to the data object, so as to add the data object to the data placement group.

[0116] Based on this, determining the storage node uniquely corresponding to the data object among at least two storage nodes included in the storage side according to the object identifier and allocating the data object to the storage node may include:

[0117] Step 11: Determine the data placement group uniquely corresponding to the data object among at least two data placement groups according to the object identifier, and add the data object to the data placement group; where one data placement group uniquely corresponds to one storage node.

[0118] Furthermore, this embodiment can perform data reorganization and request aggregation on the data objects added to the same data placement group to improve processing efficiency.

[0119] It should be noted that the number of data placement groups in this embodiment is not limited and can be set according to actual application requirements. To effectively ensure that one data placement group can uniquely correspond to one storage node and each storage node can be allocated the corresponding data placement group, this embodiment can create data placement groups according to the number of storage nodes.

[0120] Based on this, before determining the data object to be operated in the file to be operated in the client and determining the object identifier of the data object, it may further include:

[0121] Step 21: Create a data placement group according to the number of storage nodes in the storage end, and set a data placement group identifier for each data placement group.

[0122] Step 22: Determine the corresponding relationship between each data placement group and each storage node according to the data placement group identifier.

[0123] In the embodiment, the data placement group identifier is the unique identifier of the data placement group. In this embodiment, the corresponding relationship between the data placement group and the storage node can be determined according to the data placement group identifier. For example, a consistent hashing calculation can be performed on the data placement group identifier to determine the storage node number corresponding to the data placement group, so as to determine the corresponding relationship between the two. It should be noted that one data placement group needs to uniquely correspond to one storage node, but one storage node can correspond to multiple data placement groups.

[0124] S103. When the number of data objects allocated to the same storage node reaches a preset number, aggregate the operation requests corresponding to each data object in the storage node to obtain an aggregation request, and send the aggregation request to the storage node, so that the storage node batch processes the operation requests in the aggregation request.

[0125] In this step, to increase the amount of requests and data transferred in a single transmission between the client and the storage end, it can be determined whether the number of data objects allocated to the same storage node reaches the preset number. If not, the allocation can continue. If it has reached, the operation requests (such as write operation requests or read operation requests) corresponding to each data object allocated to the storage node can be aggregated to obtain an aggregation request, and the aggregation request is sent to the storage node. At this time, since the aggregation request contains multiple operation requests, this embodiment can ensure that the client simultaneously issues multiple operation requests to the storage node to improve the transmission efficiency; this embodiment can also ensure that the storage node can batch process multiple operation requests, thereby improving the processing efficiency.

[0126] Further, when a data placement group is set, it can be determined whether the number of data objects in the same data placement group reaches the preset number. If not, the data object can continue to be added to this data placement group. If it has reached, the operation requests corresponding to each data object in the data placement group can be aggregated to obtain an aggregation request, and the aggregation request is sent to the storage node corresponding to the data placement group.

[0127] Based on this, when the number of data objects assigned to the same storage node reaches a preset number, aggregating the operation requests corresponding to each data object in the storage node to obtain an aggregation request, and sending the aggregation request to the storage node may include:

[0128] Step 31: When the number of data objects in the same data placement group reaches the preset number, aggregate the operation requests corresponding to each data object in the data placement group to obtain an aggregation request.

[0129] Step 32: Send the aggregation request to the storage node corresponding to the data placement group.

[0130] It should be noted that for the convenience of aggregating request forwarding, the client can generate a device vector for the aggregation request according to the corresponding relationship between the data placement group and the storage node, and this device vector is used to record the storage node for processing this aggregation request. Subsequently, the client can send both the aggregation request and the device vector to the storage side, and the storage side will send the aggregation request to the corresponding storage node for processing according to the device vector.

[0131] For easy understanding, please refer to Figure 3 , Figure 3 , which is a schematic diagram of a request aggregation provided by an embodiment of the present application. Each "4K" in the client memory represents a data block of a data object. The data objects can be taken out in sequence according to the order of each data object in the client memory, and mapped to different data placement groups according to the object identifier (ino.index) of the data object. Subsequently, the operation requests of each data object in a data placement group can be merged into an aggregation request and sent to the corresponding storage node in the storage side at one time, and the storage node will batch process the aggregation request.

[0132] Based on the above embodiments, the present application can split the file in the client into data objects, and set object identifiers for the data objects to convert the read and write operations on the file into the read and write operations on the data objects. Subsequently, when operating on the data object, the storage node uniquely corresponding to the data object can be determined among multiple storage nodes in the storage side according to the object identifier of the data object, and the data object is assigned to the storage node. When the number of data objects assigned to the same storage node reaches the preset number, the operation requests corresponding to each data object in the storage node can be aggregated to obtain an aggregation request, and the aggregation request is sent to the storage node, so that the storage node batch processes the operation requests in the aggregation request; in this way, multiple operation requests can be aggregated into a single aggregation request and sent to the storage node, thereby improving the transmission efficiency and processing efficiency.

[0133] Based on the above embodiments, in order to improve the data transmission and processing efficiency, a numerical control separation structure can be set between the client and the storage end. Among them, the numerical control separation structure can include a data plane and a control plane. The data plane is used to transmit data to be stored or data to be read, and the control plane is used to transmit write operation requests or read operation requests to achieve separated transmission. For the data plane, considering that the data to be stored and the data to be read have a large volume, if the transmission of these data needs to pass through the client processor and the storage end processor, it is easy to reduce the transmission efficiency. Therefore, the data plane can be set between the client and the storage end based on the Remote Direct Memory Access (RDMA) technology, that is, the data to be stored or the data to be read is directly transferred between the client memory and the storage end memory through the RDMA protocol to improve the transmission efficiency. For the control plane, considering that the write operation requests or read operation requests have a small volume, the client and the storage end can establish the control plane based on the conventional network communication method (such as TCP / IP) to transmit the write operation requests or read operation requests through the client processor and the storage end processor. In order to further improve the transmission efficiency and processing efficiency in the data plane and the control plane, in this embodiment, data reorganization can be performed in the data plane and request aggregation can be performed in the control plane. First, based on the file storage scenario, an implementation of the client performing data reorganization in the data plane and request aggregation in the control plane will be introduced below.

[0134] Based on this, this method may further include:

[0135] S201. Receive the file to be stored written by the upper-layer application into the client memory.

[0136] S202. Split the file to be stored according to the preset data object size to obtain data objects, and set object identifiers for the data objects.

[0137] In steps S201 and S202, after the upper-layer application writes the file to be stored into the client memory, the client can split the file to be stored according to the preset data object size to obtain data objects, and set object identifiers for the data objects.

[0138] It should be noted that the specific value of the preset data object size is not limited in this embodiment and can be set according to actual application requirements. For example, it can be 4M.

[0139] S203. According to the object identifier, determine the data placement group uniquely corresponding to the data object in at least two data placement groups, and add the data object to the data placement group; wherein, one data placement group uniquely corresponds to one storage node.

[0140] S204. When the number of data objects in the same data placement group reaches the preset number, aggregate the data blocks corresponding to the data objects in the data placement group in the client memory to obtain aggregated write data.

[0141] When writing a file, the aggregation request may first include aggregated write data, which is used to store data blocks of each data object to be written to the storage node.

[0142] In this step, when the number of data objects in the same data placement group reaches a preset number, the data blocks corresponding to the data objects in the data placement group in the client memory can be aggregated to obtain the aggregated write data. In other words, the aggregated write data contains a preset number of data blocks.

[0143] S205. Generate a write operation request for the data object according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated write data, and the memory address of the data block corresponding to the data object.

[0144] S206. Aggregate the write operation requests corresponding to the data objects in the data placement group into write control information.

[0145] When writing a file, the aggregation request may further include write control information, which is used to store the write operation requests corresponding to the data objects.

[0146] In steps S205 and S206, after the generation of the aggregated write data is completed, a write operation request for the data object can be generated according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated write data, and the memory address of the data block corresponding to the data object. Among them, the data object name is used to distinguish different data objects, and the object identifier can be used as the data object name; the position of the data block in the aggregated write data is used to extract the data block from the aggregated write data; the memory address is used to retrieve the data block from the memory. After obtaining the write operation requests for each data object, in this embodiment, the write operation requests corresponding to the data objects in the data placement group can be aggregated into write control information.

[0147] For easy understanding, please refer to Figure 4 、 Figure 5 , Figure 4 FIG. is a schematic diagram of aggregated write data provided by an embodiment of the present application, where "4K" and "8K" both represent data blocks. Figure 5 FIG. is a schematic diagram of control information provided by an embodiment of the present application, where each square can represent a write operation request. A possible form of the write operation request is:

[0148] Data object name: 0x1000000000000000.00001234;

[0149] Data position: offset (offset value), len (data length);

[0150] RDMA memory address: client memory address.

[0151] S207. Send the write control information to the storage node through the control plane, and send the aggregated write data to the storage node through the data plane, so that the storage node performs batch writing on each data block in the aggregated write data according to each write operation request in the write control information.

[0152] In this step, after obtaining the write control information and the aggregated write data, the write control information can be sent to the storage node through the control plane, and the aggregated write data can be sent to the storage node through the data plane to achieve the separate transmission of the write control information and the aggregated write data, thereby improving the processing efficiency. The storage node can directly use the write control information to perform batch writing on each data block in the aggregated write data to improve the writing efficiency.

[0153] Based on the above embodiments, the following will introduce another implementation situation of data reorganization in the data plane and request aggregation in the control plane based on the file reading scenario. Based on this, the method may further include:

[0154] S301. Receive an asynchronous read request sent by the upper-layer application, and determine the file to be read corresponding to the asynchronous read request.

[0155] S302. Determine the data objects included in the file to be read, and determine the object identifier of the data object.

[0156] In steps S301 and S302, when an asynchronous read request is sent by the upper-layer application, the client can query the file to be read corresponding to the asynchronous read request, and determine the data objects included in the file to be read and the object identifier of the data object.

[0157] S303. According to the object identifier, determine the data placement group uniquely corresponding to the data object in at least two data placement groups, and add the data object to the data placement group; where one data placement group uniquely corresponds to one storage node.

[0158] S304. When the number of data objects in the same data placement group reaches the preset number, generate a read operation request for the data object according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object.

[0159] S305. Aggregate the read operation requests corresponding to the respective data objects in the data placement group into read control information.

[0160] When writing a file, the aggregation request may include read control information, which is used to store the read operation requests corresponding to the data objects. When the read control information is sent to the storage node, the storage node can read the data blocks corresponding to each data object according to the read control information, and aggregate these data blocks into aggregated read data according to the read control information, and return the aggregated read data to the client. In other words, the aggregated read data contains the data blocks corresponding to each data object to be read.

[0161] In steps S305 and S306, a write operation request for a data object can be generated according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object. Among them, the data object name is used to distinguish different data objects, and the object identifier can be used as the data object name; the position of the data block in the aggregated read data is used to preset the position of each data block in the aggregated read data for subsequent extraction; the memory address is used to retrieve the data block from the memory. After obtaining the read operation requests corresponding to each data object in the data placement group, in this embodiment, the read operation requests corresponding to each data object in the data placement group can be aggregated into read control information.

[0162] It should be noted that the form of the read control information is similar to that of the write control information, so the form of the read control information can also be referred to Figure 5 .

[0163] S306. Send the read control information to the storage node through the control plane, so that the storage node reads the data blocks of each data object according to the read operation requests in the read control information and sequentially aggregates them into aggregated read data control information.

[0164] S307. Receive the aggregated read data returned by the storage node through the data plane.

[0165] In steps S306 and S307, the read control information can be sent to the storage node through the control plane for the storage node to perform batch reading, and the aggregated read data returned by the storage node can also be received through the data plane to realize the separate transmission of the read control information and the aggregated read data, thereby improving the processing efficiency.

[0166] It should be noted that the form of the aggregated read data is similar to that of the aggregated write data, so the form of the aggregated read data can also be referred to Figure 4 .

[0167] S308. Extract the data blocks of each data object from the aggregated read data according to the read operation requests in the read control information.

[0168] S309. Compose the data blocks of each data object into a file to be read and return the file to be read to the upper-layer application.

[0169] In steps S308 and S309, since the read control information records the positions and memory addresses of the data blocks of each data object in the aggregated read data, the client can extract the data blocks of each data object from the aggregated read data according to the read control information in the client memory and reorganize them into the file to be read. Subsequently, the upper-layer application can be awakened, and the file to be read can be returned to the upper-layer application.

[0170] The implementation of this method on the storage side will be introduced below. Please refer to Figure 6 , Figure 6 FIG. is a flowchart of another data processing method provided by an embodiment of the present application. This method can be applied to storage nodes on the storage side, and the storage side includes at least two storage nodes. This method specifically includes:

[0171] S401. Receive an aggregation request sent by the client; wherein, the aggregation request includes a preset number of operation requests, and the operation request is a read operation request or a write operation request for a data object, and the data object is obtained by splitting the file to be operated in the client.

[0172] As described above, to increase the amount of requests and data transferred in a single time between the client and the storage side, the client can pre-divide the file to be operated to obtain data objects, and set an object identifier for each data object. Subsequently, when processing the file to be operated, the client can allocate the data objects to be operated to each storage node according to the object identifier, and when the number of data objects allocated to the same storage node reaches the preset number, aggregate the operation requests corresponding to the data objects allocated to the storage node to obtain an aggregation request, and send the aggregation request to the storage node. Therefore, in this step, the aggregation request received by the storage node can include a preset number of operation requests, and the storage node can perform batch operations on these operation requests.

[0173] S402. Create a unified transaction for the aggregation request and sequentially lock the operation requests in the aggregation request.

[0174] In this step, to perform batch processing on the operation requests, the storage node needs to create a unified transaction for the aggregation request, and then sequentially lock the operation requests in the aggregation request to ensure the effectiveness of the operations.

[0175] It should be noted that this embodiment does not limit the creation method of the transaction and the locking method of the request, and relevant storage technologies can be referred to.

[0176] S403. When the sequential locking of the operation requests is completed, perform batch processing on the operation requests in the aggregation request.

[0177] In this step, after locking each operation request, the storage node can batch process the operation requests in the aggregation request. For example, when the operation request is a write operation request, the storage node can batch write the data blocks of each data object; when the operation request is a read operation request, the storage node can batch read the data blocks of each data object. In this way, the processing efficiency of the storage node can be effectively improved.

[0178] Furthermore, to improve data transmission and processing efficiency, a numerical control separation structure can be set between the client and the storage side. The numerical control separation structure can include a data plane and a control plane. The data plane is used to transmit data to be stored or data to be read, and the control plane is used to transmit write operation requests or read operation requests to achieve separated transmission. In this embodiment, to further improve the transmission efficiency and processing efficiency in the data plane and the control plane, data reorganization can be performed in the data plane and request aggregation can be performed in the control plane. First, based on the file storage scenario, the working conditions of the storage node in the numerical control separation scenario are introduced as follows:

[0179] Based on this, the aggregation request includes aggregated write data and write control information. The aggregated write data includes the data blocks of each data object, and the write control information includes the write operation requests corresponding to each data object. The write operation request includes the data object name of the data object, the position of the data block corresponding to the data object in the aggregated write data, and the memory address of the data block corresponding to the data object;

[0180] Receiving the aggregation request sent by the client may include:

[0181] Step 41: Receive the write control information sent by the client through the control plane and receive the aggregated write data sent by the client through the data plane.

[0182] In this step, the client can send the write control information to the storage node through the control plane and send the aggregated write data to the storage node through the data plane to achieve separated transmission of the write control information and the aggregated write data, thereby improving the processing efficiency.

[0183] Batch processing of the operation requests in the aggregation request may include:

[0184] Step 51: According to the write operation requests in the write control information, batch write the data blocks in the aggregated write data.

[0185] In this step, since the write control information contains the write operation requests of each data object, and each write operation request details the position of the data block corresponding to the data object in the aggregated write data and the memory address of the data block corresponding to the data object, the storage node can directly use the write control information to batch write the data blocks in the aggregated write data to improve the write efficiency.

[0186] Next, continuing with the file reading scenario, the working conditions of the storage node in the numerical control separation scenario are introduced as follows:

[0187] Based on this, the aggregation request includes read control information, and the write control information includes each read operation request corresponding to each data object. The read operation request includes the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object;

[0188] Receiving the aggregation request sent by the client may include:

[0189] Step 61: Receive the read control information sent by the client through the control plane.

[0190] Batch processing of the operation requests in the aggregation request may include:

[0191] Step 71: According to each read operation request in the read control information, read the data blocks of each data object and sequentially aggregate them into aggregated read data;

[0192] Step 72: Return the aggregated read data to the client through the data plane, so that the client extracts the data blocks of each data object from the aggregated read data according to the read control information.

[0193] In steps 61, 71, and 72, the storage node can first receive the read control information sent by the client through the control plane. Subsequently, since the read control information contains the read operation requests for each data object, and each read operation request details the position of the data block corresponding to the data object in the aggregated read data and the memory address of the data block corresponding to the data object, the storage node can directly read the data blocks of each data object according to the read control information and aggregate these data blocks into aggregated read data according to the read control information. Subsequently, the storage node can return the aggregated read data to the client through the data plane, so that the client extracts the data blocks of each data object from the aggregated read data according to the read control information.

[0194] Based on the above embodiments, to ensure the security of data storage, when the storage node stores a data object, it is necessary to perform erasure processing on the data block of the data object to convert the data block into an erasure stripe, so as to ensure the validity of the data based on the parity blocks in the erasure stripe. However, erasure processing is prone to cause write amplification problems, which in turn affect the processing efficiency in the storage node. For this reason, this embodiment can also perform hierarchical processing on the storage space in the storage node to avoid write amplification problems through hierarchical processing. The internal processing process of the storage node is introduced below. Based on this, the storage node includes a replica pool and an erasure pool; batch processing of the operation requests in the aggregation request may include:

[0195] S501. Write the data objects corresponding to the write operation requests in the aggregation request into the replica pool.

[0196] S502. Determine whether the size of the data objects already written in the replica pool reaches the preset erasure stripe size.

[0197] S503. If so, convert the data objects already written in the replica pool into erasure stripes and write the erasure stripes into the erasure pool.

[0198] S504. If not, retain the data objects already written in the replica pool in the replica pool.

[0199] In this embodiment, the replica pool can first directly write the data blocks of the data objects without performing erasure processing. In addition, the replica pool is also used to accumulate data objects. When it is determined that the size of the data objects already written in the replica pool reaches the preset erasure stripe size, the data objects already written in the replica pool can be converted into erasure stripes and the erasure stripes can be written into the erasure pool. In other words, this embodiment can introduce a replica pool in the storage node, give priority to writing small pieces of data into the replica pool, and perform erasure processing when the small pieces of data accumulate into large pieces of data, which can avoid write amplification in the erasure pool when writing and modifying small pieces of data.

[0200] It should be noted that this embodiment does not limit the specific value of the preset erasure stripe size, and relevant erasure technologies can be referred to.

[0201] In addition, factors such as the replica pool water level, access volume, and storage - side pressure can also be comprehensively considered to control the data in the replica pool to be flushed down to the erasure pool.

[0202] Furthermore, when reading data, considering that the data objects may be stored in the replica pool, the erasure pool, or both the replica pool and the erasure pool, it is necessary to perform differential processing according to the storage location of the data objects. The following introduces the specific process of the storage node processing data reading.

[0203] Based on this, batch - processing the operation requests in the aggregation request can include:

[0204] S601. Determine the storage locations of the data objects corresponding to the read operation requests in the aggregation request.

[0205] S602. If it is determined according to the storage location that all data objects are located in the replica pool, batch - read the data blocks of the data objects from the replica pool and return them to the client.

[0206] S603. If it is determined according to the storage location that all data objects are located in the erasure pool, batch - read the data blocks of the data objects from the erasure pool and return them to the client.

[0207] S604. If it is determined according to the storage location that the data object is located in both the replica pool and the erasure code pool at the same time, then merge the data blocks of the data object in the replica pool with the data blocks of the data object in the erasure code pool, and return the merged data blocks of the data object to the client.

[0208] In steps S601 - S604, if the data object is completely stored in the replica pool or completely stored in the erasure code pool, it can be directly read out and returned from the replica pool or the erasure code pool. If the data object is partially stored in the replica pool and partially stored in the erasure code pool, it is necessary to merge the data blocks of the data object in the replica pool with the data blocks of the data object in the erasure code pool, and return the merged data blocks of the data object to the client.

[0209] Of course, after completing the data merging, in addition to returning the merged data blocks to the client, the merged data blocks can also be asynchronously flushed to the erasure code pool.

[0210] Based on this, after returning the merged data blocks of the data object to the client, it may further include:

[0211] Step 81: Update the corresponding erasure code stripe in the erasure code pool by using the merged data blocks of the data object.

[0212] Furthermore, this embodiment can also automatically migrate hot data from the erasure code pool to the replica pool and migrate cold data from the replica pool to the erasure code pool according to the access frequency of the data object.

[0213] Based on this, this method may further include:

[0214] Step 91: Count the access frequency of each data object and determine the hot data objects according to the access frequency;

[0215] Step 92: Copy the hot data objects from the erasure code pool to the replica pool.

[0216] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general - purpose hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0217] Next, the data processing device, electronic device, computer program product, and computer - readable storage medium provided by the embodiments of the present application will be introduced.

[0218] Please refer to Figure 7 , Figure 7 which is the structural block diagram of a data processing device provided by an embodiment of the present application. The device is applied to the client and may include:

[0219] A data object determination module 701, configured to determine a data object to be operated on in a file to be operated on in a client, and determine an object identifier of the data object; wherein, the data object is obtained by splitting the file to be operated on, and the operation corresponding to the data object is a read operation or a write operation.

[0220] An allocation module 702, configured to determine a storage node uniquely corresponding to the data object from at least two storage nodes included in a storage end according to the object identifier, and allocate the data object to the storage node.

[0221] A request aggregation and distribution module 703, configured to, when the number of data objects allocated to the same storage node reaches a preset number, aggregate operation requests corresponding to the data objects in the storage node to obtain an aggregation request, and send the aggregation request to the storage node, so that the storage node batch processes the operation requests in the aggregation request.

[0222] Optionally, the allocation module 702 may include:

[0223] A placement group allocation sub-module, configured to determine a data placement group uniquely corresponding to the data object from at least two data placement groups according to the object identifier, and add the data object to the data placement group; wherein, one data placement group uniquely corresponds to one storage node.

[0224] The request aggregation and distribution module 703 includes:

[0225] An aggregation sub-module, configured to, when the number of data objects in the same data placement group reaches a preset number, aggregate operation requests corresponding to the data objects in the data placement group to obtain an aggregation request.

[0226] A distribution sub-module, configured to send the aggregation request to the storage node corresponding to the data placement group.

[0227] Optionally, the operation request is a write operation request, and the aggregation request includes aggregated write data and write control information; the aggregation sub-module includes:

[0228] A data aggregation unit, configured to aggregate data blocks corresponding to the data objects in the client memory in the data placement group to obtain aggregated write data.

[0229] A write request generation unit, configured to generate a write operation request for the data object according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated write data, and the memory address of the data block corresponding to the data object.

[0230] A write request aggregation unit, configured to aggregate the write operation requests corresponding to the data objects in the data placement group into write control information.

[0231] The distribution sub-module includes:

[0232] A write request sending unit, configured to send write control information to a storage node through a control plane, and send aggregated write data to the storage node through a data plane, so that the storage node performs batch writing on each data block in the aggregated write data according to each write operation request in the write control information.

[0233] Optionally, the file to be operated on is a file to be stored; the data object determination module 701 includes:

[0234] A file receiving sub-module, configured to receive the file to be stored written by an upper-layer application into the client memory;

[0235] A file splitting sub-module, configured to split the file to be stored according to a preset data object size to obtain data objects, and set object identifiers for the data objects.

[0236] Optionally, the operation request is a read operation request, and the aggregation request includes read control information; the aggregation sub-module includes:

[0237] A read request generation unit, configured to generate a read operation request for a data object according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object;

[0238] A read request aggregation unit, configured to aggregate each read operation request corresponding to each data object in a data placement group into read control information;

[0239] The sending sub-module includes:

[0240] A read request sending unit, configured to send the read control information to a storage node through a control plane, so that the storage node reads the data blocks of each data object according to each read operation request in the read control information and sequentially aggregates them into aggregated read data;

[0241] The apparatus may further include:

[0242] A data receiving module, configured to receive the aggregated read data returned by the storage node through a data plane;

[0243] A data extraction module, configured to extract the data blocks of each data object from the aggregated read data according to each read operation request in the read control information.

[0244] Optionally, the file to be read is a file to be read;

[0245] The data object determination module 701 includes:

[0246] A request receiving sub-module, configured to receive an asynchronous read request sent by an upper-layer application and determine the file to be read corresponding to the asynchronous read request;

[0247] An object determination sub-module, configured to determine data objects included in a file to be read and determine object identifiers of the data objects;

[0248] The apparatus may further include:

[0249] A data assembly module, configured to form data blocks of each data object into a file to be read and return the file to be read to an upper-layer application.

[0250] Optionally, the apparatus may further include:

[0251] A data placement group creation module, configured to create data placement groups according to the number of storage nodes in a storage side and set data placement group identifiers for each data placement group;

[0252] A data placement group mapping module, configured to determine the corresponding relationship between each data placement group and each storage node according to the data placement group identifier.

[0253] For the description of the features in the corresponding embodiment of the data processing apparatus applied to a client, reference may be made to the relevant description in the corresponding embodiment of the data processing method applied to a client, which will not be elaborated here.

[0254] Please refer to Figure 8 , Figure 8 , which is a structural block diagram of another data processing apparatus provided in an embodiment of the present application. The apparatus is applied to a storage node in a storage side, and the storage side includes at least two storage nodes. The apparatus may include:

[0255] A request receiving module 801, configured to receive an aggregation request sent by a client; wherein, the aggregation request includes a preset number of operation requests, and the operation requests are read operation requests or write operation requests for data objects, and the data objects are obtained by splitting a file to be operated in the client;

[0256] A locking module 802, configured to create a unified transaction for the aggregation request and sequentially lock the operation requests in the aggregation request;

[0257] A request batch processing module 803, configured to batch process the operation requests in the aggregation request when the sequential locking of the operation requests is completed.

[0258] Optionally, the aggregation request includes aggregation write data and write control information. The aggregation write data includes data blocks of each data object, and the write control information includes write operation requests corresponding to each data object. The write operation requests include the data object name of the data object, the position of the data block corresponding to the data object in the aggregation write data, and the memory address of the data block corresponding to the data object;

[0259] The request receiving module 801 includes:

[0260] A write request receiving sub-module, configured to receive, through a control plane, write control information sent by a client, and receive, through a data plane, aggregated write data sent by the client;

[0261] A request batch processing module 803, including:

[0262] A write request processing sub-module, configured to perform batch writing on each data block in the aggregated write data according to each write operation request in the write control information.

[0263] Optionally, the aggregation request includes read control information, the write control information includes each read operation request corresponding to each data object, and the read operation request includes the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object;

[0264] A request receiving module 801, including:

[0265] A read request receiving sub-module, configured to receive, through a control plane, read control information sent by a client;

[0266] A request batch processing module 803, including:

[0267] A read request processing sub-module, configured to read the data blocks of each data object according to each read operation request in the read control information and sequentially aggregate them into aggregated read data;

[0268] A data aggregation sub-module, configured to return the aggregated read data to the client through a data plane, so that the client extracts the data blocks of each data object from the aggregated read data according to the read control information.

[0269] Optionally, the storage node includes a replica pool and an erasure code pool;

[0270] A request batch processing module 803, including:

[0271] A replica pool writing sub-module, configured to write the data objects corresponding to each write operation request in the aggregation request into the replica pool;

[0272] An erasure code processing sub-module, configured to determine whether the size of the data objects already written in the replica pool reaches a preset erasure code stripe size; if so, convert the data objects already written in the replica pool into erasure code stripes and write the erasure code stripes into the erasure code pool; if not, retain the data objects already written in the replica pool in the replica pool.

[0273] A request batch processing module 803, including:

[0274] A storage location determination sub-module, configured to determine the storage location where the data objects corresponding to each read operation request in the aggregation request are located;

[0275] The first reading sub-module is configured to, if it is determined according to the storage location that the data objects are all located in the replica pool, batch-read the data blocks of the data objects from the replica pool and return them to the client;

[0276] The second reading sub-module is configured to, if it is determined according to the storage location that the data objects are all located in the erasure code pool, batch-read the data blocks of the data objects from the erasure code pool and return them to the client;

[0277] The third reading sub-module is configured to, if it is determined according to the storage location that the data objects are located in both the replica pool and the erasure code pool, merge the data blocks of the data objects in the replica pool with the data blocks of the data objects in the erasure code pool, and return the merged data blocks of the data objects to the client.

[0278] Optionally, the request batch processing module 803 includes:

[0279] The update sub-module is configured to update the corresponding erasure code stripe in the erasure code pool by using the merged data blocks of the data objects.

[0280] Optionally, the apparatus includes:

[0281] The statistics module is configured to count the access frequencies of the data objects and determine the hot data objects according to the access frequencies;

[0282] The migration module is configured to copy the hot data objects from the erasure code pool to the replica pool.

[0283] For the description of the features in the corresponding embodiments of the data processing apparatus applied to the storage node, reference can be made to the relevant description in the corresponding embodiments of the data processing method applied to the storage node, which will not be elaborated here.

[0284] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above data processing method embodiments.

[0285] Please refer to Figure 9 , Figure 9 , which is a structural block diagram of an electronic device provided by an embodiment of the present application. An embodiment of the present application provides an electronic device 10, including a processor 11 and a memory 12; wherein, the memory 12 is used to store a computer program; the processor 11 is configured to execute the data processing method provided in the foregoing embodiment when executing the computer program.

[0286] For the specific process of the above data processing method, reference can be made to the corresponding content provided in the foregoing embodiments, and details will not be elaborated herein.

[0287] Moreover, as a carrier for storing resources, the memory 12 can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc., and the storage method can be transient storage or permanent storage.

[0288] In addition, the electronic device 10 further includes a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16; among them, the power supply 13 is used to provide operating voltage for each hardware device on the electronic device 10; the communication interface 14 can create a data transmission channel between the electronic device 10 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed on it here; the input / output interface 15 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0289] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above-described data processing method embodiments when running.

[0290] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0291] An embodiment of the present application also provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described data processing method embodiments.

[0292] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described data processing method embodiments.

[0293] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0294] The above has introduced in detail a data processing method, device, distributed storage system, equipment, and program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can also be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A data processing method, characterized in that, Applied to a client, the method includes: Determine a data object to be operated on in the file to be operated on in the client, and determine an object identifier of the data object; wherein, the data object is obtained by splitting the file to be operated on, and the operation corresponding to the data object is a read operation or a write operation; Determine, according to the object identifier, a storage node uniquely corresponding to the data object among at least two storage nodes included in the storage end, and allocate the data object to the storage node; When the number of data objects allocated to the same storage node reaches a preset number, aggregate each operation request corresponding to each data object in the storage node to obtain an aggregate request, and send the aggregate request to the storage node, so that the storage node batch processes the operation requests in the aggregate request.

2. The data processing method according to claim 1, wherein Determine, according to the object identifier, a storage node uniquely corresponding to the data object among at least two storage nodes included in the storage end, and allocate the data object to the storage node, including: Determine, according to the object identifier, a data placement group uniquely corresponding to the data object among at least two data placement groups, and add the data object to the data placement group; wherein, one data placement group uniquely corresponds to one storage node; When the number of data objects allocated to the same storage node reaches a preset number, aggregate each operation request corresponding to each data object in the storage node to obtain an aggregate request, and send the aggregate request to the storage node, including: When the number of data objects in the same data placement group reaches the preset number, aggregate each operation request corresponding to each data object in the data placement group to obtain the aggregate request; Send the aggregate request to the storage node corresponding to the data placement group.

3. The data processing method according to claim 2, characterized in that The operation request is a write operation request, and the aggregate request includes aggregate write data and write control information; Aggregate each operation request corresponding to each data object in the data placement group to obtain the aggregate request, including: Aggregate data blocks corresponding to each data object in the data placement group in the client memory to obtain the aggregate write data; Generate a write operation request for the data object according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregate write data, and the memory address of the data block corresponding to the data object; Aggregate each write operation request corresponding to each data object in the data placement group into the write control information; Send the aggregate request to the storage node corresponding to the data placement group, including: Send the write control information to the storage node through the control plane, and send the aggregate write data to the storage node through the data plane, so that the storage node batch writes each data block in the aggregate write data according to each write operation request in the write control information.

4. The data processing method according to claim 3, wherein The file to be operated on is a file to be stored; Determine a data object to be operated on in the file to be operated on in the client, and determine an object identifier of the data object, including: Receive the file to be stored written by the upper-layer application to the client memory; Split the file to be stored according to a preset data object size to obtain the data object, and set the object identifier for the data object.

5. The data processing method according to claim 2, characterized in that The operation request is a read operation request, and the aggregation request contains read control information; Aggregate the operation requests corresponding to the data objects in the data placement group to obtain the aggregation request, including: Generate a read operation request for the data object according to the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object; Aggregate the read operation requests corresponding to the data objects in the data placement group into the read control information; Send the aggregation request to the storage node corresponding to the data placement group, including: Send the read control information to the storage node through the control plane, so that the storage node reads the data blocks of the data objects according to the read operation requests in the read control information and aggregates them sequentially into aggregated read data; After sending the read control information to the storage node through the control plane, it further includes: Receive the aggregated read data returned by the storage node through the data plane; Extract the data blocks of the data objects from the aggregated read data according to the read operation requests in the read control information.

6. The data processing method according to claim 5, wherein The file to be read is the file to be read; Determine the data object to be operated on in the file to be operated on according to the file to be operated on in the client, and determine the object identifier of the data object, including: Receive an asynchronous read request sent by the upper-layer application, and determine the file to be read corresponding to the asynchronous read request; Determine the data objects included in the file to be read, and determine the object identifiers of the data objects; After extracting the data blocks of the data objects from the aggregated read data, it further includes: Form the file to be read with the data blocks of the data objects, and return the file to be read to the upper-layer application.

7. The data processing method according to claim 2, wherein Before determining the data object to be operated on in the file to be operated on according to the file to be operated on in the client and determining the object identifier of the data object, it further includes: Create the data placement group according to the number of storage nodes in the storage end, and set a data placement group identifier for each data placement group; Determine the corresponding relationship between each data placement group and each storage node according to the data placement group identifier.

8. A data processing method, characterized in that, Applied to the storage nodes in the storage end, the storage end includes at least two storage nodes, and the method includes: Receive an aggregation request sent by the client; wherein, the aggregation request contains a preset number of operation requests, the operation request is a read operation request or a write operation request for a data object, and the data object is obtained by splitting the file to be operated on in the client; Create a unified transaction for the aggregation request, and lock the operation requests in the aggregation request sequentially; When the sequential locking of the operation requests is completed, batch-process the operation requests in the aggregation request.

9. The data processing method according to claim 8, characterized in that, The aggregation request includes aggregated write data and write control information. The aggregated write data includes each data block of each of the data objects. The write control information includes each write operation request corresponding to each of the data objects. The write operation request includes the data object name of the data object, the position of the data block corresponding to the data object in the aggregated write data, and the memory address of the data block corresponding to the data object; Receiving an aggregation request sent by a client, including: Receiving the write control information sent by the client through the control plane, and receiving the aggregated write data sent by the client through the data plane; Batch processing the operation requests in the aggregation request, including: According to each write operation request in the write control information, batch writing each data block in the aggregated write data.

10. The data processing method according to claim 8, wherein The aggregation request includes read control information. The write control information includes each read operation request corresponding to each of the data objects. The read operation request includes the data object name of the data object, the position of the data block corresponding to the data object in the aggregated read data, and the memory address of the data block corresponding to the data object; Receiving an aggregation request sent by a client, including: Receiving the read control information sent by the client through the control plane; Batch processing the operation requests in the aggregation request, including: According to each read operation request in the read control information, reading the data blocks of each of the data objects and sequentially aggregating them into the aggregated read data; Returning the aggregated read data to the client through the data plane, so that the client extracts the data blocks of each of the data objects from the aggregated read data according to the read control information.

11. The data processing method according to any one of claims 8 to 10, characterized in that, The storage node includes a replica pool and an erasure code pool; Batch processing the operation requests in the aggregation request, including: Writing the data objects corresponding to each write operation request in the aggregation request into the replica pool; Judging whether the size of the data objects already written in the replica pool reaches a preset erasure code stripe size; If so, converting the data objects already written in the replica pool into erasure code stripes and writing the erasure code stripes into the erasure code pool; If not, retaining the data objects already written in the replica pool in the replica pool.

12. The data processing method according to claim 11, wherein Batch processing the operation requests in the aggregation request, including: Determining the storage locations of the data objects corresponding to each read operation request in the aggregation request; If it is determined according to the storage location that all the data objects are located in the replica pool, batch reading the data blocks of the data objects from the replica pool and returning them to the client; If it is determined according to the storage location that all the data objects are located in the erasure code pool, batch reading the data blocks of the data objects from the erasure code pool and returning them to the client; If it is determined according to the storage location that the data objects are located in both the replica pool and the erasure code pool at the same time, merging the data blocks of the data objects in the replica pool with the data blocks of the data objects in the erasure code pool, and returning the merged data blocks of the data objects to the client.

13. The data processing method according to claim 12, wherein After returning the merged data blocks of the data objects to the client, it further includes: Update the corresponding erasure stripe in the erasure pool by using the merged data block of the data object.

14. The data processing method according to claim 12, characterized in that, Further comprising: Count the access frequencies of the data objects and determine the hot data objects according to the access frequencies; Copy the hot data objects from the erasure pool to the replica pool.

15. A data processing device, characterized in that, When applied to a client, the apparatus comprises: A data object determination module, configured to determine a data object to be operated on in a file to be operated on in the client and determine an object identifier of the data object; wherein, the data object is obtained by splitting the file to be operated on, and the operation corresponding to the data object is a read operation or a write operation; An allocation module, configured to determine, according to the object identifier, a storage node uniquely corresponding to the data object from at least two storage nodes included in a storage end, and allocate the data object to the storage node; A request aggregation and sending module, configured to, when the number of data objects allocated to the same storage node reaches a preset number, aggregate operation requests corresponding to the data objects in the storage node to obtain an aggregation request, and send the aggregation request to the storage node, so that the storage node processes the operation requests in the aggregation request in batches.

16. A data processing device, characterized in that, When applied to a storage node in a storage end, the storage end includes at least two such storage nodes, and the apparatus comprises: A request receiving module, configured to receive an aggregation request sent by a client; wherein, the aggregation request includes a preset number of operation requests, the operation request is a read operation request or a write operation request for a data object, and the data object is obtained by splitting a file to be operated on in the client; A locking module, configured to create a unified transaction for the aggregation request and sequentially lock the operation requests in the aggregation request; A request batch processing module, configured to, when the sequential locking of the operation requests is completed, batch process the operation requests in the aggregation request.

17. A distributed storage system, characterized in that, Comprising: A client and a storage end, the storage end includes at least two storage nodes; The client is configured to execute the data processing method according to any one of claims 1 to 7; The storage node is configured to execute the data processing method according to any one of claims 8 to 14.

18. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the data processing method according to any one of claims 1 to 7 or the data processing method according to any one of claims 8 to 14 when executing the computer program.

19. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, the data processing method according to any one of claims 1 to 7 or the data processing method according to any one of claims 8 to 14 is implemented.

20. A non-volatile computer-readable storage medium, characterized in that, Computer executable instructions are stored in the non-volatile computer readable storage medium, and when the computer executable instructions are loaded and executed by the processor, the data processing method according to any one of claims 1 to 7 or the data processing method according to any one of claims 8 to 14 is implemented.