Data processing method and related device thereof
By cached the pending data in the data forwarder of the storage cluster, and splicing and forwarding it after the storage server generates redundant data, the problem of repeated overhead transmission bandwidth of the same data in the storage cluster is solved, and network performance and service efficiency are improved.
Patent Information
- Application Number
- PCT/IB2024/060749
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
The same data in the storage cluster repeatedly overhead transmission bandwidth, resulting in a degradation of network performance.
The preset cache space in the data forwarder caches the pending data and sends the pending data to the storage server through the data forwarder. The storage server uses erasure code to encode the data to be processed, generates redundant data, and only transmits the redundant data to the data forwarder. The data forwarder splices and splits the pending data with redundant data, and forwards it to multiple storage servers for storage.
Reduces the number of pending data transmissions from the storage server to the data forwarder, reduces network bandwidth consumption, and improves the service efficiency of the storage cluster.
Smart Images

Figure IB2024060749_08052025_PF_FP_ABST
Abstract
Description
[0001] TECHNICAL FIELD: The present disclosure relates to the field of cloud technology, and in particular to data processing methods and related devices. Background: Storage clusters offer high scalability, high availability, and strong data consistency, leading to a rapid and continuous growth in their scope and scale of application. To improve reliability, storage clusters no longer perform local write operations, meaning both writes and reads generate network traffic. This significantly increases network bandwidth consumed by storage clusters. Therefore, network performance driven by this increased bandwidth becomes a key factor in determining storage cluster performance. The architecture of a storage cluster includes multiple storage racks connected via aggregation layer switches (PSWs). These racks include access layer switches (ASWs), data forwarders, and storage servers. Data forwarders, such as CXL forwarders, interconnect storage servers using the Compute Express Link (CXL) protocol. CXL, short for Compute Express Link, is a multi-protocol interconnect technology bus that provides the PCIE-like CXL.io protocol. As an emerging interconnect, CXL offers advantages such as cache consistency, simplicity, and high speed, facilitating resource pooling. The data to be processed is transmitted sequentially through the aggregation layer switch (PSW), the access layer switch (ASW), and the data forwarder to a storage server in the storage cluster (for example, storage server i in storage rack x, denoted as storage server xi). Storage server xi applies EC encoding to the received data to be processed, generating redundant data. EC stands for Erasure Code. Erasure codes achieve higher data reliability with less redundancy. As a coding technique, erasure codes can add m copies of data (the m copies are the generated redundant data) to k copies of the original data to be processed, and can restore the original data from any k of the k+m copies. Using erasure codes to store data in the storage cluster improves storage reliability. After generating redundant data, storage server xi concatenates the redundant data and the data to be processed to form a codeword (codeword). This codeword is uploaded to the data forwarder, which then splits the codeword into multiple data segments and sends the data segments to multiple other storage servers for storage. In the above process, the data to be processed is transmitted multiple times between storage server x1 and the data forwarder within a short period of time. This repeated transmission of the same data consumes bandwidth, increasing network bandwidth consumption. SUMMARY OF THE INVENTION The data processing method and related devices provided by embodiments of the present invention at least address the problem of repeated transmission bandwidth consumption for the same data in a storage cluster.According to one aspect of the present invention, a data processing method is provided, which is applied to a storage cluster, wherein the storage cluster includes multiple storage racks, the storage racks include a data forwarder and multiple storage servers, and the multiple storage servers are interconnected through the data forwarder. The method includes: caching data to be processed in a preset cache space in the data forwarder, and sending the data to be processed to a first storage server in the storage cluster through the data forwarder; encoding the data to be processed using an erasure code by the first storage server to obtain redundant data, and transmitting the redundant data to the data forwarder; splicing the data to be processed and the redundant data by the data forwarder to obtain spliced data, dividing the spliced data into multiple data slices, and forwarding the data slices to a second storage server in the storage cluster respectively, so that the second storage server stores the data slices based on a predetermined storage operation. In some embodiments, before caching the data to be processed in a preset cache space in the data forwarder, the method further includes: dividing the data to be processed into multiple original data blocks and caching the blocks in the cache space, so that the multiple original data blocks are sent to the first storage server via the data forwarder; wherein the step of encoding the data to be processed using an erasure code by the first storage server includes: encoding the multiple original data blocks using the erasure code by the first storage server to obtain the redundant data. In some embodiments, the storage operation includes log storage and flush storage, and the storage server is provided with a log storage space for log storage and a flush storage space for flush storage. The step of storing the data shards by the second storage server based on a predetermined storage operation includes: writing the data shards to the log storage space by the second storage server based on the log storage, and deleting the data shards from the log storage space after the data shards are written to the flush storage space by the second storage server or another second storage server based on the flush storage. In some embodiments, when the storage operation includes flushing storage, the step of encoding the data to be processed by the first storage server using an erasure code further includes: after the first storage server cross-combines the data to be processed and one or more other data to be processed to obtain multiple combined data blocks, encoding the combined data blocks using the erasure code to obtain the redundant data, wherein the combined data blocks include multiple original data blocks, and different original data blocks belong to different data to be processed.In some embodiments, when caching the data to be processed in a preset cache space in the data forwarder, the method further includes: determining a cache address of the data to be processed in the cache space, so as to send both the data to be processed and the cache address to the first storage server; wherein the step of the data forwarder splicing the data to be processed with the redundant data to obtain spliced data includes: receiving the cache address corresponding to the redundant data sent by the first storage server; determining the data to be processed corresponding to the redundant data in the cache space based on the cache address, so as to obtain the spliced data after splicing the data to be processed with the redundant data. In some embodiments, the step of the data forwarder splicing the to-be-processed data with the redundant data to obtain spliced data further includes: receiving multiple cache addresses corresponding to the redundant data sent by the first storage server, wherein different cache addresses belong to different to-be-processed data, and the first storage server sends multiple cache addresses when first sending the redundant data corresponding to one of the combined data blocks; and determining, based on the multiple cache addresses, multiple original data blocks corresponding to the redundant data in the cache space, to obtain the spliced data by splicing the multiple original data blocks with the redundant data. In some embodiments, after the second storage server stores the data shard based on a predetermined storage operation, the method further includes: deleting the to-be-processed data corresponding to the data shard from the cache space. In some embodiments, the down-flush storage space stores the data slices according to pre-divided storage blocks, and the method further includes: when invalid data exists in a target storage block in the down-flush storage space of any second storage server, recycling the target storage block; during recycling, transmitting non-invalid data to a third storage server through the data forwarder to store the non-invalid data in a termination storage block created in the target down-flush storage space of the third storage server, the invalid data being the data slice deleted or updated in the target storage block, and the second storage server and the third storage server are located in the same storage rack.In some embodiments, before storing the non-invalidated data in the terminated storage block created in the target downlink storage space of the third storage server, the method further includes: copying the non-invalidated data to a reclaim cache space created in the data forwarder, so that when another storage block is reclaimed the next time, if the non-invalidated data in the other storage block exists in the reclaim cache space, the non-invalidated data is forwarded from the reclaim cache space to the other terminated storage block via the data forwarder. In some embodiments, the processing priority of the reclaim process is lower than the processing priority of the pending data by the data forwarder and the storage server. According to another aspect of the present invention, a data processing device is provided, configured to execute the data processing method. According to another aspect of the present invention, a storage cluster is provided, comprising the data processing device. According to another aspect of the present invention, an electronic device is provided, comprising: a processor and a memory storing a program, wherein the program comprises instructions that, when executed by the processor, cause the processor to execute the data processing method. According to another aspect of the present invention, a non-transitory machine-readable medium storing computer instructions is provided. The computer instructions are configured to cause the computer to execute the data processing method. Advantageous effects of embodiments of the present invention: In embodiments of the present invention, before a storage server encodes the data to be processed using erasure coding, the data to be processed is cached in a data forwarder. After the storage server completes the encoding process, the storage server only needs to transmit the redundant data obtained by encoding to the data forwarder. The data forwarder then concatenates the cached data to be processed and the redundant data, splits them, and distributes them for storage. This reduces the number of transmissions of the data to be processed from the storage server to the data forwarder, reduces network bandwidth consumption, and improves the service efficiency of the storage cluster. Details of one or more embodiments of the present invention are set forth in the following figures and description to make other features, objects, and advantages of the present invention more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the embodiments of the present invention or technical solutions in the prior art, the following briefly describes the figures required for use in the embodiments or description of the prior art. Obviously, the figures described below are merely some embodiments of the present invention. Persons skilled in the art can derive other embodiments based on these figures without inventive effort.Figure 1 is a flow chart of a data processing method according to an embodiment of the present invention; Figure 2 is a flow chart of a data forwarder splicing data to be processed and redundant data according to an embodiment of the present invention; Figure 3 is a flow chart of a data forwarder splicing a combined data block composed of multiple original data blocks with corresponding redundant data according to an embodiment of the present invention; Figure 4 is a diagram of data processing after caching user data in a CXL forwarder's buffer space according to an embodiment of the present invention; Figure 5 is a diagram of data processing after dividing user data into original data blocks according to an embodiment of the present invention; Figure 6 is a diagram of scheduling for equal scheduling and within-rack valid data rescheduling according to an embodiment of the present invention; Figure 7 is a comparison of transmission paths for equal scheduling and within-rack valid data rescheduling according to an embodiment of the present invention; Figure 8 is a diagram of valid data reclaim processing in a CXL forwarder's reclaim buffer space according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS The present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the figures and embodiments of this embodiment are for illustrative purposes only and are not intended to limit the scope of protection of this embodiment. In a storage cluster, data to be processed may be transmitted multiple times between a storage server x1 and a data forwarder within a short period of time. This repeated transmission of the same data consumes bandwidth, increasing network bandwidth consumption. Therefore, a first embodiment of the present invention provides a data processing method applicable to the storage cluster. The storage cluster includes multiple storage racks, each of which includes a data forwarder and multiple storage servers. The multiple storage servers are interconnected via the data forwarders. The storage racks are connected via access layer switches and aggregation layer switches. As shown in FIG1 , the method includes the following steps: Step S11: Cache the data to be processed in a preset cache space in the data forwarder, and send the data to the first storage server in the storage cluster via the data forwarder. The data forwarder in this embodiment of the present invention, such as a CXL forwarder that interconnects the storage servers using the CXL protocol, offers advantages such as cache consistency, simplicity, and high speed. Step S12: The first storage server uses an erasure code to encode the data to be processed to obtain redundant data, and transmits the redundant data to the data forwarder.In this embodiment of the present invention, the first storage server that encodes the data to be processed only needs to transmit the redundant data to the data forwarder, eliminating the need to retransmit the data to be processed to the data forwarder. This reduces bandwidth overhead for the data to be processed from the storage server to the data forwarder. Step S13: The data forwarder concatenates the data to be processed with the redundant data to obtain concatenated data. After the concatenated data is divided into multiple data slices, the data slices are forwarded to the second storage server in the storage cluster. The second storage server then stores the data slices based on a predetermined storage operation. The storage servers in the storage cluster encode the received data to be processed and also receive data slices corresponding to the data to be processed that has been encoded by other storage servers. As can be seen, in this embodiment of the present invention, before a storage server encodes the data to be processed using erasure coding, the data to be processed is cached in a data forwarder. After the storage server completes the encoding process, the storage server only needs to transmit the redundant data obtained by encoding to the data forwarder. The data forwarder then concatenates the cached data to be processed with the redundant data and splits and stores them. This reduces the number of transmissions of the data to be processed from the storage server to the data forwarder, reduces network bandwidth consumption, and improves the service efficiency of the storage cluster. In step S11, before caching the data to be processed in a preset cache space in the data forwarder, the method provided in this embodiment of the present invention further includes: splitting the data to be processed into multiple original data blocks and caching them in the cache space, so that the multiple original data blocks can be sent to the first storage server via the data forwarder. Furthermore, the step of encoding the data to be processed using erasure coding by the first storage server includes: encoding the multiple original data blocks using erasure coding by the first storage server to obtain redundant data. The chunking operation of dividing the data to be processed into multiple original data blocks here refers to adjusting irregular or large blocks of data to be processed into original data blocks (user bits) for subsequent EC encoding. In other words, a large block of data is first divided into multiple equal parts, each part consisting of k bits. Each EC encoding generates m parity bits (i.e., redundant data) from k (k is a positive integer greater than 1) user bits. The first storage server transmits only the redundant data consisting of m parity bits to the data forwarder. The data forwarder concatenates the user bits and m parity bits into an EC codeword (i.e., concatenated data) with a code length of n=k+m. The EC codeword is then sliced and the resulting data slices are sent to multiple second storage servers for storage. In other words, each second storage server stores one data slice.In this embodiment of the present invention, storage operations include log storage and flush storage. Because storage servers in a storage cluster also receive and store other data shards while performing encoding processing, the storage servers in this embodiment of the present invention are provided with log storage space for log storage and flush storage space for flush storage. In step S13, the step of storing the data shards by the second storage server based on the predetermined storage operation includes: writing the data shards to the log storage space by the second storage server based on log storage; and deleting the data shards from the log storage space after the second storage server or another second storage server writes the data shards to the flush storage space based on flush storage. Thus, this embodiment of the present invention improves the write performance of the storage cluster by adopting a dual-write mode of log storage and flush storage while reducing network bandwidth consumption. When the storage operation includes a downlink storage, step S12 of encoding the pending data using an erasure code by the first storage server further includes: after the first storage server cross-combines the pending data with one or more other pending data to obtain multiple combined data blocks, encoding the combined data blocks using an erasure code to obtain redundant data. The combined data blocks include multiple original data blocks, and different original data blocks belong to different pending data. Specifically, in this embodiment of the present invention, three adjacent pending data may be cross-combined. For example, the three pending data are A, B, and C, and each pending data is divided into three parts. For example, data A may be divided into A1, A2, and A3. Similarly, data B may be divided into B1, B2, and B3. Similarly, data C may be divided into C1, C2, and C3. During cross-combination, A1, B1, and C1 are combined to form a combined data block, which is then flushed and encoded using EC coding (i.e., the combined data block to be flushed and stored is encoded using erasure coding) via the first storage server, generating redundant data PJL. The first storage server only needs to transmit PJ1 back to the data converter, without transmitting the combined data block consisting of A1, B1, and C1. The data converter then concatenates the cached A1, B1, and C1 with PJ1 to generate corresponding spliced data, and then sends the resulting data shards to multiple storage servers for storage. Similarly, the combined data block consisting of A2, B2, and C2 generates redundant data PJ2, and only PJ2 is uploaded. The combined data block consisting of A3, B3, and C3 generates PJ3, and only redundant data PJ3 is uploaded.Thus, the embodiments of the present invention significantly reduce the number of transmissions of combined data blocks between the storage server and the data forwarder, thereby reducing network bandwidth consumption. Furthermore, by changing the data format of the to-be-processed data through a cross-combination approach, the mapping is broken up, avoiding overlap in the data formats of log storage and download storage, thereby enhancing data consistency. In step S11, when caching the to-be-processed data within a preset cache space in the data forwarder, the method provided by the embodiments of the present invention further includes: determining a cache address of the to-be-processed data within the cache space, and transmitting both the to-be-processed data and the cache address to the first storage server. As shown in FIG2 , in step S13, the step of the data forwarder splicing the to-be-processed data with the redundant data to obtain spliced data includes: Step S21: Receiving the cache address corresponding to the redundant data sent by the first storage server. Specifically, after encoding the to-be-processed data to obtain redundant data, the first storage server does not need to retransmit the to-be-processed data; instead, it transmits only the cache address of the to-be-processed data corresponding to the redundant data to the data forwarder along with the redundant data, thereby reducing network bandwidth consumption associated with transmitting the to-be-processed data. Step S22: Determine the pending data corresponding to the redundant data in the cache space based on the cache address, and concatenate the pending data with the redundant data to obtain concatenated data. After receiving the redundant data and the cache address, the data forwarder reads the cached pending data according to the cache address and concatenates the pending data with the redundant data. As can be seen, in this embodiment of the present invention, after caching the pending data in the cache space of the data forwarder, the cache address of the pending data in the cache space is determined. Therefore, when the first storage server encodes the pending data, the cache address is also sent to the first memory. Consequently, the first memory only needs to transmit the obtained redundant data and the corresponding cache address to the data forwarder, without having to transmit the pending data to the data forwarder. After receiving the redundant data and the cache address, the data forwarder concatenates the pending data read from the cache space based on the cache address with the redundant data, and then divides the obtained concatenated data into multiple data slices and distributes them to multiple second storage servers for storage. As shown in FIG3 , in step S13, the step of the data forwarder splicing the to-be-processed data with the redundant data to obtain the spliced data further includes: step S31: receiving multiple cache addresses corresponding to the redundant data sent by the first storage server, wherein different cache addresses belong to different to-be-processed data, and the first storage server sends the multiple cache addresses when sending the redundant data corresponding to one of the combined data blocks for the first time.As a result, the first storage server not only eliminates the need to transmit the combined data block to the data forwarder, but also, because the first storage server transmits the multiple cache addresses corresponding to the redundant data corresponding to the combined data block when it first transmits the redundant data corresponding to the combined data block, it does not need to transmit the cache addresses corresponding to the combined data block again when sending the redundant data corresponding to other combined data blocks. This further reduces network bandwidth overhead and improves the service efficiency of the storage cluster. Step S32: Determine multiple original data blocks corresponding to the redundant data in the cache space based on the multiple cache addresses, and splice the multiple original data blocks with the redundant data to obtain spliced data. That is, after receiving the redundant data of the combined data block, the data forwarder simply splices the multiple original data blocks constituting the combined data block, read from the cache space based on the received multiple cache addresses, with the redundant data. For example, if the cache addresses of data A, B, and C are the first, second, and third addresses, respectively, then after the combined data block consisting of A1, B1, and C1 is flushed and EC-encoded by the first storage server to generate redundant data PJ1, the first storage server need not transmit the combined data block consisting of A1, B1, and C1 to the data forwarder; instead, it only needs to transmit PJ1, the first address, the second address, and the third address to the data converter. Subsequently, after the first storage server encodes the combined data block consisting of A2, B2, and C2 to generate redundant data PJ2, it only needs to upload PJ2. After encoding the combined data block consisting of A3, B3, and C3 to generate PJ3, it only needs to upload redundant data PJ3, eliminating the need to upload the first, second, and third addresses corresponding to the combined data block, further reducing bandwidth overhead. In step S13, after the second storage server stores the data shards based on a predetermined storage operation, the method provided by this embodiment of the present invention further includes deleting the to-be-processed data corresponding to the data shards from the cache space. That is, after the data to be processed is stored in the second storage server in the form of data shards, the embodiment of the present invention deletes the data to be processed from the cache space to release memory, thereby improving the throughput performance of the cache space, ensuring that subsequent incoming data to be processed can be cached smoothly, and further ensuring the service efficiency of the storage cluster.The flush storage space stores data slices according to pre-divided storage blocks. The method provided by an embodiment of the present invention further includes: when invalid data exists in a target storage block in the flush storage space of any second storage server, reclaiming the target storage block. During the reclaim process, non-invalid data (or valid data) is transmitted via a data forwarder to a third storage server for storage in a termination storage block created in the target flush storage space of the third storage server. Invalid data refers to data slices that have been deleted or updated in the target storage block. The second storage server and the third storage server are located in the same storage rack. After the reclaim process is complete, the target storage block is deleted in its entirety to free up storage space. Since the storage servers corresponding to the termination storage block and the target storage block are located in the same storage rack, the transmission path for reclaiming valid data is greatly shortened, network traffic is not generated, and network transmission pressure during the data reclaim process is significantly reduced. Before storing the non-invalidated data in the terminating storage block created in the target downlink storage space of the third storage server, the method provided in this embodiment of the present invention further includes copying the non-invalidated data to a reclaim cache space created in the data forwarder. This allows the data forwarder to forward the non-invalidated data from the reclaim cache space to the other terminating storage block if the non-invalidated data in the other storage block exists in the reclaim cache space during the next reclaim process on another storage block. Thus, in this embodiment of the present invention, valid data is cached by creating a new reclaim cache space in the data forwarder. When valid data to be subsequently reclaimed exists in the reclaim cache space, the data forwarder forwards the valid data directly from the reclaim cache space to the separately determined terminating storage block, eliminating the need to read the data from the storage block where the valid data resides. Obtaining valid data from the reclaim cache space in the data forwarder significantly shortens the path required to read valid data from the storage block, improves data read performance, and further reduces network bandwidth overhead. In this embodiment of the present invention, the priority of the reclaim process is lower than the priority of the pending data processed by the data forwarder and storage server. Thus, by setting processing priorities, embodiments of the present invention prioritize the transmission of pending data within each storage server and data forwarder, while executing the reclaim process at a lower priority, thereby improving the storage cluster's external service efficiency. As can be seen above, embodiments of the present invention utilize data forwarder interconnection to achieve storage cluster resource pooling while also simplifying and improving data transmission paths, reclaim processing scheduling, priority determination, and computational operations by integrating front-end and back-end operations within the storage cluster.When writing pending data to the storage cluster, multiple transmissions from the storage server to the data forwarder are eliminated. A small cache space is used to cache pending data, enabling data assembly and segmentation on the forwarder side. During the garbage collection (GC) phase, valid data is replicated within the storage rack, replacing global cross-rack replication. This ensures data consistency and reliability while shortening the data transmission path, reducing transmission latency, and reducing bandwidth overhead. Because the collected valid data already has redundancy, no further EC encoding is required, eliminating the need for EC encoding. By distinguishing between recycling traffic and pending data processing traffic, SLA (Service-Level Agreement) priority is achieved, improving the storage cluster's external service efficiency. By setting up a recycling cache space in the forwarder and executing the recycling process synchronously, data transmission pressure is alleviated when the recycling process is triggered, reducing access to the storage server and shortening the transmission path to only the forwarder to the final storage block storing valid data. A second embodiment of the present invention provides a data processing device for executing a data processing method. For details about the data processing method, please refer to the first embodiment of the present invention, which will not be further described herein. The data processing device provided in this embodiment of the present invention caches the data to be processed in a data forwarder before a storage server encodes the data to be processed using erasure coding. After the storage server completes the encoding process, the storage server only needs to transmit the encoded redundant data to the data forwarder. The data forwarder then concatenates the cached data to be processed with the redundant data and splits and distributes them for storage. This reduces the number of transmissions of the data to be processed from the storage server to the data forwarder, reduces network bandwidth consumption, and improves the service efficiency of the storage cluster. A third embodiment of the present invention provides a storage cluster, comprising the data processing device provided in the second embodiment of the present invention. A fourth embodiment of the present invention provides an electronic device, comprising: a processor; and a memory storing a program. The program comprises instructions. When executed by the processor, the instructions cause the processor to execute the data processing method provided in the first embodiment of the present invention. A fifth embodiment of the present invention provides a non-transitory machine-readable medium storing computer instructions, where the computer instructions are used to cause a computer to execute the data processing method provided by the first embodiment of the present invention.A sixth embodiment of the present invention provides a computer program product, including a computer program. When executed by a processor of a computer, the computer program causes the computer to perform the data processing method provided in the first embodiment of the present invention. A seventh embodiment of the present invention, building on the aforementioned six embodiments, provides an application embodiment of the data processing method, where the data forwarders in a storage cluster are CXL forwarders and the storage cluster uses a dual-write mode of log storage and flush storage to store user data. This application embodiment describes how to cache user data (i.e., pending data) before distribution and storage using CXL forwarders, thereby reducing the number of user data transmissions, saving transmission bandwidth overhead, and improving storage cluster performance. As shown in Figure 4, user data to be written to the storage cluster is cached in the cache space of the CXL forwarder, and the cache address of the user data in the cache space is sent along with the user data to storage server i. (This embodiment of the present invention uses server i as the receiving end to illustrate this application embodiment. The storage server can be a physical server or a virtual server, and is therefore applicable to improving system performance under resource pooling.) Storage server i performs log EC encoding on user data to obtain log redundant data PJ (i.e., storage server i uses erasure coding to encode user data for log storage). It then transmits only the log redundant data PJ and the cache address to the CXL forwarder. The CXL forwarder reads the user data from the cache based on the cache address, concatenates the log redundant data PJ with the user data to obtain concatenated data, and then divides the concatenated data into slices, which are sent to multiple servers for logging. Since storage server i no longer needs to upload user data to the CXL forwarder, bandwidth overhead for user data from storage server i to the CXL forwarder is reduced, improving the service efficiency of the storage cluster. Similar to log EC encoding, storage server i performs down-flush EC encoding (i.e., storage server i uses erasure coding to encode user data for down-flush storage). The difference lies in the modified EC encoding data format, but similarly reduces transmission bandwidth overhead. Specifically, Figure 5 illustrates the implementation process of an application embodiment of the present invention. User data entering storage server i is divided into raw data blocks of equal size according to the specified data length. The raw data blocks are then written to the CXL forwarder's cache for caching. The cache address is then used as a unique identifier for the user data. After entering storage server i, each raw data block is directly log EC encoding, generating log redundant data PJ, which is then transmitted back to the CXL forwarder along with the corresponding cache address.The CXL forwarder concatenates user data read from the cache space based on the cache address with the PJ to obtain concatenated data. It then segments the concatenated data and distributes the resulting data shards to multiple storage servers for log storage. When downloading EC encoding, storage server i cross-combines multiple user data blocks to enhance data consistency and prevent overlap between the data formats of log storage and downloading user data. EC encoding is then performed on the cross-combined combined data blocks. For example, storage server i may cross-combine three adjacent user data blocks. If the three user data blocks are A, B, and C, each user data block is divided into three equal parts. For example, data A is segmented into A1, A2, and A3. Similarly, data B is segmented into B1, B2, and B3. Data C is segmented into C1, C2, and C3. The cache addresses of A, B, and C in the cache space are the first, second, and third addresses, respectively. When storage server i cross-combines user data A, B, and C, it combines A1, B1, and C1 into a combined data block and then performs EC encoding on the combined data block (i.e., erasure coding is used on the combined data block to be flushed and stored). After generating redundant data PJ1, storage server i only needs to transmit PJ1 and the cache addresses of user data A, B, and C back to the CXL forwarder, without transmitting the combined data block A1, B1, and C1. The CXL forwarder reads A1, B1, and C1 from the cache space based on the cache addresses, then concatenates A1, B1, and C1 with PJ1 to generate the corresponding concatenated data. The resulting data shards are then sent to multiple storage servers for storage. Similarly, storage server i encodes the combined data block consisting of A2, B2, and C2 to generate redundant data PJ2. Only PJ2 needs to be uploaded, and neither the cache address nor the combined data block needs to be uploaded. Storage server i encodes the combined data block consisting of A3, B3, and C3 to generate redundant data PJ3. In this case, only redundant data PJ3 needs to be uploaded. Thus, this embodiment of the present invention significantly reduces the number of combined data block transmissions between the storage server and the data forwarder, thereby reducing network bandwidth consumption. Furthermore, by changing the data format of the processed data through cross-combination to break up the mapping, it avoids duplicate data formats between log storage and download storage, thereby enhancing data consistency.After the user data is encoded with a correction code, the resulting data fragments are written to the log storage space and the flush storage space. This frees up the CXL forwarder's cache space (i.e., the user data already stored in the cache space is deleted) to be used as a cache for the next set of user data. Once the user data is written to the storage cluster, the cache space is reclaimed, improving cache throughput. Another innovation in the application embodiments of the present invention in saving transmission bandwidth lies in the data recycling process (hereinafter referred to as "GC operation"). Data fragments are written sequentially to the log storage space and then cyclically overwritten (log data is deleted once it is written to the flush storage space). The flush storage space stores received data fragments in pre-divided storage blocks. Because the flush storage space holds a large amount of data, it requires continuous data recycling and GC operations. GC operations merge valid data into newly created storage blocks and delete invalid data from old storage blocks (valid data from old blocks is copied to new storage blocks, and then the old storage blocks are deleted entirely), thereby freeing up storage space. Invalid data refers to old data in a storage block that has been deleted or updated. The upper portion of FIG. 6 illustrates a common scenario of current GC operations in a storage cluster (as an example). In this common scenario, the storage disks under all storage servers within the storage cluster perform equal scheduling. As a result, it is common for source storage blocks (target storage blocks being reclaimed through GC operations) and terminated storage blocks (storage blocks about to be written to through GC operations) to appear in different storage racks. This results in a long transmission path for valid data from the source storage blocks to the terminated storage blocks under this equal scheduling. When the storage cluster stores a large amount of data and the probability of triggering GC operations is high, this will significantly increase the bandwidth consumption caused by reclaiming valid data. In this regard, the lower portion of Figure 4 is a schematic diagram of optimized traffic scheduling during background garbage collection (GC) operations in the storage cluster. In the lower portion of the figure, if a GC operation is required on storage block D, a new termination storage block D1 is created at storage server 2 in storage rack X, to which storage block D belongs. The valid data in storage block D is then written to termination storage block D1 via the CXL forwarder. Compared to equal scheduling, this approach completes data recovery and consolidation within the rack, ensuring distributed data reliability while reducing data transmission volume.As shown in the lower portion of Figure 4 , the selected storage blocks to be reclaimed typically belong to an EC group, which is distributed across multiple storage racks. User data belonging to an EC group is encoded and stored together. Modifying only a portion of the user data requires re-encoding of the data shards corresponding to the data, and the existing EC protection becomes invalid. Therefore, the application embodiment provided by the present invention performs GC operations on the storage blocks containing the data shards belonging to an EC group together. All data shards stored in these storage blocks belong to their respective EC groups to ensure data consistency and reliability. An EC group corresponds to a group of storage blocks, and the data in an EC group is protected by the same EC encoding. Storage space is allocated in the storage server of the storage rack where valid data resides and is set as the address of the terminating storage block. Valid data from the source storage block (D and E in the lower portion of Figure 4 ) is read out and transmitted via a CXL forwarder to the location of the terminating storage block. Invalid data is not reclaimed. The data in the source storage block is then deleted to free up space. This optimizes data flow and reduces duplicate transmissions during frequent GC operations, shortening the transmission path of valid data and significantly reducing bandwidth overhead. This approach is suitable for storage clusters with large data volumes and a high probability of triggering GC operations. Figure 7 compares the transmission path of valid data in Figure 6 with that under equal scheduling. The upper portion of Figure 7 shows the transmission path of valid data under equal scheduling, while the lower portion shows the optimized transmission path for valid data reclaimed and scheduled within the same rack, according to an embodiment of the present invention. GC operations include GC reading valid data and GC writing valid data to the terminating storage block. Because GC data traffic in this embodiment of the present invention only sequentially passes through the source storage block, CXL forwarder, and terminating storage block (the source storage block, CXL forwarder, and terminating storage block belong to the same storage rack), no network traffic is generated, reducing network transmission pressure. Furthermore, existing storage clusters reuse the EC encoding process used for user data writes when performing GC writes. While this eliminates the need for additional GC write operations, the drawback is that the valid data from GC operations already has EC redundancy, eliminating the need for further EC encoding. Furthermore, the reclaim traffic generated during GC operations is not differentiated from the user data write traffic. This makes it difficult to implement more refined priority policies across the frontend and backend for writing user data to the storage cluster (network transmission, encoding, storage, etc.). Since GC operations are backend operations of the storage cluster, while user data writes directly demonstrate the storage cluster's service capabilities, prioritizing reclaim traffic below that of user data writes can improve customer service agreements (SLAs).In the lower portion of Figure 7, valid data passes through the CXL forwarder without EC encoding. At the same time, traffic marking is performed on both read and write GC operations, which assigns a processing priority to the packets used for reclaimed processing. As a result, user data is prioritized across the storage servers and CXL forwarders, while GC operations are executed at a lower processing priority. In an application embodiment of the present invention, the data replication scope of GC operations is optimized from the entire storage cluster to individual storage racks, reducing the overall data transmission volume. Therefore, as shown in FIG8 , providing a reclaim buffer space in the CXL forwarder to cache valid data to be reclaimed can improve the data hit rate in GC operations. Specifically, valid data is copied to the reclaim buffer space created in the CXL forwarder. When another storage block is next reclaimed, if valid data exists in the reclaim buffer space in the other storage block, the CXL forwarder forwards the valid data from the reclaim buffer space to a newly determined terminating storage block in the same rack. This eliminates the need to read the valid data from the source storage block and allows the data to be directly transferred from the CXL forwarder's reclaim buffer space to the terminating storage block (such as the terminating storage block created in storage server n in FIG8 ). This further shortens the data transmission path and reduces data transmission overhead. Especially when the cluster's storage watermark is high, GC operations are frequently triggered. Cache valid data in the CXL forwarder reduces accesses to the source storage block. Transferring the valid data held in the cache to the final storage block shortens data transmission latency and saves a data transmission path (i.e., the path from the storage server where the source storage block resides to the CXL forwarder). The storage cluster's management node transmits a eviction instruction (the GC instruction shown in Figure 8) over the network, specifying the source storage block and the final storage block. If a GC cache hit occurs (i.e., valid data from the source storage block exists in the eviction cache), a new space is created in the eviction cache to cache the new valid data. The valid data from the source storage block already cached in the eviction cache is quickly copied to this newly created space. If the same valid data is subsequently transferred, the valid data is retrieved from the CXL forwarder's eviction cache, significantly shortening the path required to read the valid data from the storage disk. This improves data read performance during GC operations.If the reclaimed cache space misses (i.e., the reclaimed cache space does not cache the valid data to be reclaimed), cache scheduling rules (such as LRU scheduling, which stands for least recently used and is a common memory scheduling method that prioritizes data with the oldest last access time when the cache is full) are used to determine whether to create new space in the reclaimed cache space to cache the valid data read from the source storage block. This valid data enters the CXL forwarder and is then transmitted and written to the final storage block. A CRC check (Cyclic Redundancy Check, a commonly used error-checking code in data communications) is performed at each step of the read, transmit, and write process to ensure data consistency and accuracy. Application embodiments of the present invention utilize CXL forwarder interconnection to achieve storage cluster resource pooling while simplifying and improving the data transmission path, reclaim processing scheduling, priority determination, and computational operations by integrating front-end and back-end operations of the storage cluster. When writing pending data to the storage cluster, multiple transmissions from the storage server to the CXL forwarder are eliminated. Small-capacity cache space is used to cache pending data, enabling data combination and segmentation on the CXL forwarder. During the data recovery phase, valid data is replicated within the storage rack, replacing global cross-storage rack replication. This ensures data consistency and reliability, shortens the data transmission path, reduces transmission latency, and reduces bandwidth overhead along the transmission path. Because the recovered valid data already has redundancy, no further EC encoding is required, saving EC encoding operations. By distinguishing between recovery processing traffic and user data write traffic, priority setting is achieved to ensure SLAs and improve the storage cluster's external service efficiency. By providing recovery cache space in the CXL forwarder and executing the recovery process synchronously, data transmission pressure is alleviated when the recovery process is triggered, reducing access to the storage server, further shortening the transmission path, and thus further reducing data transmission overhead. The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.In the context of embodiments of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. Machine-readable signal media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. It should be noted that the term "including" and its variations used in embodiments of the present invention are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." The modifiers "one" and "a plurality" in the embodiments of the present invention are illustrative and non-restrictive. Those skilled in the art will understand that, unless the context clearly indicates otherwise, they should be understood to mean "one or more." The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in the embodiments of the present invention are all authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or deny. The various steps described in the method implementations provided in the embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method implementations may include additional steps and / or omit steps. The scope of protection of the present invention is not limited in this respect.
[0002] The term "embodiment" in this specification refers to specific features, structures, or characteristics described in conjunction with the embodiment that may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it imply that these embodiments are mutually exclusive, independent, or optional. Each embodiment in this specification is described in a related manner, and reference is made to the common and similar parts of each embodiment. In particular, the device, equipment, and system embodiments are described more briefly because they are generally similar to the method embodiments. For relevant parts, reference is made to the description of the method embodiments. The above-described embodiments represent only a few implementation methods of the present invention, and while the description is relatively specific and detailed, this should not be construed as limiting the scope of patent protection. It should be noted that variations and modifications are possible for those skilled in the art without departing from the spirit of the present invention, and such variations and modifications are within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
Claims 1. A data processing method, the method being applied to a storage cluster, the storage cluster comprising a plurality of storage racks, the storage racks comprising a data forwarder and a plurality of storage servers, the plurality of storage servers being interconnected via the data forwarder, the method comprising: The data to be processed is cached in a preset cache space in a data forwarder, and the data to be processed is sent to a first storage server in a storage cluster through the data forwarder; the data to be processed is encoded by the first storage server using an erasure code to obtain redundant data, and the redundant data is transmitted to the data forwarder; the data to be processed and the redundant data are spliced by the data forwarder to obtain spliced data, and after the spliced data is divided into multiple data slices, the data slices are forwarded to the second storage server in the storage cluster respectively, so that the data slices are stored by the second storage server based on a predetermined storage operation.
2. The method according to claim 1, wherein: Before caching the data to be processed in a preset cache space in the data forwarder, the method further includes: dividing the data to be processed into multiple original data blocks and then caching them in the cache space, so as to send the multiple original data blocks to the first storage server through the data forwarder; the step of encoding the data to be processed by the first storage server using an erasure code includes: encoding the multiple original data blocks by the first storage server using the erasure code to obtain the redundant data.
3. The method according to claim 2, wherein: The storage operation includes log storage and flush storage, the storage server is provided with a log storage space for log storage and a flush storage space for flush storage, and the step of storing the data slices based on a predetermined storage operation through the second storage server includes: writing the data slices into the log storage space based on the log storage through the second storage server, and after writing the data slices into the flush storage space based on the flush storage through the second storage server or another second storage server, deleting the data slices in the log storage space.
4. The method according to any one of claims 2 to 3, wherein: The storage operation includes refreshing the storage. The step of encoding the data to be processed by the first storage server using an erasure code also includes: after the first storage server cross-combines the data to be processed and one or more other data to be processed to obtain multiple combined data blocks, encoding the combined data blocks using the erasure code to obtain the redundant data, wherein the combined data block includes multiple original data blocks, and different original data blocks belong to different data to be processed.
5. The method according to claim 4, wherein: When caching the data to be processed in a preset cache space in the data forwarder, the method also includes: determining a cache address of the data to be processed in the cache space, so as to send both the data to be processed and the cache address to the first storage server; wherein the step of the data forwarder splicing the data to be processed with the redundant data to obtain spliced data includes: receiving the cache address corresponding to the redundant data sent by the first storage server; determining the data to be processed corresponding to the redundant data in the cache space based on the cache address, so as to obtain the spliced data after splicing the data to be processed with the redundant data.
6. The method according to claim 5, wherein: The step in which the data forwarder splices the data to be processed and the redundant data to obtain the spliced data also includes: receiving multiple cache addresses corresponding to the redundant data sent by the first storage server, wherein different cache addresses belong to different data to be processed, and the first storage server sends multiple cache addresses when sending the redundant data corresponding to one of the combined data blocks for the first time; determining multiple original data blocks corresponding to the redundant data in the cache space based on the multiple cache addresses respectively, so as to obtain the spliced data after splicing the multiple original data blocks with the redundant data.
7. The method according to claim 3, wherein: After the data shard is stored by the second storage server based on a predetermined storage operation, the method further includes: deleting the to-be-processed data corresponding to the data shard from the cache space.
8. The method according to claim 3, wherein: The down-flush storage space stores the data slices according to pre-divided storage blocks, and the method further includes: when invalid data exists in a target storage block in the down-flush storage space of any of the second storage servers, recycling the target storage block, and during recycling, transmitting non-invalid data to a third storage server through the data forwarder, so as to store the non-invalid data in the target storage block of the third storage server; The terminated storage block created in the storage space is flushed, the invalid data is the data slice deleted or updated in the target storage block, and the second storage server and the third storage server are located in the same storage rack.
9. The method according to claim 8, wherein: Before storing the non-invalid data in a termination storage block created in the target down-flush storage space of the third storage server, the method further includes: copying the non-invalid data to a recovery cache space created in the data forwarder, so that when another storage block is reclaimed next time, if the non-invalid data in another storage block exists in the recovery cache space, the non-invalid data is forwarded from the recovery cache space to another termination storage block through the data forwarder.
10. The method according to claim 8, wherein: The processing priority of the recycling process is lower than the processing priority of processing the to-be-processed data by the data forwarder and the storage server.
11. A data processing device, wherein the data processing device is used to execute the data processing method according to any one of claims 1 to 10.
12. A storage cluster, comprising the data processing device according to claim 11.
13. An electronic device, comprising: A processor, and a memory storing a program, wherein the program includes instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 10.
14. A non-transitory machine-readable medium storing computer instructions for causing the computer to execute the method of any one of claims 1 to 10. 17
Citation Information
Patent Citations
Redundant data encoding method for untrusted environment and storage medium
CN111475839A
Data processing method in storage cluster, storage cluster and equipment
CN115113979A
Recovering Error Corrected Data
US20210365337A1
Redundant data calculation method and apparatus
US20220253356A1