A distributed storage system read optimization method based on erasure code

By performing erasure coding calculations on the client and cache non-full striped data, the read request strategy of the distributed storage system is optimized, and the problems of amplification and delay of read operations are solved, and the read performance and reliability are improved.

CN117873378BActive Publication Date: 2025-05-06CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311717429.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-05-06
Estimated Expiration
2043-12-14

AI Technical Summary

Technical Problem

In distributed storage systems, read operations have problems with amplification and high latency, especially in the case of non-full striped reading, resulting in insufficient read performance and reliability.

Method used

Erasing code calculation is performed on the client and non-full striped data is cached, and read request strategies are optimized, including cached read, partial read and full striped read, reducing read request and read bandwidth, and improving read performance and reliability.

Benefits of technology

By caching non-full striped data on the client and optimizing read strategies, the read latency is significantly reduced and the read performance and reliability is improved, ensuring read optimization without sacrificing write performance and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117873378B_ABST
    Figure CN117873378B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed storage system read optimization scheme based on erasure code, which belongs to the field of computer storage technology and distributed storage technology, and includes the following steps: S1, try "cache read" first, S2, try "partial read" if the cache does not hit, S3, if the condition of partial read is not met, perform "full stripe read", once the read range hits the cache interval, the hit part does not need to send a read request to the data node, S4, wait for the "partial read" or "full stripe read" request to return successfully. A data storage method is proposed for performing erasure code calculation on the client, caching non-full stripe data, and simplifying the read and write process. By caching non-full stripe data on the client, reducing read requests and read bandwidth, the read delay can be reduced. Based on the above data storage method, different read strategies are specified on the client for different read scenarios, which improves the performance and reliability of reading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer storage technology-distributed storage technology, and specifically to a distributed storage system read optimization method based on erasure codes. Background Art

[0002] Distributed storage is an architecture different from traditional centralized storage. In order to achieve more flexible scalability and larger storage scale, it adopts a decentralized networking method, connected through internal switches, and provides a unified storage resource pool based on distributed software. For the protection of stored data, traditional centralized storage uses redundant dual controllers and RAID technology to ensure data reliability; while for larger-scale distributed storage systems, most of them use multiple copies and erasure coding technology to ensure data reliability.

[0003] The main advantage of multiple copies is high performance, but low storage space utilization and high cost; erasure coding technology sacrifices some CPU computing resources and network load, improves storage space utilization, and provides the reliability of approximate copies. Therefore, erasure coding technology has been widely used in distributed storage systems.

[0004] For M+N erasure codes, after erasure code encoding and calculation for M data blocks, N check blocks will be generated, and then the data blocks of these M+N shards will be saved to M+N data nodes. By obtaining the data of any M shards, M data blocks can be restored according to the decoding operation of the erasure code, which means that the loss of up to N data blocks can be tolerated. Since the complete stripe needs to be gathered to calculate the final erasure code when writing, for writing of non-full stripes, after the stripe is gathered, the process of reading-EC encoding-writing needs to be performed, which causes the problem of write amplification. The industry has made some optimizations to address the write amplification problem, such as sending the data of the non-full stripe to one of the data nodes, and then sending it to the corresponding data node and N check data nodes according to the position of each stripe unit data in the stripe, so as to convert the non-full stripe data into N+1 copies for storage, avoiding the write amplification problem of writing non-full stripe data. The read process also has the problems of amplification and high latency, which need to be optimized. For example, to read data that does not fill up a stripe, the read interval needs to be aligned with the stripe width, which expands the range of the read request interval. Moreover, the completion of the read operation depends on the slowest responding data node, that is, the success of the read operation requires that all requests from the M shards be successfully returned, which increases the read latency. Summary of the invention

[0005] The invention provided by the present invention aims to provide a distributed storage system read optimization method based on erasure codes. In order to optimize the performance and reliability of reading, the present invention places the EC calculation on the client, and caches the non-full stripe data on the client after persisting it on the storage node. In this way, the data cached by the client corresponds to the data of the non-full stripe part in the user data block, that is, the "tail" of the data block. The data not cached by the client corresponds to the data of the full stripe part of the user. Based on the above data storage method, the present invention formulates the read strategy with the highest performance and reliability by parsing the read request and the state of the reference data node, thereby optimizing the performance and reliability of reading.

[0006] In order to achieve the above effect, the present invention provides the following technical solution: a distributed storage system read optimization method based on erasure code, comprising the following steps:

[0007] S1. Try "cache read" first.

[0008] S2. If the cache does not hit, try "partial read".

[0009] S3: If the partial read condition is not met, a "full stripe read" is performed. Once the read range hits the cache interval, the hit part does not need to send a read request to the data node.

[0010] S4: Wait until the "partial read" or "full stripe read" request returns successfully.

[0011] S5: The data of the non-full stripe is taken out from the cache and appended to the end. This is "cache read" + "partial read" or "cache read" + "full stripe read".

[0012] S6. When a read failure occurs, a retry is performed. The strategy for the retry read is "redundant read". Of course, for the cache interval that hits, it is "cache read" + "redundant read".

[0013] Furthermore, the following steps are included: according to the operation steps in S1, the data to be written by the client process is divided into a full stripe part and a non-full stripe part, the full stripe part is encoded and calculated with an erasure code, the non-full stripe part is first placed in the client's temporary cache, and then appended to the end of the shard request, the end of the data shard request saves the data belonging to this shard in the non-full stripe data, if there is no data in this I / O, it will not be appended, and the end of the verification data node request saves the full amount of non-full stripe data that needs to be sent this time, and when the write request sent is successful, the client's non-full stripe cache is updated to the latest non-full stripe data.

[0014] Furthermore, the following steps are included: according to the operation steps in S1, the cache read, if the read range covers the cache part, then the data of the non-full stripe is directly taken from the client cache, which is equivalent to a cache hit, and there is no need to send a read request to the data node, which can greatly reduce the delay.

[0015] Further, the following steps are included: According to the operation steps in S1, the consistency of the cache needs to be guaranteed by a write mark: when the stripe where the cache is located is being written, the previous non-full stripe of the data in the cache needs to be read after the write is completed and updated.

[0016] Further, the steps include: according to the operation steps in S2, the partial read: if the read length is less than a stripe width, only send read requests to the relevant data nodes, read as many as needed, and reduce the number of read requests sent. Since the read request is sent asynchronously, the completion of the read request depends on the data node with the slowest response. Therefore, reducing the read request can also optimize the overall delay.

[0017] Furthermore, the following steps are included: according to the operation steps in S2, if it is found that the data sharding node related to the partial read is unavailable, the partial read method is abandoned and "redundant read" is performed directly. The partial read request will only be sent to the data sharding node, so there is no need to perform stripe alignment or EC decoding calculation.

[0018] Further, the following steps are included: according to the operation steps in S3, for the full stripe read, if the read length is greater than a stripe width, each data node needs to send a read request.

[0019] Furthermore, the following steps are included: according to the operation steps in S3, in order to simplify the reading process and be able to perform EC decoding on these stripes, the read range will be stripe aligned, and then M available data node addresses are obtained according to the cluster view, and a read request is sent to these M data nodes.

[0020] Furthermore, the following steps are included: According to the operation steps in S3, for the sake of performance and reliability, the cluster management node will give priority to the data sharding nodes, the data nodes with low load and availability, and the full stripe read will only perform erasure code decoding when there is an abnormality in the data sharding node and the data of the verification sharding node is read, otherwise only the data is spliced ​​to minimize the overhead of EC decoding calculation.

[0021] Further, the following steps are included: according to the operation steps in S6,

[0022] S601. If "partial read" or "full stripe read" fails, or if the relevant data shard node is found to be unavailable when attempting "partial read", "redundant read" is performed.

[0023] S602, redundant read will send read requests to a maximum of M+N data nodes (excluding nodes that have been found to be unavailable), and will start parsing data as long as the number of successful read responses reaches M.

[0024] S603: If the M responses are all data fragment information, they are concatenated in order.

[0025] S604: If there is verification data sharding information, erasure code decoding is performed on the client.

[0026] S605: If necessary, append the non-full stripe data in the client cache to the end.

[0027] S606. If it is found during the parsing process that the data format does not meet the requirements (it may be a data node failure or the data length is inconsistent with the expectation), the response is discarded and the read request of other responses continues to be waited for. After M responses are met, the parsing continues until the complete data is parsed.

[0028] S607: If the number of successful read responses is less than M at this time, try to restore as much data as possible using replicas.

[0029] S608. If all available read responses are data shard information, then splicing is started from the first data shard in sequence. Once the response information of the corresponding shard is missing, parsing is stopped, and the data finally read will be less than expected.

[0030] S609: If the available read response contains the check shard information, determine whether the read range falls entirely within the non-full stripe portion. If so, retrieve the required data from the check shard response, which is the complete data expected to be read.

[0031] The present invention provides a distributed storage system read optimization method based on erasure codes, which has the following beneficial effects: the distributed storage system read optimization scheme based on erasure codes, without sacrificing reliability and write performance, proposes a data storage method for performing erasure code calculations on the client, caching non-full stripe data, and simplifying the read and write processes; by caching non-full stripe data on the client, reducing read requests and read bandwidth, the read delay can be reduced; based on the above data storage method, different read strategies are specified on the client for different read scenarios, thereby improving the read performance and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A schematic diagram of a storage system based on erasure codes of a distributed storage system read optimization method based on erasure codes of the present invention;

[0033] Figure 2A schematic diagram of a flow chart of a distributed storage system read optimization method based on erasure codes according to the present invention. DETAILED DESCRIPTION

[0034] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments.

[0035] The present invention provides a technical solution:

[0036] Example 1, please refer to Figure 1-2 ,

[0037] A distributed storage system read optimization method based on erasure code includes the following steps: S1, giving priority to "cache reading", the data to be written by the client process is divided into a full stripe part and a non-full stripe part, the full stripe part is subjected to erasure code encoding calculation, the non-full stripe part is first placed in the client temporary cache, and then appended to the end of the shard request, the end of the data shard request saves the data belonging to the current shard in the non-full stripe data, if there is no data in this I / O, it will not be appended, and the end of the verification data node request saves the full amount of non-full stripe data that needs to be sent this time, when the write request sent is successful, the non-full stripe cache of the client is updated to the latest non-full stripe data, cache reading, if the read range covers the cache part, then the non-full stripe data is directly from the client The cache is retrieved from the client cache, which is equivalent to a cache hit. There is no need to send a read request to the data node, which can greatly reduce the latency. The consistency of the cache needs to be guaranteed by the write tag: when the stripe where the cache is located is being written, the previous non-full stripe of the data in the cache needs to wait for the write to be completed and updated before reading. S2. If the cache does not hit, try "partial read". Partial read: If the read length is less than the width of a stripe, only send a read request to the relevant data node, and read as many as needed to reduce the number of read requests sent. Since the read request is sent asynchronously, the completion of the read request depends on the data node with the slowest response. Therefore, reducing the read request can also optimize the overall latency. If it is found that the data sharding node related to the partial read is unavailable, abandon the partial read method and proceed directly. "Redundant read" is performed. Some read requests will only be sent to the data sharding nodes, so there is no need to perform stripe alignment or EC decoding calculation. S3. If the conditions for partial read are not met, "full stripe read" is performed. Once the read range hits the cache interval, the hit part does not need to send a read request to the data node. For full stripe read, if the read length is greater than a stripe width, each data node needs to send a read request. In order to simplify the read process and be able to perform EC decoding on these stripes, the read range will be stripe aligned, and then M available data node addresses will be obtained according to the cluster view, and read requests will be sent to these M data nodes. For performance and reliability, the cluster management node will give priority to data sharding nodes, low-load and available data nodes. Point, full stripe reading only performs erasure code decoding when there is an abnormality in the data shard node and the data of the verification shard node is read. Otherwise, it only splices the data to minimize the overhead of EC decoding calculation. S4. After the request of "partial read" or "full stripe read" returns successfully, S5. Take out the data of the non-full stripe from the cache and append it to the end. At this time, it is "cache read" + "partial read" or "cache read" + "full stripe read". S6. When a read failure occurs, retry is performed. The retry read strategy is "redundant read". Of course, for the cache interval that hits, it is "cache read" + "redundant read". S601. If "partial read" or "full stripe read" fails, or when trying "partial read", it is found that the relevant data shard node is unavailable,Then "redundant read" is performed. S602. Redundant read will send read requests to a maximum of M+N data nodes (excluding nodes that have been found to be unavailable). As long as the number of successful read responses is M, data parsing will begin. S603. If the M responses are all data sharding information, they are spliced ​​in order. S604. If there is verification data sharding information, erasure code decoding is performed on the client. S605. If necessary, the non-full stripe data in the client cache is appended to the end. S606. If the data format is found to be inconsistent with the requirements during the parsing process (it may be a data node failure or the data length is inconsistent with the expectation), the response is discarded and the read request of other responses is continued to wait. After M responses are met, parsing continues until parsing. The complete data is obtained, S607. If the number of successful read responses is less than M at this time, try to use copies to restore as much data as possible, S608. If the available read responses are all data shard information, start splicing from the first data shard in sequence. Once the response information of the corresponding shard is missing, stop parsing, and the final data read will be less than expected. S609. If the available read response contains the check shard information, determine whether the read range all falls on the non-full stripe part. If so, take out the required data from the check shard response, which is the complete data expected to be read. It can be seen from the above process that the above data storage method can significantly improve the performance and reliability of reading. Since the non-full stripes are cached, it also makes The reading process is simpler. It should be noted that "full stripe reading" and "redundant reading" have their own advantages and disadvantages, and the two can replace each other. Compared with "redundant reading", "full stripe reading" has the best performance when the cluster is normal, because it will select M data sharding nodes. When the request returns, the client only needs to splice the data and does not need to perform erasure code decoding calculations. "Redundant reading" sends read requests to M+N nodes. The first M read success responses returned are not necessarily all data sharding nodes. Once there is a verification data node, the client needs to perform an EC decoding calculation, which increases the reading overhead. In the case of abnormal data sharding nodes in the cluster, "redundant reading" is obviously more advantageous in performance, because "full stripe reading" selects The M data nodes selected are not necessarily the first M data nodes with the fastest response among the available nodes. The disadvantage of "full stripe reading" is that its reliability is worse than that of "redundant reading", because the probability of receiving M successful read responses after sending M read requests is obviously lower than the probability of receiving M successful read responses after sending a maximum of M+N read requests. Therefore, in actual use, it is necessary to select a suitable read strategy between full stripe reading and redundant reading according to the actual test results. The present invention puts the data of non-full stripes in the client process, and the calculation of the erasure code is completed on the client. After the erasure code of the full stripe part is calculated, the original data of the non-full stripe is appended to the end of the sent data of each data node, and the full amount of non-full stripe data is appended to the end of the sent data of the verification data node.Finally, they are sent to the corresponding data nodes respectively. When the user data completes the original non-full stripe interval, EC encoding is performed on the client to convert the non-full stripe data into full stripe data storage. If there is still new non-full stripe data in this I / O, the new non-full stripe data will replace the old non-full stripe data in the client cache after the persistence of the user data is completed. In this way, the non-full stripe data is persisted in the form of N+1 copies, which not only caches the non-full stripe data, simplifies the reading and writing process, but also avoids the copying of non-full stripe data between data nodes. Whether it is full stripe data or non-full stripe data, it can tolerate the loss of N data blocks and meet EC N+M data protection requirements are met, and EC calculation is performed on the client to make the overall read and write process clearer. Data storage service nodes do not need to forward data, which simplifies the I / O process of the data storage node. Based on the above-mentioned EC+cache non-full stripe storage method on the client, the present invention formulates different read strategies for different I / Os, and optimizes the read performance and reliability. In terms of performance, by analyzing the offset and length of the client read request, the client cache is used to optimize the read performance: first check whether the read hits the cached part of the data. If it hits this part of the data, this part of the data is directly read from the cache. The delay of the cache hit is much lower than the network communication, and there is no need to send a request to the corresponding data node. For small I / O, the number of read requests sent is reduced, and for large I / O, the read bandwidth can also be reduced. The present invention calls it "cache read". Secondly, for reads with a length less than a stripe width, stripe alignment is not performed, but requests are sent to the required data nodes based on the actual offset and length. This also reduces the number of read requests and has a certain probability of reducing the delay in read completion (for example, the response of the node that does not need to send the read request). The response speed is relatively slow. When this reading method is used, there is no need to send a read request to the node, which reduces the delay of user-side reading). The present invention calls it "partial reading". Again, for readings greater than one stripe width, it is necessary to send read requests to at least M data nodes. Since erasure code decoding may be involved, stripe alignment is required. The present invention calls it "full stripe reading". Before full stripe reading, the cluster metadata management node needs to screen the data nodes to which the read request is to be sent. The screening principle is: nodes with normal disk status are given priority, data sharding nodes are given priority, and nodes with low load are given priority. In this way, the decoding operation of the erasure code (direct splicing of data nodes) can be avoided as much as possible, and the delay of the read request response can be reduced. In terms of reliability, on the one hand, the above-mentioned cache reading and partial reading methods improve the success rate of reading by reducing the number of read requests. On the other hand, for read requests that have been responded to, if a read failure occurs, a retry will be performed. The retry method is to send a read request to all data nodes. As long as the number of successful read responses received meets M and the conditions for restoring the complete stripe data are met, splicing or decoding will be attempted until the complete data is obtained.Even if there are less than M successful read responses, we will try to restore the data in the form of copies: for example, if the data that the user needs to read happens to fall within a non-full stripe interval, and the verification data node successfully returns the data, the data that the user needs can be retrieved from the data returned by the verification node, which is called "redundant read" in the present invention.

[0038] Example 2, please refer to Figure 1-2 ,

[0039] A distributed storage system read optimization method based on erasure coding comprises the following steps:

[0040] S1. Deploy 6 data nodes in EC 4+2 mode. The fault domain is Host level. The stripe unit size is 4k, the stripe width size is 16k, and the 6 data nodes are called d0, d1, d2, d3, d4, and d5 respectively. Among them, d0-d3 are data sharding nodes, and d4-d5 are verification data nodes. User data is abstracted into data blocks, and data blocks are striped according to the stripe width: EC needs to be calculated once for user data of each stripe width, and the calculated verification data is stored in the corresponding data node, which is the d4-d5 node in this example.

[0041] S2. Deploy a cluster metadata management node m0, to which d0-d5 regularly reports node load, disk information and other status.

[0042] S3. Start a client c0 process to handle the read and write requests from the user side. The erasure code calculation plug-in e0 is integrated into the c0 process. For a logical data block chunk0, the client applies for a 16k temporary cache tmp_buf and a 16k non-full stripe cache buf. When this logical data block is no longer written, tmp_buf and buf are released.

[0043] S4, when writing 0-9527Bytes of chunk0, it does not meet the stripe width of 16k, and will first put 9527Bytes of data into tmp_buf, then send 4096Bytes of data to d0, 5431Bytes of data to d1, and 9527Bytes of data to d4 and d5. When c0 receives the write success response from d0, d1, d4, and d5, c0 will hand over the data in tmp_buf to buf for management, and mark tmp_buf as empty.

[0044] S5. When writing 9527-16387Bytes of chunk0, the 9527Bytes in tmp plus the first 6857Bytes to be written can make up a stripe width, and erasure code encoding calculation is required. The remaining 3Bytes are written into the tmp_buf temporary cache, and 4096+3Bytes of data are sent to d0, 4096Bytes of data are sent to d1-d3, and 4096+3Bytes of data are sent to d4-d5. The data node needs to process the data sent by c0 according to the size of chunk0 and its own shard ID. When d0-d5 all return success, update c0's buf to the latest 3Bytes of non-full stripe data.

[0045] S6. If you want to read the data of 16384-16386Bytes of chunk0 at this time, read it directly from buf without sending any read request. This is "cache read".

[0046] S7: If you want to read 100-4196Bytes of chunk0, c0 will only send a request to d0 to read 3996Bytes, and send a request to d1 to read 100Bytes. This is a "partial read".

[0047] S8. If you want to read 16380-16387Bytes of data in chunk0, c0 will only send a request to d3 to read 4Bytes, and the last 3Bytes will be directly taken from buf after d3 returns the request. This is "cache read" + "partial read".

[0048] S9. If the cluster is normal, and you want to read 0-16384Bytes of data in chunk0, send a read request of 4096Bytes to d0-d3. After the read is successful, the data is directly spliced. If there is an abnormal data node in the cluster, such as d1, c0 will obtain the 4 best data nodes from m0. For example, d0, d2, d3, and d4 are obtained. Then c0 sends a read request of 4096Bytes to these nodes respectively. After the read is successful, the erasure code decoding calculation needs to be performed. The above two scenarios are "full stripe read".

[0049] S10. If the cluster is normal, to read 0-16387Bytes of data in chunk0, a read request of 4096Bytes is sent to d0-d3. After all the read requests are successfully returned, 3Bytes of data are taken out from buf and appended to the end. If there is an abnormal data node in the cluster, such as d1, c0 will obtain the 4 best data nodes from m0. For example, d0, d2, d3, and d4 are obtained. Then c0 sends a read request of 4096Bytes to these nodes respectively. After all the read requests are successfully returned, erasure code decoding calculation is required, and then 3Bytes of data are taken out from buf and appended to the end. The above two scenarios are "cache read" + "full stripe read".

[0050] S11. If the first read fails, a retry will be performed. The retry will send a read request to all available data nodes. For example, a read request is sent to all 6 data nodes. When c0 receives 4 successful read responses, it starts to parse the data. At this time, it will decide whether to use data splicing or erasure code decoding calculation based on the properties of the 4 shards. When an error occurs in the parsing, the response with incorrect data format will be discarded, and the data will continue to be parsed after the remaining read requests return successfully until it succeeds. This is a "redundant read". If the read also hits the non-full stripe data in the cache, it is a "cache read" + "redundant read".

[0051] S12. If this user data block is no longer written, or is actively closed (write-forbidden), the cache corresponding to the data block will be released, "cache read" will become invalid, and the other three read strategies can still be used.

[0052] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A distributed storage system read optimization method based on erasure code, characterized in that: The data to be written by the client process is divided into a full stripe part and a non-full stripe part. The full stripe part is calculated by erasure coding. The non-full stripe part is first placed in the client's temporary cache and then appended to the end of the shard request. The end of the data shard request saves the data belonging to this shard in the non-full stripe data. If there is no non-full stripe part in this I / O, it will not be appended to the end of the shard request. The end of the verification data node request saves the full amount of non-full stripe data that needs to be sent this time. When the write request sent is successful, the client's non-full stripe cache is updated to the latest non-full stripe data. The data reading stage includes the following steps: S1. Prioritize cache read, where if the read request covers the cache interval of the client's temporary cache, directly read the client cache; S2. If the read request does not cover the cache interval of the client's temporary cache, a partial read is attempted. The partial read means that if the read length is less than one stripe width, the read request is only sent to the relevant data shard node. If it is found that the data shard node related to the partial read is unavailable, the partial read method is abandoned and redundant read is performed directly. The partial read request is only sent to the data shard node; S3. If the length of the read request is greater than the width of a stripe, each data shard node needs to send a read request. In this case, a full stripe read is performed. Once the range of the full stripe read hits the cache interval, the hit part does not need to send a read request to the data shard node. S4, after the partial read or full stripe read request returns successfully, take out the data of the non-full stripe from the cache and append it to the end of the data returned successfully by the request; S5. If a partial read or full stripe read fails, a retry is performed. The retry read strategy is redundant read. Redundant read sends read requests to all data shard nodes and verification data nodes. When the redundant read request covers the cache interval of the client's temporary cache, the covered part uses cache read.

2. According to claim 1, a distributed storage system read optimization method based on erasure coding is characterized in that: The following steps are involved: According to the operation steps in S1, the consistency of the cache needs to be guaranteed by the write mark: when the stripe where the cache is located is being written, the data of the previous non-full stripe of the data in the cache needs to be read after the write is completed and updated.

3. According to claim 2, a distributed storage system read optimization method based on erasure coding is characterized in that: The following steps are involved: According to the operation steps in S2, if it is found that the data sharding node related to the partial read is unavailable, the partial read method is abandoned and redundant read is performed directly. The partial read request will only be sent to the data node, so there is no need to perform stripe alignment or EC decoding calculation.

4. According to claim 3, a distributed storage system read optimization method based on erasure coding is characterized in that: The following steps are involved: According to the operation steps in S3, in order to simplify the reading process and be able to perform EC decoding on the stripes, the read range will be stripe aligned, and then all available data shard node addresses will be obtained according to the cluster view, and read requests will be sent to all available data shard nodes.

5. According to claim 4, a distributed storage system read optimization method based on erasure coding is characterized in that: The following steps are involved: According to the operation steps in S3, the cluster management node will give priority to the data shard nodes with low load and availability. Full stripe reading will only perform erasure code decoding when there is an abnormality in the data shard node and the data of the verification data node is read. Otherwise, it only splices the data to minimize the overhead of EC decoding calculation.

6. The method for read optimization of a distributed storage system based on erasure coding according to claim 5, characterized in that: The method comprises the following steps: according to the operation steps in S5, S501. If a partial read or full stripe read fails, or if a related data shard node is found to be unavailable when a partial read is attempted, a redundant read is performed; S502, redundant read sends read requests to all data shard nodes, excluding nodes that have been found to be unavailable. As long as the number of successful read responses received meets the number of full stripe shards, data parsing begins; S503: If the responses to the number of full-strip fragments are all data fragment information, splice them in order; S504: If there is verification data sharding information, erasure code decoding is performed on the client; S505: If all read requests are successfully returned, the non-full stripe data in the client temporary cache is appended to the end of the successfully returned data; S506. If the data format does not meet the requirements during the data parsing process after the successful read response, the response with the data format not meeting the requirements is discarded, and the read request of other responses is continued to be waited for. After the number of full stripe shards is met, the parsing is continued until the complete data is parsed; S507: If the number of successful read responses is less than the number of full stripe shards, try to restore the data using replicas; S508: If all available read responses are data shard information, then splicing starts from the first data shard in sequence. Once the response information of the corresponding shard is missing, parsing is stopped, and the data finally read will be less than expected. S509: If the available read response contains the verification shard information, determine whether the read range falls entirely within the non-full stripe portion. If so, retrieve the required data from the verification shard response, which is the complete data expected to be read.

Citation Information

Patent Citations

  • Data storage method and system and storage medium

    CN111400083A

  • Erasure code data storage method and device, equipment and medium

    CN115268773A