A method and device for extending an erasure code storage system
By designing a new expansion mechanism in the erasure code storage system and updating the verification block with local data blocks, the problem of low expansion efficiency in the existing technology is solved, and a more efficient data transmission and expansion process is achieved.
Patent Information
- Application Number
- CN202111459202.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-02
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-12-02
AI Technical Summary
The existing erasure code storage system is less efficient in the expansion process, resulting in poor transmission parallelism and longer expansion process.
A new expansion mechanism is designed to reduce the amount of data transmission by utilizing locally stored data blocks to improve the expansion efficiency by performing data block migration and verification block updates in parallel.
It effectively reduces the data transmission of verification block updates, improves the expansion efficiency and effect, supports continuous expansion and reduces the adjustment overhead of the storage system during the next expansion.
Smart Images

Figure CN114237970B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of financial technology (Fintech), and in particular, to a method and device for extending an erasure code storage system. Background Art
[0002] With the development of computer technology, more and more technologies are being applied in the financial field. Traditional finance is gradually transforming into financial technology. However, due to the security and real-time requirements of the financial industry, higher requirements are also placed on technology.
[0003] Currently, storage systems are deployed on a large number of storage nodes and are the main backbone for supporting various upper-layer applications such as information retrieval, machine learning, video streaming, etc. In order to ensure the reliability of data in storage systems, storage systems often rely on replication and erasure coding technologies, both of which require storing additional data redundancy in advance so that the system can use redundancy to recover lost data. Compared with replication, erasure coding technology can achieve higher data reliability at the same storage overhead.
[0004] Moreover, with the continuous growth of data, higher requirements are placed on the scalability of the storage system. Specifically, the implementation of the storage expansion function requires the storage system to perform two operations: data relocation and check block update. However, in the solutions provided in the prior art, the expansion of the storage system during data relocation and check block update inevitably causes a large amount of data transmission, resulting in poor transmission parallelism and a long expansion process, that is, poor expansion efficiency and effect. Summary of the invention
[0005] The present invention provides a method and device for expanding an erasure code storage system, which are used to solve the problem of low expansion efficiency of the erasure code storage system in the prior art.
[0006] In a first aspect, the present invention provides a method for extending an erasure code storage system, the method comprising: determining data in the storage system, encoding the data, and dispersing and storing the data in each node to obtain spatial position distribution information of the each node; based on expansion requirement information, determining the number of newly added nodes on each stripe, and based on the newly added number of nodes and the spatial position distribution information, determining the extended node information on each stripe; the stripe comprises data blocks and check blocks having a coding relationship; based on the extended node information and the least common multiple rule, determining an extended group, and splitting the extended group to obtain a target group comprising a plurality of selected stripes; the extended group is composed of a plurality of stripes that meet the conditions that the expansion requirement can be completed and the spatial position distribution law remains unchanged; executing an expansion algorithm on the target group to obtain a corresponding target extended group, the target extended group comprising extended data blocks and extended check blocks.
[0007] In the above method, a new expansion mechanism is proposed to reduce traffic and explore the expansion mechanism of transmission parallelism in continuous scaling. In this expansion mechanism, a new stripe layout is designed, which uses locally stored data blocks to update the check blocks, thereby reducing the data transmission for checking block updates. Therefore, the data transmission for checking block updates can be reduced, thereby improving the expansion efficiency.
[0008] Optionally, the data is encoded and stored in various nodes in a dispersed manner to obtain spatial position distribution information of the various nodes, including: dividing the data into K data blocks of the same size; K is a positive integer greater than 1; performing intra-domain matrix operations on the K data blocks and a preset coding matrix to obtain M check blocks; M is a positive integer greater than 1 and less than K; the K data blocks and the M check blocks constitute multiple stripes; dispersing the data blocks and check blocks on the same stripe on different K+M nodes, determining the distribution information of the K data blocks and the M check blocks on each node, and obtaining the spatial position distribution information based on the distribution information.
[0009] In the above method, specific data processing and decentralized storage of data blocks and check blocks are provided. Based on this method, a good implementation basis can be provided for the subsequent expansion and update of check blocks and data blocks based on the new stripe layout, thereby improving the expansion efficiency.
[0010] Optionally, based on the newly added number of nodes and the spatial position distribution information, the extended node information on each stripe is determined, including: based on the spatial position distribution information, determining the first number of nodes storing data blocks on each stripe, and the second number of nodes storing check blocks on each stripe; adding the first number of nodes and the newly added number of nodes to obtain a third number of nodes, and using the third number of nodes as the number of expanded storage data blocks on each stripe; and using the second number of nodes as the number of expanded storage check blocks on each stripe to determine the extended node information on each stripe.
[0011] Based on the above method, the expansion node information on each stripe and the number of expanded storage data blocks and storage check blocks on each stripe can be accurately and quickly determined. In this way, a basis is provided for filling data in subsequent data blocks and check blocks, thereby quickly realizing the migration of data blocks and the update of check blocks, and improving the expansion efficiency.
[0012] Optionally, based on the extended node information and the least common multiple rule, an extended group is determined, and the extended group is split to obtain a target group including strips with corresponding relationships, including: based on the extended node information and the least common multiple rule, the extended group is determined; the extended group includes V extended strips; the V extended strips are split to determine P basic groups and R adjustment groups; each of the basic groups includes Vp basic strips, and each of the adjustment groups includes Vr adjustment strips; P and R are positive integers greater than 1; K basic strips are selected from the basic group, and D adjustment strips are selected from the adjustment group, and the target group is determined based on the K basic strips and the D adjustment strips; the target group includes K+D strips.
[0013] Based on the above method, the expansion group can be split to determine a basic group including strips that need to be updated according to the data blocks on the newly added nodes, and an adjustment group including strips that send data blocks to the basic group, so that rapid migration of data blocks and rapid update of check blocks can be achieved based on the adjustment group and the basic group.
[0014] Optionally, the least common multiple rule is determined using the following formula:
[0015] V=LCM(K,K+D+1)(K+D)(K+1) / K
[0016] The LCM() is used to represent the function of finding the least common multiple, k is used to represent the number of nodes storing data blocks on each stripe before expansion, and d is used to represent the number of newly added nodes.
[0017] Based on the above method, the number of stripes included in the extended group can be accurately and quickly determined.
[0018] Optionally, an expansion algorithm is executed on the target group to obtain a corresponding target expansion group, wherein the target expansion group includes an expansion data block and an expansion check block, including: numbering K+D stripes in any of the target groups, and numbering K+M+D nodes after the storage system is expanded; calculating the difference check blocks of the data blocks on the adjustment stripes in the first K+1 nodes, and updating the first check block of the basic stripe on the same node based on the difference check blocks; transmitting the data blocks on the adjustment stripes to the basic stripe on the same node in a round-robin mode to obtain an initial expansion group after expansion; and performing preset operations on the initial expansion group to obtain a corresponding target expansion group.
[0019] Based on the above method, the data block migration and check block update of the erasure code storage system expansion are performed in parallel. That is, during the expansion process, some nodes are scheduled to perform data block migration operations, and at the same time, transmission tasks are assigned to another part of the nodes to perform check block update operations. In this way, the expansion efficiency can be improved.
[0020] Optionally, after obtaining the target extension group, the method further includes: determining the logical relationship of the strips corresponding to the target extension group, and the first spatial distribution information corresponding to each extension data block and the extension check block; and adjusting the order of the logical relationship according to the spatial distribution information so that the first spatial distribution information is the same as the logical layout of the spatial distribution information.
[0021] Based on the above method, it is possible to support the erasure code storage system to perform the next expansion without adjusting the spatial distribution, thereby reducing unnecessary overhead. In addition, a function of supporting the continuous expansion of the erasure code storage system is also provided.
[0022] In a second aspect, the present invention provides a device for extending an erasure code storage system, the device comprising: a first processing unit, for determining data in the storage system, encoding the data, and dispersively storing the data in each node to obtain spatial position distribution information of the each node; a second processing unit, for determining the number of newly added nodes on each stripe based on expansion requirement information, and determining the extended node information on each stripe based on the newly added number of nodes and the spatial position distribution information; the stripe comprises data blocks and check blocks having a coding relationship; a third processing unit, for determining an extended group based on the extended node information and a least common multiple rule, and splitting the extended group to obtain a target group comprising a plurality of selected stripes; the extended group is composed of a plurality of stripes that meet the conditions that the expansion requirement can be completed and the spatial position distribution law remains unchanged; an obtaining unit, for executing an expansion algorithm on the target group to obtain a corresponding target extended group, the target extended group comprising an extended data block and an extended check block.
[0023] Optionally, the first processing unit is used to: divide the data into K data blocks of the same size; K is a positive integer greater than 1; perform intra-domain matrix operations on the K data blocks and a preset coding matrix to obtain M check blocks; M is a positive integer greater than 1 and less than K; the K data blocks and the M check blocks constitute multiple stripes; disperse the data blocks and check blocks on the same stripe on different K+M nodes, determine the distribution information of the K data blocks and the M check blocks at each node, and obtain the spatial position distribution information based on the distribution information.
[0024] Optionally, the second processing unit is used to: determine the number of first nodes storing data blocks on each stripe and the number of second nodes storing check blocks on each stripe based on the spatial position distribution information; add the first number of nodes and the newly added number of nodes to obtain a third number of nodes, and use the third number of nodes as the number of expanded storage data blocks on each stripe; and use the second number of nodes as the number of expanded storage check blocks on each stripe to determine the extended node information on each stripe.
[0025] Optionally, the third processing unit is used to: determine an extended group based on the extended node information and the least common multiple rule; the extended group includes V extended strips; split the V extended strips to determine P basic groups and R adjustment groups; each of the basic groups includes Vp basic strips, and each of the adjustment groups includes Vr adjustment strips; P and R are positive integers greater than 1; select K basic strips from the basic group and D adjustment strips from the adjustment group, and determine a target group based on the K basic strips and the D adjustment strips; the target group includes K+D strips.
[0026] Optionally, the least common multiple rule is determined using the following formula:
[0027] V=LCM(K,K+D+1)(K+D)(K+1) / K
[0028] The LCM() is used to represent the function of finding the least common multiple, k is used to represent the number of nodes storing data blocks on each stripe before expansion, and d is used to represent the number of newly added nodes.
[0029] The optional acquisition unit is specifically used to: number the K+D stripes in any of the target groups, and number the K+M+D nodes after the storage system is expanded; calculate the difference check blocks of the data blocks on the adjusted stripes in the first K+1 nodes, and update the first check block of the basic stripe on the same node based on the difference check blocks; transfer the data blocks on the adjusted stripes to the basic stripe on the same node in a round-robin mode to obtain the initial extended group after expansion; perform preset operations on the initial extended group to obtain the corresponding target extended group.
[0030] Optionally, the device also includes an adjustment unit, which is used to: determine the logical relationship of the strips corresponding to the target extension group, and the first spatial distribution information corresponding to each extended data block and the extended check block; according to the spatial distribution information, adjust the order of the logical relationship so that the first spatial distribution information is the same as the logical layout of the spatial distribution information.
[0031] The beneficial effects of the above-mentioned second aspect and each optional device of the second aspect can refer to the beneficial effects of the above-mentioned first aspect and each optional method of the first aspect, and will not be repeated here.
[0032] In a third aspect, the present invention provides a computer device, including a program or an instruction, which, when executed, is used to execute the above-mentioned first aspect and each optional method of the first aspect.
[0033] In a fourth aspect, the present invention provides a storage medium, comprising a program or an instruction, which, when executed, is used to execute the above-mentioned first aspect and each optional method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below.
[0035] Figure 1 Schematic diagram of the data block migration phase of the process of extending the erasure code RS(2,1,4) for traditional storage systems;
[0036] Figure 2 Schematic diagram of the check block update phase of the process of extending the erasure code RS(2,1,4) for traditional storage systems;
[0037] Figure 3 A schematic diagram of an optional application scenario provided for an embodiment of the present invention; Figure 4 A schematic diagram of the architecture of an optional erasure code storage system provided by an embodiment of the present invention;
[0038] Figure 5 A schematic flow chart of the steps of a method for extending an erasure code storage system provided by an embodiment of the present invention;
[0039] Figure 6 A schematic diagram of the encoding process of the erasure code RS(k,m) in a stripe provided by an embodiment of the present invention;
[0040] Figure 7 A schematic diagram of erasure code storage distribution of erasure code RS(2,2) and erasure code RS(3,2) in an erasure code storage system provided in an embodiment of the present invention;
[0041] Figure 8 A schematic diagram of a check block update and data block relocation parallelism algorithm for erasure code RS (2, 1, 4) provided in an embodiment of the present invention;
[0042] Fig. 9 A schematic diagram of a workflow diagram of an extension process of an erasure code (2, 2, 3) provided in an embodiment of the present invention;
[0043] Fig.10 A schematic diagram of a result graph of an experiment on the influence of different bandwidths provided by an embodiment of the present invention;
[0044] Fig.11 A schematic diagram of a result graph of an experiment on the impact of data blocks of different sizes provided by an embodiment of the present invention;
[0045] Fig.12 A schematic diagram of a result graph of an experiment on the impact of different numbers of newly added nodes on a test provided by an embodiment of the present invention;
[0046] Fig.13 A schematic diagram of a result diagram of a numerical analysis test of an extended process flow rate under a general configuration of an erasure code storage system provided in an embodiment of the present invention;
[0047] Fig.14 A schematic diagram of a numerical analysis experimental result diagram of the impact of different numbers of newly added nodes on traffic bandwidth during the expansion process of an erasure code storage system provided by an embodiment of the present invention;
[0048] Fig.15 A schematic diagram of a result diagram of a numerical analysis of bandwidth utilization during different expansion processes of an erasure code storage system provided by an embodiment of the present invention;
[0049] Fig.16 A schematic diagram of the structure of an apparatus for extending an erasure code storage system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0050] In order to better understand the above-mentioned technical scheme, the above-mentioned technical scheme will be described in detail below in conjunction with the accompanying drawings and specific implementation methods of the specification. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical scheme of the present invention, rather than limitations on the technical scheme of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments may be combined with each other.
[0051] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the images used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0052] To facilitate understanding of the technical solution provided by the embodiment of the present invention, some key terms or processes used in the embodiment of the present invention are explained here:
[0053] 1. Erasure Code (EC): It is a forward error correction technology (FEC) that can add m copies of data to n copies of original data, and can restore the original data through any n copies of data among n+m copies. That is, if any data less than or equal to m copies fails, it can still be restored through the remaining data. It is mainly used to avoid packet loss in network transmission, and storage systems use it to improve storage reliability.
[0054] 2. There are three main types of applications of erasure code technology in distributed storage systems: array erasure code, RS (Reed-Solomon) erasure code and LDPC (Low Density Parity Check Code) erasure code. In the embodiments of the present invention, the expansion of RS erasure code corresponding to distributed storage system is mainly described.
[0055] The design concept of the embodiment of the present invention is briefly introduced below:
[0056] See also Figure 1 As shown, it is a schematic diagram of the data block migration phase of the traditional storage system with extended erasure code parameters (2,1,4) in the prior art. And, please refer to Figure 2 As shown, it is a schematic diagram of the check block update phase of the traditional storage system extended erasure code with parameters (2,1,4) in the prior art. Figure 1 and Figure 2 The S in the figure is used to represent the stripe, N is used to represent the node, D is used to represent the data block, and P is used to represent the check block.
[0057] Obviously, the migration of data blocks and the updating of check blocks in the prior art inevitably cause a large amount of data transmission, resulting in poor transmission parallelism and a long expansion process, that is, poor expansion efficiency and effect.
[0058] In view of this, the present invention provides a method for extending an erasure code storage system, which proposes a new extension mechanism, the purpose of which is to reduce traffic and explore an extension mechanism for transmission parallelism in continuous scaling. In this extension mechanism, a new stripe layout is designed, which uses locally stored data blocks to update the check blocks, thereby reducing the data transmission for checking block updates. It can be seen that the method for extending an erasure code storage system provided by the present invention can reduce the data transmission for checking block updates, thereby improving the expansion efficiency.
[0059] After introducing the design concept of the embodiment of the present invention, the following briefly introduces the application scenarios to which the technical solution of the extended erasure code storage system in the embodiment of the present invention is applicable. It should be noted that the application scenario described in the embodiment of the present invention is to more clearly illustrate the technical solution of the embodiment of the present invention, and does not constitute a limitation on the technical solution provided by the embodiment of the present invention. Ordinary technicians in this field can know that with the emergence of new application scenarios, the technical solution provided by the embodiment of the present invention is also applicable to similar technical problems.
[0060] The storage system provided in the embodiment of the present invention can be applied to most storage systems that need to perform storage expansion functions. The storage system is, for example, a business order storage system, or a transaction data storage system, etc. Figure 3 As shown, a scenario diagram provided by an embodiment of the present invention. In the scenario diagram, multiple electronic devices 301 are respectively deployed with proxy nodes and a metadata server 302 is deployed with a global coordinator. The electronic device 301 can communicate with the metadata server 302 of the global coordinator, for example, directly or indirectly connected through wired or wireless communication, and the present invention does not limit it. Among them, electronic devices 301-1, electronic devices 301-2, ..., electronic devices 301-n can be deployed by different proxy nodes.
[0061] In the embodiment of the present invention, the electronic device 301 may be a server, for example, but is not limited thereto. Each electronic device 301 may include one or more processors 3011, a memory 3012, and an I / O interface 3013 for interacting with other servers.
[0062] In an embodiment of the present invention, the metadata server 302 deployed with a global coordinator can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0063] In this scenario, the metadata server 302 deployed with a global coordinator is responsible for managing the metadata of the stripe. In addition, each round of transmission tasks can be issued to each electronic device 301 to perform data block migration or check block update operations. When each electronic device 301 transmits this confirmation signal to the metadata server 302 deployed with a global coordinator, the metadata server 302 deployed with a global coordinator can execute the next round of transmission commands to each electronic device 301.
[0064] In this scenario, each electronic device 301 needs to receive the transmission command sent by the coordinator, parse the transmission command and execute the task content of the transmission command. Specifically, after each electronic device 301 sends the data block or check block to be sent to the corresponding electronic device 301, the electronic device 301 will send a confirmation signal to the metadata server 302 deployed with the global coordinator, informing the metadata server 302 deployed with the global coordinator that the transmission is completed, so that it can prepare to execute the next round of transmission commands.
[0065] See also Figure 4 , which is a schematic diagram of the architecture of the erasure code storage system provided by an embodiment of the present invention. The metadata server can issue a command to update the check block in the proxy node in the existing node, and issue a command to migrate the data block of the proxy node in the existing node to the proxy node in the newly added node.
[0066] Of course, the method provided in the embodiment of the present invention is not limited to Figure 1 The application scenarios shown may also be used in other possible application scenarios, and the embodiments of the present invention are not limited thereto.
[0067] To further illustrate the method for extending the erasure code storage system provided by the embodiment of the present invention, a detailed description is given below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiment of the present invention provides method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or no creative labor. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiment of the present invention. In the actual processing process or when the device is executed, the method can be executed in the method sequence shown in the embodiment or drawings or in parallel (for example, in an application environment of a parallel processor or multi-threaded processing).
[0068] The following combination Figure 5 The method flow chart shown illustrates the method for extending the erasure code storage system in the embodiment of the present invention. The method flow of the embodiment of the present invention is introduced below.
[0069] See also Figure 5 As shown, it is an implementation flow chart of the method for extending the erasure code storage system provided by an embodiment of the present invention. The method can be executed by a metadata server. The specific implementation process is as follows:
[0070] Step 501: determine the data in the storage system, encode the data, and store the data in various nodes in a dispersed manner to obtain the spatial location distribution information of each node.
[0071] In an embodiment of the present invention, the metadata server can select an RS-type erasure code that meets the fault tolerance requirements and storage efficiency of the storage system based on the reliability requirements and storage overhead limitations of the storage system, and use the RS-type erasure code as data in the storage system.
[0072] In an embodiment of the present invention, the metadata server may divide the data into K data blocks of the same size, wherein K is a positive integer greater than 1. Then, the K data blocks may be subjected to an intra-domain matrix operation with a preset coding matrix to obtain M check blocks, wherein M is a positive integer greater than 1 and less than K. In addition, the K data blocks and the M check blocks may constitute a plurality of stripes.
[0073] Furthermore, the data blocks and check blocks on the same stripe may be dispersed on different K+M nodes, the distribution information of the K data blocks and the M check blocks on each node may be determined, and the spatial position distribution information may be obtained based on the distribution information.
[0074] In the embodiment of the present invention, the parameters of the RS erasure code include three parameters, which are represented by K, M, and W, for example, where K indicates that the RS erasure code has K data blocks, M indicates that the RS erasure code has M check blocks, and W is used to characterize the number of bits corresponding to the RS erasure code; where W can generally be 4, 8, 16, or 32. In the embodiment of the present invention, w=8 is used as an example for explanation below.
[0075] In an embodiment of the present invention, the metadata server may obtain a check block based on the parameters of the RS erasure code and a preset coding matrix. Specifically, the metadata server may perform intra-domain matrix operations on the above-mentioned K data blocks and the generated preset coding matrix limited to the Galois field, thereby obtaining M check blocks. Exemplarily, the check block may be obtained by performing bitwise operations on the data blocks and the numbers of the preset coding matrix. The aforementioned preset coding matrix may be a Vandermonde matrix or a Cauchy matrix, which is not limited in the embodiment of the present invention.
[0076] For example, see Figure 6 As shown, Figure 6 A schematic diagram of the process of encoding RS erasure codes provided in an embodiment of the present invention. The metadata server can determine a coding matrix based on a unit matrix and a generator matrix, and determine the coding matrix as a preset coding matrix, and then multiply the preset coding matrix by k data blocks to obtain m check blocks, so that the k data blocks and the m data blocks can be stored on k+m nodes.
[0077] In the embodiment of the present invention, according to the parameters of the erasure code and the preset coding matrix, K data blocks are coded to generate M corresponding check blocks, which are represented by a binary group (K, M).
[0078] In the embodiment of the present invention, after determining the data block and the check block having a coding relationship with the data block, the stripe can be determined based on the data block and the check block having a coding relationship with the data block. Further, the data blocks and check blocks of the same stripe can be dispersed and stored in different nodes.
[0079] Specifically, in an embodiment of the present invention, a spatial distribution scheme for dispersing data blocks and check blocks on the same stripe on different K+M nodes is set as follows: 1 check block and K data blocks are stored on the first K+1 nodes; wherein the positions of the check blocks are arranged diagonally on the K+1 stripes; and check blocks other than 1 check block are stored on the last M-1 nodes.
[0080] In an embodiment of the present invention, the metadata server can determine the spatial location distribution information according to the distribution of data blocks and check blocks in different nodes. Specifically, the spatial location distribution information can be understood as the location information of data blocks in stripes and nodes, and the location information of check blocks in stripes and nodes.
[0081] For example, see Figure 7 As shown, Figure 7 The parameters of the RS erasure code provided in the embodiment of the present invention are (2, 2) and (3, 2). S is used to represent the stripe, N is used to represent the node, D is used to represent the data block, and P is used to represent the check block. Figure 7 ,, the storage distribution of the data blocks and the check blocks corresponding to the data blocks in the embodiment of the present invention can be clearly known.
[0082] Step 502: Based on the expansion requirement information, determine the number of newly added nodes on each stripe, and based on the number of newly added nodes and spatial position distribution information, determine the extended node information on each stripe; the stripe includes data blocks and check blocks having a coding relationship.
[0083] In an embodiment of the present invention, based on the spatial position distribution information, the number of first nodes storing data blocks on each stripe and the number of second nodes storing check blocks on each stripe are determined; the number of first nodes and the number of newly added nodes are added to obtain the third number of nodes, and the third number of nodes is used as the number of expanded storage data blocks on each stripe; and the number of second nodes is used as the number of expanded storage check blocks on each stripe to determine the extended node information on each stripe.
[0084] Exemplarily, based on the system adjustment reliability and stripe length requirements, the number of newly added nodes is determined to be D, the number of first nodes storing data blocks on each stripe is K, and the number of second nodes storing check blocks on each stripe is M, so that the number of expanded storage data blocks on each stripe can be determined to be K+D, and the number of expanded storage check blocks on each stripe can be determined to be M.
[0085] Step 503: Based on the extended node information and the least common multiple rule, determine the extended group, and split the extended group to obtain a target group including multiple selected strips; the extended group is composed of multiple strips that meet the conditions of being able to complete the expansion requirements and the unchanged spatial position distribution law.
[0086] In an embodiment of the present invention, the metadata server can determine an extension group based on the extension node information and the least common multiple rule, wherein the extension group is composed of multiple stripes that meet the conditions that the extension requirement can be met and the spatial position distribution law remains unchanged, and the extension group includes V extension stripes. It can be seen that V extension stripes can meet the conditions that the extension requirement and the spatial position distribution law remain unchanged.
[0087] Specifically, the least common multiple rule is determined using the following formula:
[0088] V=LCM(K,K+D+1)(K+D)(K+1) / K
[0089] Wherein, the LCM() is used to represent the function of obtaining the least common multiple, K is used to represent the number of nodes storing data blocks on each stripe before expansion; and D is used to represent the number of newly added nodes.
[0090] In an embodiment of the present invention, after the metadata server determines V extended strips, it can split the V extended strips to determine P basic groups and R adjustment groups; each basic group includes Vp basic strips, and each adjustment group includes Vr adjustment strips; P and R are positive integers greater than 1.
[0091] In an embodiment of the present invention, the metadata server can represent the V stripes into two types of groups, namely the aforementioned basic group and the adjusted group, according to the different functions of the data blocks and the check blocks in the V extended stripes. Among them, the check blocks of the stripes in the basic group need to be updated according to the data blocks on the newly added nodes, and the data blocks of the stripes in the adjusted group are sent to the stripes in the basic group.
[0092] Specifically, Vp basic strips and Vr adjustment strips need to satisfy the equation: Vp:Vr=K:D, so Vp and Vr can be determined based on the following formula:
[0093] Vp = LCM(K, K + D + 1)(K + 1); Vr = LCM(K, K + D + 1)D(K + 1) / K.
[0094] In an embodiment of the present invention, during the expansion process, the data blocks of the adjustment group can be correspondingly transmitted to the metadata server in the stripe of the stripe basic group with a corresponding relationship. Specifically, according to the stripe order, it can be determined that K(K + 1) basic stripes in the basic group correspond to D(K + 1) stripes in the adjustment group. Therefore, the corresponding relationship can be expressed as: in the basic group: {(i - 1)K(K + 1) + 1, (i - 1)K(K + 1) + 2,..., iK(K + 1)}, these K(K + 1) stripes, and in the adjustment group: {(i - 1)D(K + 1) + 1, (i - 1)D(K + 1) + 2,..., iD(K + 1)} these D(K + 1) stripes have a corresponding relationship. Where 0 < i < LCM(K, K + D + 1) / K. It can be seen that there are LCM(K, K + D + 1) / K pairs of corresponding stripes in every V stripes.
[0095] In an embodiment of the present invention, after determining P basic groups and R adjustment groups and their respective corresponding stripes, K basic stripes can be selected from the K(K + 1) stripes in the basic group according to the stripe order for each group. And, from the D(K + 1) stripes in the adjustment group corresponding to the K(K + 1) stripes, D adjustment stripes can be continuously selected at intervals of K adjustment stripes according to the stripe order. Therefore, the above (K + D)(K + 1) stripes can be formed into (K + 1) small groups, so as to determine the target group based on K basic stripes and D adjustment stripes; where each target group includes K + D stripes.
[0096] It can be seen that any target group includes: independently selecting K basic stripes of {(i - 1)K + 1, (i - 1)K + 2,..., iK} from the K(K + 1) stripes in the basic group, and D adjustment stripes selected from {D(K + 1) - i, (D - 1)(K + 1) - i,..., (K + 1) - i} stripes in the adjustment group having a corresponding relationship with it.
[0097] In an embodiment of the present invention, after determining the target group, step 504 is executed for each target group: performing an expansion algorithm on the target group to obtain a corresponding target expansion group, and the target expansion group includes expansion data blocks and expansion check blocks.
[0098] Specifically, determining the corresponding target expansion group can adopt but is not limited to the following steps:
[0099] Step a: Number the K + D stripes in any target group, and number the K + M + D nodes after the storage system is expanded.
[0100] In the embodiment of the present invention, K+D stripes in any target group may be numbered as {1, 2, ..., K+D}, and K+M+D nodes may be numbered as {1, 2, ..., K+M+D}.
[0101] Step b: Calculate the difference check blocks of the data blocks on the adjustment stripes in the first K+1 nodes, and update the first check block of the basic stripe on the same node based on the difference check blocks.
[0102] In the embodiment of the present invention, the first check block of the first K stripes can be updated first. Specifically, the check block can be updated by calculating the difference check block. Specifically, the difference check blocks of the data blocks on the {K+1, K+2, ..., K+D} adjustment stripes in the first K+1 nodes can be calculated, and the first check block of the basic stripe on the same node can be updated based on the difference check block.
[0103] Step c: The data blocks on the adjusted stripe are transferred to the basic stripe on the same node in a round-robin mode to obtain the expanded initial extension group.
[0104] In the embodiment of the present invention, for the data block migration part, the data blocks on the adjustment stripes {K+1, K+2, ..., K+D} can be transmitted in a round-robin manner to the basic stripes {1, 2, ..., K}. Specifically, in each round, D data blocks on D nodes are selected from the K nodes storing data blocks in order. In the i-th round, D nodes {i, i+1, ..., i+D-1} are selected, and the D data blocks on the D adjustment stripes {K+1, K+2, ..., K+D} are correspondingly transmitted to D newly added nodes. When i+D-1>K, starting from the first node storing data blocks, the node selection continues.
[0105] Step d: Perform a preset operation on the initial expansion group to obtain a corresponding target expansion group.
[0106] In an embodiment of the present invention, the number of data blocks transmitted to the new node is counted by means of a global counter. After each (K+1)D data blocks are sent to the newly added node, the first check block of the subsequent D basic strips and the check block of the corresponding node adjustment stripe are logically replaced.
[0107] Specifically, the above process can be expressed as: in the i-th basic stripe that undergoes logical position replacement, the first check block of the stripe and the i-th data block of the corresponding node adjustment stripe are replaced, and the above data block migration algorithm is also executed after the replacement. The value range of i is greater than 0 and less than D.
[0108] It can be seen that after K rounds, when executing the data block migration algorithm, only D nodes of the K nodes that originally stored the data blocks were occupied in each round, and the remaining (KD) nodes were idle.
[0109] In the embodiment of the present invention, the following operations may be performed to update the (M-1) check blocks in the baseband stripe except the first check block:
[0110] Step 1: There are M nodes in the adjustment stripe storing check blocks, of which (M-1) nodes overlap with the nodes that only store check blocks in the basic stripe. In each round, the linear combination of (M-1) check blocks is transmitted from these (M-1) nodes to other (M-1) nodes.
[0111] Specifically, the process can be expressed as follows: in the i-th round, the (M-1) nodes transmit a linear combination of the check blocks to the nodes that are i positions away from each other. When the position of the node exceeds (M-1), starting from the first node, continue to select nodes for transmission. It can be seen that it ends after (M-2) rounds, where the value range of i is greater than 0 and less than M-2.
[0112] Step 2: There is one storage check block node and (KM) storage data block nodes in the adjustment stripe that do not overlap with the remaining (M-1) storage check block nodes in the basic stripe. In each round, (M-1) nodes are selected from these (K-M+1) nodes to transmit the linear combination of data blocks or the linear combination of check blocks to the remaining (M-1) storage check block nodes in the basic stripe.
[0113] Specifically, the process can be expressed as follows: in the i-th round, select {i, i+1, ..., i+M-1} nodes from the (K-M+1) nodes to transmit the linear combination of data blocks or the linear combination of check blocks to the nodes corresponding to {1, 2, ..., M-1} basic stripes storing check blocks. It can be seen that the process ends after (K-M+1) rounds. The value range of i is greater than 0 and less than K-M+1.
[0114] In summary, after (K-1) rounds, the update check block operation is completed by collecting the required linear combination of data blocks and the linear combination of check blocks, and the update check block operation can be completed by the update algorithm designed by the present invention.
[0115] It should be noted that, in the embodiment of the present invention, the aforementioned restriction condition on the update part of the check block is: each round of data block migration requires D storage data block nodes and the maximum (M-1) storage data block nodes required in Step 2 for updating the check block cannot overlap, that is, the inequality is required to be satisfied: K is greater than or equal to D+M-1.
[0116] In the embodiment of the present invention, it is necessary to limit each node to full-duplex transmission and reception, that is, in each round, each node can only receive and send one block at the same time, and it is necessary to maximize the node utilization rate in each round based on a preset algorithm. In the actual implementation process, when determining the number of new nodes, the data center manager needs to determine that the parameters K, M, and D meet this restriction condition.
[0117] In an embodiment of the present invention, the last (M-1) nodes in the basic stripe that only store check blocks, after receiving the linear combination of (KM) data blocks and the linear combination of (M-1) check blocks from other nodes and the linear combination of check blocks on the adjusted stripe calculated on the own node, can calculate the difference check blocks corresponding to the last (M-1) check blocks on the basic stripe through the erasure code decoding algorithm, and then perform an XOR operation on the check blocks of the basic stripe and the calculated difference check blocks, so as to calculate the updated check blocks, namely the extended check blocks.
[0118] In the embodiment of the present invention, when all stripes of an expansion group are updated, the corresponding target expansion group can be obtained. In addition, the logical relationship order of the stripes can be adjusted according to the spatial distribution information to meet the overall spatial distribution plan before the expansion. In this way, when the storage system performs the next expansion, there is no need to adjust the spatial distribution to bring unnecessary overhead.
[0119] In the embodiment of the present invention, please refer to Figure 8 , Figure 8 This is a schematic diagram of the RS (2,1,4) check block update and data block relocation parallelism algorithm provided by the present invention. Fig. 9 , Fig. 9 A schematic diagram of the process of extending RS (2, 1, 4) provided in an embodiment of the present invention.
[0120] In the embodiment of the present invention, the logical relationship of the stripes corresponding to the target expansion group and the first spatial distribution information corresponding to each expansion data block and the expansion check block can also be determined. Then, according to the spatial distribution information, the logical relationship order is adjusted so that the logical layout of the first spatial distribution information is the same as that of the spatial distribution information. In this way, when the storage system performs the next expansion, there is no need to adjust the spatial distribution to bring unnecessary overhead.
[0121] In the embodiment of the present invention, when the stripes of all basic groups have completed the expansion operation, all data blocks and check blocks of the adjustment group in the storage system can also be deleted, so as to reduce the waste and consumption of resources as much as possible.
[0122] It can be seen that the method for extending an erasure code storage system provided by an embodiment of the present invention has a low input / output, or I / O, overhead during the expansion process, which reduces the amount of data that needs to be read and written during the expansion process and the amount of data transmitted in the network. In addition, the time delay of the expansion process is short. On the basis of low I / O overhead and full-duplex communication, a new update check block algorithm increases the available bandwidth resources in the storage system, and reduces the time delay of the expansion process by executing the scheduling expansion algorithm in parallel. In addition, continuous expansion can also be supported, that is, after a single expansion process is completed, the overall spatial distribution of the present invention is consistent with that before the expansion. Therefore, when the storage system performs the next expansion, there is no need to adjust the spatial distribution to bring unnecessary overhead.
[0123] In the specific implementation process, the solution proposed in the embodiment of the present invention was tested. Specifically, the solution proposed in the embodiment of the present invention was tested in two testing methods: real platform research and simulation experiment.
[0124] Method 1: Testing the solution proposed in the embodiment of the present invention based on a real platform.
[0125] In an embodiment of the present invention, the specific experimental environment includes 19 virtual servers of the ecs.g6.large type, each of which is equipped with 2vCPU (2.5GHz Intel Xeon Platinum 8269CY) and 8GB of memory. And 40GB of storage, the operating system running is Ubuntu18.04. The maximum network bandwidth between any two servers is about 3Gb / s. One of the 19 servers serves as a global coordinator, and the remaining 18 servers are agents running the server program of the present invention. Among them, the default setting of the experiment is that the block size is 64MB, the erasure code scheme is RS (6,3) and RS (10,4), and the number of new nodes varies according to different experiments.
[0126] Specifically, each experiment was repeated multiple times, and the measured parameter was the time consumption of the expansion process, that is, the time it takes to transmit all blocks to the corresponding nodes. Specifically, the expansion time is defined as the average stripe expansion time consumption. The shorter the average expansion time consumption, the higher the efficiency of the expansion process.
[0127] In addition, the test adopts a comparative experiment to compare two advanced erasure code storage system expansion mechanisms, Scale-RS and NCScale. In actual implementation, it can also be tested or used in other experimental test environments, and the embodiment of the present invention does not limit this.
[0128] In the specific implementation process, the expansion time when the network bandwidth changes from 1Gb / s to 2Gb / s can be measured. Specifically, the test results are as follows: Fig.10 See Fig.10 Among the three expansion mechanisms, the method provided by the embodiment of the present invention requires the least expansion traffic and improves the parallelism of transmission compared with the other two mechanisms. In general, when the network bandwidth is 1Gb / s, the present invention reduces 49.8% and 58.9% on average compared with Scale-RS and NCScale. And, when the network bandwidth increases to 2Gb / s, the average reduction is 50.8% and 58.8% respectively.
[0129] Obviously, when the bandwidth increases, the average expansion time of the present invention is less than that of Scale-RS and NCScale. It can be seen that the expansion performance of the method provided by the present invention is better than that of Scale-RS and NCScale.
[0130] In the specific implementation process, different block sizes can also be tested, such as the expansion time from 32MB to 64MB. During this test, the network bandwidth can be set to 3Gb / s. Fig.11 As shown in the figure, the expansion time increases with the increase of block size, and the method provided by the present invention shortens the scaling time by 49.1-53.0% and 24.1-76.9% respectively compared with Scale-RS and NCScale. In addition, it can be seen that the method provided by the present invention and Scale-RS have achieved quite stable performance in the continuous expansion process, while the scaling time of NCScale increases significantly in the second expansion operation, i.e., the (8,3,10) expansion process.
[0131] In the specific implementation process, the impact of the number of newly added nodes (i.e., the number of newly added nodes mentioned above, i.e., parameter D) on the scaling time can also be tested and studied. Specifically, the network bandwidth can be fixed at 3Gb / s, and the parameter D can be studied from 2 to 3. Please refer to Figure 12. Under different numbers of newly added nodes, the average expansion time of the three mechanisms is not significantly affected. The most fundamental reason is that we make all methods have transmission parallelism to achieve fair comparison, so that the newly added D nodes can receive the migrated data in parallel. The method provided by the present invention reduces the expansion time of the Scale-RS and NCScale mechanisms by 49.8-51.4% and 23.6-76.3%, respectively, significantly improving the expansion efficiency.
[0132] Method 2: Testing the solution proposed in the embodiment of the present invention based on simulation testing.
[0133] In the specific implementation process, a traffic simulation test under a general configuration can be performed. For example, see Fig.13As shown, the test evaluates the traffic of the successive expansion process of different expansion mechanisms, and considers two cases, RS (6, 3) and RS (10, 4), and sets the value of parameter d to 2.
[0134] Please continue reading Fig.13 It can be seen that under different expansion process parameters, the solution provided by the present invention performs well in the continuous expansion process. Compared with Scale-RS, when expanding from RS(6,3) and RS(10,4), the solution provided by the present invention reduces the expansion traffic by 22.9-26.7% and 19.4-21.7% respectively. Compared with NCScale, the expansion traffic is reduced by 8.3% to 62.8%, that is, the solution provided by the present invention reduces resource consumption.
[0135] In a specific implementation process, a simulation test can be performed to determine the effect of the number of expansion nodes on the expansion bandwidth. The experiment measures the effect of adding different numbers of nodes on the expansion process efficiency of the present invention. For example, see Fig.14 As shown, the two parameter cases of RS (6, 3) and RS (10, 4) are used before expansion, and then the number of newly added nodes (i.e., parameter D) is changed from 2 to 10. It can be seen that the expansion traffic increases with the number of new nodes added. The reason for this is that adding more nodes requires the transmission of more blocks for relocation and check block update. However, the solution provided by the present invention still maintains the advantage of compressed expansion traffic. Compared with Scale-RS and NCScale, the expansion process of the present invention can reduce the expansion traffic by 35.2% and 38.1% respectively on average, that is, the solution provided by the present invention reduces resource consumption.
[0136] In the specific implementation process, different simulation experiments of average bandwidth utilization can be performed. The experiment ultimately evaluates the average bandwidth utilization, and the average bandwidth utilization is defined as the ratio of the average amount of data transmitted per time unit to the theoretical maximum amount of data that can be transmitted per time unit for data block relocation and check block update.
[0137] See also Fig.15 As shown, it can be seen that compared with Scale-RS and NCScale, the solution provided by the present invention achieves a near-optimal bandwidth utilization. In particular, the solution provided by the present invention achieves a bandwidth utilization of 96.7% in the expansion process RS(18,4,20). On average, the bandwidth utilization of the present invention is 41.7-46.7% and 61.9-78.3% higher than that of Scale-RS and NCScale, respectively.
[0138] In summary, the present invention proposes a fast and continuous expansion mechanism to address the phenomenon that the expansion process of the erasure code storage system has high I / O consumption, low bandwidth utilization, and increased consumption due to continuous expansion. The present invention analyzes the expansion process of the erasure code storage system from a continuous perspective, and designs a new spatial distribution scheme and check block update algorithm to increase node bandwidth utilization and block transmission execution. The present invention reduces the expansion process time and bandwidth flow consumption on the basis of ensuring system reliability.
[0139] like Fig.16 As shown, the present invention provides a device for extending an erasure code storage system, the device comprising: a first processing unit 1601, used to determine the data in the storage system, encode the data, and store the data in each node in a dispersed manner to obtain spatial position distribution information of the nodes; a second processing unit 1602, used to determine the number of newly added nodes on each stripe based on the extension requirement information, and determine the extended node information on each stripe based on the newly added number of nodes and the spatial position distribution information; the stripe includes a data block and a check block having a coding relationship; a third processing unit 1603, used to determine an extended group based on the extended node information and the least common multiple rule, and split the extended group to obtain a target group including a plurality of selected stripes; the extended group is composed of a plurality of stripes that meet the conditions that the extension requirement can be completed and the spatial position distribution law remains unchanged; an obtaining unit 1604, used to execute an extension algorithm on the target group to obtain a corresponding target extended group, the target extended group including an extended data block and an extended check block.
[0140] Optionally, the first processing unit 1601 is used to: divide the data into K data blocks of the same size; K is a positive integer greater than 1; perform intra-domain matrix operations on the K data blocks and a preset coding matrix to obtain M check blocks; M is a positive integer greater than 1 and less than K; the K data blocks and the M check blocks constitute multiple stripes; disperse the data blocks and check blocks on the same stripe on different K+M nodes, determine the distribution information of the K data blocks and the M check blocks at each node, and obtain the spatial position distribution information based on the distribution information.
[0141] Optionally, the second processing unit 1602 is used to: determine the number of first nodes storing data blocks on each stripe and the number of second nodes storing check blocks on each stripe based on the spatial position distribution information; add the first number of nodes and the newly added number of nodes to obtain a third number of nodes, and use the third number of nodes as the number of expanded storage data blocks on each stripe; and use the second number of nodes as the number of expanded storage check blocks on each stripe to determine the extended node information on each stripe.
[0142] Optionally, the third processing unit 1603 is used to: determine an extended group based on the extended node information and the least common multiple rule; the extended group includes V extended strips; split the V extended strips to determine P basic groups and R adjustment groups; each of the basic groups includes Vp basic strips, and each of the adjustment groups includes Vr adjustment strips; P and R are positive integers greater than 1; select K basic strips from the basic group and select D adjustment strips from the adjustment group, and determine a target group based on the K basic strips and the D adjustment strips; the target group includes K+D strips.
[0143] Optionally, the least common multiple rule is determined using the following formula:
[0144] V=LCM(K,K+D+1)(K+D)(K+1) / K
[0145] The LCM() is used to represent the function of finding the least common multiple, k is used to represent the number of nodes storing data blocks on each stripe before expansion, and d is used to represent the number of newly added nodes.
[0146] The optional obtaining unit 1604 is specifically used to: number the K+D stripes in any of the target groups, and number the K+M+D nodes after the storage system is expanded; calculate the difference check blocks of the data blocks on the adjusted stripes in the first K+1 nodes, and update the first check block of the basic stripe on the same node based on the difference check blocks; transfer the data blocks on the adjusted stripe to the basic stripe on the same node in a round-robin mode to obtain the initial extended group after expansion; perform preset operations on the initial extended group to obtain the corresponding target extended group.
[0147] Optionally, the device also includes an adjustment unit, which is used to: determine the logical relationship of the strips corresponding to the target extension group, and the first spatial distribution information corresponding to each extended data block and the extended check block; according to the spatial distribution information, adjust the order of the logical relationship so that the first spatial distribution information is the same as the logical layout of the spatial distribution information.
[0148] An embodiment of the present invention provides a computer device, including a program or an instruction. When the program or the instruction is executed, it is used to execute a method for extending an erasure code storage system and any optional method provided in an embodiment of the present invention.
[0149] An embodiment of the present invention provides a storage medium, including a program or an instruction. When the program or the instruction is executed, it is used to execute a method for extending an erasure code storage system and any optional method provided in an embodiment of the present invention.
[0150] Finally, it should be noted that: It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0151] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0152] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0153] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for extending an erasure code storage system, characterized in that: The method comprises: Determine the data in the storage system, encode the data, and store the data in a dispersed manner in each node to obtain spatial position distribution information of each node; Based on the expansion demand information, determine the number of nodes newly added on each stripe, and based on the number of newly added nodes and the spatial position distribution information, determine the extended node information on each stripe; the stripe includes a data block and a check block having a coding relationship; Based on the expansion node information and the least common multiple rule, an expansion group is determined, and the expansion group is split to obtain a target group including multiple selected strips; the expansion group is composed of multiple strips that meet the conditions that the expansion requirements can be completed and the spatial position distribution law remains unchanged; An expansion algorithm is executed on the target group to update the check blocks stored locally in each node to obtain a corresponding target expansion group, wherein the target expansion group includes an expansion data block and an expansion check block.
2. The method according to claim 1, characterized in that Encoding the data and dispersively storing the data in each node to obtain spatial position distribution information of each node includes: Divide the data into K data blocks of the same size; K is a positive integer greater than 1; Performing an intra-domain matrix operation on the K data blocks and a preset coding matrix to obtain M check blocks; M is a positive integer greater than 1 and less than K; the K data blocks and the M check blocks constitute a plurality of stripes; The data blocks and check blocks on the same strip are dispersed on different K+M nodes, the distribution information of the K data blocks and the M check blocks on each node is determined, and the spatial position distribution information is obtained based on the distribution information.
3. The method according to claim 1 or 2, characterized in that Determining the extended node information on each stripe based on the number of newly added nodes and the spatial position distribution information includes: Based on the spatial position distribution information, determine the number of first nodes storing data blocks on each stripe and the number of second nodes storing check blocks on each stripe; Add the first number of nodes and the newly added number of nodes to obtain the third number of nodes, and use the third number of nodes as the number of expanded storage data blocks on each stripe; and use the second number of nodes as the number of expanded storage check blocks on each stripe to determine the expanded node information on each stripe.
4. The method according to claim 2, characterized in that Based on the extended node information and the least common multiple rule, an extended group is determined, and the extended group is split to obtain a target group including stripes having a corresponding relationship, including: Based on the extended node information and the least common multiple rule, an extended group is determined; the extended group includes V extended stripes; The V extended strips are split to determine P basic groups and R adjustment groups; each of the basic groups includes Vp basic strips, and each of the adjustment groups includes Vr adjustment strips; P and R are positive integers greater than 1; K basic stripes are selected from the basic group, and D adjustment stripes are selected from the adjustment group, and a target group is determined based on the K basic stripes and the D adjustment stripes; the target group includes K+D stripes.
5. The method according to claim 4, characterized in that The least common multiple rule is determined using the following formula: Wherein, the LCM() is used to represent the function for obtaining the least common multiple, K is used to represent the number of nodes storing data blocks on each stripe before expansion; and D is used to represent the number of newly added nodes.
6. The method according to claim 4, characterized in that Execute an expansion algorithm on the target group to obtain a corresponding target expansion group, wherein the target expansion group includes an expansion data block and an expansion check block, including: Numbering the K+D stripes in any of the target groups, and numbering the K+M+D nodes after the storage system is expanded; Calculate the difference check blocks of the data blocks on the adjustment stripe in the first K+1 nodes, and update the first check block of the basic stripe on the same node based on the difference check blocks; Transmitting the data blocks on the adjusted stripe to the basic stripe on the same node in a round-robin mode to obtain an expanded initial extended group; A preset operation is performed on the initial extended group to obtain a corresponding target extended group.
7. The method according to claim 1, characterized in that After obtaining the target extended group, the method further includes: Determine the logical relationship of the stripes corresponding to the target extended group, and first spatial distribution information corresponding to each extended data block and extended check block; According to the spatial distribution information, the logical relationship order is adjusted so that the logical layout of the first spatial distribution information is the same as that of the spatial distribution information.
8. A device for extending an erasure code storage system, characterized in that: The device comprises: A first processing unit is used to determine the data in the storage system, encode the data, and store the data in various nodes in a dispersed manner to obtain spatial position distribution information of the various nodes; A second processing unit is used to determine the number of nodes newly added on each stripe based on the expansion demand information, and determine the extended node information on each stripe based on the number of newly added nodes and the spatial position distribution information; the stripe includes a data block and a check block having a coding relationship; A third processing unit is used to determine an extension group based on the extension node information and the least common multiple rule, and split the extension group to obtain a target group including multiple selected strips; the extension group is composed of multiple strips that meet the conditions that the extension requirement can be completed and the spatial position distribution law remains unchanged; The obtaining unit is used to execute an expansion algorithm on the target group to update the check blocks stored locally in each node and obtain a corresponding target expansion group, wherein the target expansion group includes an expansion data block and an expansion check block.
9. A computer device, characterized in that: The method comprises a program or an instruction, and when the program or the instruction is executed, the method according to any one of claims 1 to 7 is performed.
10. A storage medium, characterized in that: The method comprises a program or an instruction, and when the program or the instruction is executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Patent Citations
Multi-cloud storage system expansion method based on RAID4 (Redundant Array of Independent Disks)
CN106027653A
Fault-tolerant coding method, device and system for improving expandability of data deduplication system
CN111831223A