Single block repair cross-rack flow optimization method for piggyback coding

By building a target map and adjusting the connectivity components, the single-block repair of cross-rack traffic encoded by piggyback is optimized, solving the problem of large cross-rack traffic overhead and improving repair performance.

CN120011132AActive Publication Date: 2025-05-16HUAZHONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510206377.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-16
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

When performing single-block repair of piggyback encoding, cross-rack traffic overhead is high, limiting the improvement of overall performance of the data center.

Method used

By constructing a target graph, a graph communication algorithm is performed based on the coupling degree between data blocks, connecting components are obtained, and adjusted according to the number of racks to obtain uniform grouping, and finally the group with the smallest traffic overhead across racks is selected as the target grouping.

Benefits of technology

It realizes the reduction of cross-rack traffic overhead during single-block repair, optimizes cross-rack traffic, and improves repair performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011132A_ABST
    Figure CN120011132A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer storage, and particularly discloses a piggyback coded single block repair cross-rack flow optimization method, which comprises the following steps that: based on each data block in a piggyback coding result, a target map is constructed, a vertex in the target map represents the data block, and a weight value of an edge is a coupling degree between two data blocks; aiming at each coupling degree, executing a graph connectivity algorithm to obtain a group of connected components corresponding to each coupling degree; according to the number of the racks, each group of connected components is adjusted, corresponding uniform groups are obtained, the number of the data block sets in the uniform groups is equal to the number of the racks, and the difference of the number of the data blocks between the different data block sets in the uniform groups is smaller than a preset threshold value; and in the plurality of uniform groups, selecting one item with the minimum cross-rack traffic overhead as a target group. Through the method and the device, the single block repair for piggyback coding is realized, the cross-rack flow is optimized, and the repair performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer storage technology, and more specifically, relates to a piggyback-coded single-block repair cross-rack traffic optimization method. Background Art

[0002] With the rapid development of the digital age, data centers, as the core hub for information storage and processing, have seen an exponential increase in demand for computing power and space. In this context, server racks are playing an increasingly important role in data center architecture.

[0003] From the perspective of the data center layout structure, a single server rack is like an efficient and coordinated "small computing cluster" with multiple servers deployed inside. These servers in the same rack are closely connected to each other, and high-speed bandwidth is used to achieve fast and smooth data transmission, which greatly improves the execution efficiency and response speed of computing tasks and effectively promotes the efficient operation of the data center.

[0004] However, when data interaction involves servers in different racks, problems arise. Limited by factors such as network topology and physical connection paths, the bandwidth speed between servers in different racks is relatively low, and the data transmission efficiency is significantly reduced, which not only takes more time but also occupies a lot of resources. This bottleneck has become a key factor limiting the overall performance improvement of data centers.

[0005] It can be seen that when performing single-block repair of piggyback coded data blocks, in order to improve the repair efficiency, the cross-rack traffic should be reduced. For single-block repair of piggyback coding, how to optimize the cross-rack traffic is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0006] In view of the defects of the prior art, the purpose of this application is to optimize the cross-rack traffic for a single-block repair of piggyback coding, aiming to solve the problem of high cross-rack traffic overhead in existing repair solutions.

[0007] To achieve the above objectives, in a first aspect, the present application provides a piggyback-coded single-block repair cross-rack traffic optimization method, the method comprising: Based on each data block in the piggyback encoding result, a target graph is constructed, where the vertices in the target graph represent data blocks. Assuming that any two vertices in the target graph represent a first data block and a second data block, if the single-block repair of the first data block requires the second data block to provide x sub-blocks and the single-block repair of the second data block requires the first data block to provide y sub-blocks, then an edge is configured between the two vertices, and the weight of the edge is the coupling degree between the first data block and the second data block, and the coupling degree is the minimum value between x and y, where x and y are positive numbers; For each coupling degree, the graph connectivity algorithm is executed to obtain a set of connected components corresponding to each coupling degree. If the weight of the edge is less than the current coupling degree, the corresponding edge is an invalid edge, and the invalid edge does not participate in the execution process of the graph connectivity algorithm; According to the number of racks, each group of connected components is adjusted (split, merged, etc.) to obtain a corresponding uniform grouping, where the uniform grouping includes multiple data block sets, the number of data block sets in the uniform grouping is equal to the number of racks, and the difference in the number of data blocks between different data block sets in the uniform grouping is less than a preset threshold; Among multiple uniform groupings, the one with the smallest cross-rack traffic overhead is selected as the target grouping (the target grouping is used to indicate a scheme for uniformly distributing multiple data blocks in the piggyback encoding result to each rack).

[0008] Here, the target grouping is exemplified. When the number of racks is 2, if the target grouping is [1 25], [3 4 6], it means that data blocks 1, 2, and 5 are placed in rack 1, and data blocks 3, 4, and 6 are placed in rack 2.

[0009] Specifically, by constructing a target graph, the weight of the edge between vertices can be used to represent the coupling degree between two data blocks; then, for each coupling degree, a graph connectivity algorithm can be executed to obtain a set of connected components corresponding to each coupling degree, a set of connected components includes one or more connected components, and a connected component is specifically a data block set composed of one or more data blocks. For each execution of the graph connectivity algorithm, if the weight of the edge is less than the current coupling degree, the corresponding edge is an invalid edge, and the invalid edge does not participate in the execution process of the graph connectivity algorithm; then, in order to evenly distribute the data blocks to a specified number of racks, each group of connected components can be adjusted according to the number of racks to obtain the corresponding uniform grouping; then, the single-block repair cross-rack traffic overhead corresponding to different uniform groups can be compared, and the one with the smallest single-block repair cross-rack traffic overhead is selected as the target grouping.

[0010] Since target grouping is the item with the smallest cross-rack traffic overhead for single-block repair, it can ensure that after multiple data blocks in the piggyback encoding result are evenly distributed to each rack according to the target grouping, the cross-rack traffic can be minimized when performing single-block repair.

[0011] Therefore, the method provided in the present application can reduce the cross-rack traffic overhead of a single-block repair during single-block repair, implement single-block repair for piggyback encoding, optimize cross-rack traffic, and improve repair performance.

[0012] In a possible implementation, the target graph is constructed based on each data block in the piggyback encoding result, including: For each data block in the piggyback encoding result, the single-block repair data corresponding to each data block is counted, where the single-block repair data indicates the number of sub-blocks required to be provided by data blocks other than the corresponding data block in order to repair the corresponding data block; Based on the single-block repair data corresponding to each data block, graph modeling is performed to obtain the target graph.

[0013] Here, a single-block repair data is exemplified, assuming that the number of sub-blocks required from data blocks 1 to 6 to repair data block 1 is [0, 2, 1, 1, 1]. [0, 2, 1, 1, 1, 1] specifically means that data block 1 needs to provide 0 sub-blocks, data block 2 needs to provide 2 sub-blocks, and so on, data block 6 needs to provide 1 sub-block. [0, 2, 1, 1, 1, 1] is the single-block repair data corresponding to data block 1.

[0014] In a possible implementation, the above-mentioned counting of single-block repair data corresponding to each data block includes: Based on the functional relationship between the original data block and the check data block, the single block repair data corresponding to the data block in the piggyback encoding result is counted.

[0015] For example, if repairing a data block requires z sub-blocks from any n data blocks among m data blocks, then the number of sub-blocks of each of the m data blocks required to repair the data block is recorded. Sub-blocks.

[0016] In a possible implementation, adjusting each group of connected components according to the number of racks to obtain corresponding uniform groupings includes: When the number of connected components in a group of connected components is greater than the number of racks, the connected components in the corresponding group are merged, and the difference in the number of data blocks between different connected components in the corresponding group is eliminated to obtain the corresponding uniform grouping; When the number of connected components in a group of connected components is less than the number of racks, the connected components in the corresponding group are divided, and the difference in the number of data blocks between different connected components in the corresponding group is eliminated to obtain the corresponding uniform grouping; When the number of connected components in a group of connected components is equal to the number of racks, the difference in the number of data blocks between different connected components in the corresponding group is eliminated to obtain the corresponding uniform grouping.

[0017] For example, assuming that the number of racks is 2, if a group of connected components is [3 4], [5 6], [1 2], then the number of connected components in the group is 3, which is greater than the number of racks. Then, the connected components in the group are merged, and the difference in the number of data blocks between different connected components in the group is eliminated, and the corresponding uniform grouping is obtained as [3 4 1], [5 6 2].

[0018] Exemplarily, assuming that the number of racks is 2, if a group of connected components is [1 2 3 4 5 6], the number of connected components in the group is 1, which is less than the number of racks. The connected components in the group are then divided, and the difference in the number of data blocks between different connected components in the corresponding group is eliminated, and the corresponding uniform grouping is obtained as [1 2 3][4 5 6].

[0019] For example, assuming that the number of racks is 2, if a group of connected components is [1 2], [3 4 5 6], then the number of connected components in the group of connected components is 2, which is equal to the number of racks, thereby eliminating the difference in the number of data blocks between different connected components in the corresponding group, and obtaining the corresponding uniform grouping of [1 2 3] [4 5 6].

[0020] In a possible implementation, among the multiple uniform groups, selecting a single-block repair cross-rack traffic overhead with the smallest amount as the target group includes: Calculate the single-block repair cross-rack traffic overhead corresponding to each uniform grouping; Based on the single-block repair cross-rack traffic overhead corresponding to each uniform grouping, a uniform grouping with the minimum overhead is selected as the target grouping.

[0021] In a possible implementation, the above calculation of the single-block repair cross-rack traffic overhead corresponding to each uniform grouping includes: When uniform grouping is used to evenly distribute multiple data blocks in the piggyback encoding result to each rack, the single block repair cross-rack traffic corresponding to each data block is counted; Based on the single-block repair cross-rack traffic corresponding to all data blocks, a sum calculation is performed to obtain the single-block repair cross-rack traffic overhead corresponding to the uniform grouping.

[0022] In a second aspect, the present application provides a piggyback-coded single-block repair cross-rack traffic optimization device, comprising: A graph construction module is used to construct a target graph based on each data block in the piggyback encoding result, where the vertices in the target graph represent data blocks. Assuming that any two vertices in the target graph represent a first data block and a second data block, if single-block repair of the first data block requires the second data block to provide x sub-blocks and single-block repair of the second data block requires the first data block to provide y sub-blocks, then an edge is configured between the two vertices, and the weight of the edge is the coupling degree between the first data block and the second data block, where the coupling degree is the minimum value between x and y, and x and y are positive numbers; The coupling degree analysis module is used to execute the graph connectivity algorithm for each coupling degree and obtain a set of connectivity components corresponding to each coupling degree. If the weight of an edge is less than the current coupling degree, the corresponding edge is an invalid edge and does not participate in the execution process of the graph connectivity algorithm. A uniform grouping acquisition module is used to adjust each group of connected components according to the number of racks to obtain a corresponding uniform grouping, wherein the uniform grouping includes multiple data block sets, the number of data block sets in the uniform grouping is equal to the number of racks, and the difference in the number of data blocks between different data block sets in the uniform grouping is less than a preset threshold; The target group selection module is used to select the one with the smallest cross-rack traffic overhead from multiple uniform groups as the target group.

[0023] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.

[0024] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.

[0025] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.

[0026] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0027] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art: Since target grouping is the item with the lowest cross-rack traffic overhead for single-block repair, it can ensure that after multiple data blocks in the piggyback encoding results are evenly distributed to each rack according to the target grouping, cross-rack traffic can be reduced to a minimum when performing single-block repair, thereby achieving single-block repair for piggyback encoding, optimizing cross-rack traffic, and improving repair performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flow chart of a method for optimizing cross-rack traffic with single-block repair using piggyback coding provided in an embodiment of the present application; Figure 2 It is a schematic diagram of the piggyback framework provided in an embodiment of the present application; Figure 3 It is a schematic diagram of graph modeling based on each data block in the piggyback encoding result provided by an embodiment of the present application; Figure 4 It is a structural schematic diagram of a single-block repair cross-rack traffic optimization device for piggyback coding provided in an embodiment of the present application; Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to facilitate a clearer understanding of the embodiments of the present application, some relevant background knowledge is first introduced as follows.

[0030] Using erasure coding When storing data, a file is divided into The original data blocks are encoded into Total data blocks, of which, , is the number of check blocks, A collection of total data blocks is called a stripe. The data block is invalid. When any of the data blocks fails, it is necessary to read data blocks, which can be original data blocks or check blocks. Then, through the erasure code decoding algorithm, data blocks to repair the failed data blocks.

[0031] The piggyback framework proposes the concept of sub-strips based on erasure codes. Specifically, the original stripe is divided into t sub-strips (each data block is divided into t sub-blocks), and the check sub-block (the sub-block in the check data block is called the check sub-block) is obtained from the data sub-block (the sub-block in the original data block is called the data sub-block) through a functional relationship. Figure 2 Shown is the piggyback frame An example of encoding.

[0032] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0033] The terms "first" and "second" in the specification and claims of this application are used to distinguish different objects rather than to describe a specific order of the objects. For example, a first data block and a second data block are used to distinguish different data blocks rather than to describe a specific order of the data blocks.

[0034] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0035] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.

[0036] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0037] Figure 1 is a flow chart of a method for optimizing cross-rack traffic flow using a single-block repair using piggyback coding provided in an embodiment of the present application, such as Figure 1 As shown, the method specifically includes the following steps S1 to S5.

[0038] Step S1, obtaining the relationship between other data blocks required for single-block repair of each data block.

[0039] The relationship between other data blocks required for repairing a single data block is specifically as follows: an assumption is made for each data block, assuming that the data block is invalid and needs to be repaired, and the number of sub-blocks of other data blocks required to be used is determined based on a functional relationship (coding function) between the original data block and the check data block.

[0040] Step S2, graph modeling the relationship and obtaining the coupling degree between data blocks.

[0041] Step S2 is specifically as follows: construct a graph with data blocks as vertices of the graph; if two data blocks need each other during repair, add an edge between the vertices corresponding to the two data blocks, and use the minimum number of sub-blocks required by the two data blocks (referred to as coupling degree in this application) as the weight of the edge.

[0042] The coupling degree is specifically: when data block A is repaired, a sub-block in data block B is required, and when data block B is repaired, a sub-block in data block A is required, then there is a coupling relationship between data block A and data block B; further, if repairing data block A requires x sub-blocks of data block B, and repairing data block B requires y sub-blocks of data block A, then the coupling degree between data block A and data block B is the minimum value of x and y.

[0043] Step S3: for different coupling degrees, a graph connectivity algorithm is used to obtain a set of connected components corresponding to each coupling degree.

[0044] A connected component includes one or more connected components. A connected component refers to the largest subset of all vertices in a graph that are connected by paths. A connected component is specifically a data block set consisting of one or more data blocks.

[0045] Step S3 is specifically as follows: in the process of running the graph connectivity algorithm for a specific coupling degree, the following rules are set: for the weight of the edge, when the weight is greater than or equal to the current coupling degree, the edge is judged as a valid edge; conversely, if the weight of the edge is less than the coupling degree, it is regarded as an invalid edge, and only valid edges participate in the execution of the graph connectivity algorithm. Since different coupling degrees will produce different sets of valid edges, each specific coupling degree will correspond to a unique set of connected components. If the number of coupling degrees used is n, then n groups of connected components can be obtained. For example, the coupling degrees used are 1, 1.6 and 2, and the number of coupling degrees used is 3. Accordingly, 3 groups of connected components can be obtained.

[0046] Step S4, adjusting (splitting or merging) each group of connected components according to the number of racks to obtain a corresponding uniform grouping (for each group of connected components).

[0047] The uniform grouping includes multiple data block sets, the number of data block sets in the uniform grouping is equal to the number of racks, and the difference in the number of data blocks between different data block sets in the uniform grouping is less than a preset threshold.

[0048] In step S4, the requirement is specifically met as follows: if the data blocks are required to be evenly distributed in x racks, then x even data block sets are required ("even" means that the difference in the number of data blocks between different data block sets is less than a preset threshold).

[0049] For example, the preset threshold can be configured as 1, the number of racks is 2, and the number of data blocks is 8. To evenly distribute 8 data blocks to 2 racks, two even data block sets are required. The number of data blocks in one data block set is 4, and the number of data blocks in the other data block set is also 4. The difference in the number of data blocks between different data block sets is , the difference is less than the preset threshold.

[0050] For example, the preset threshold can be configured as 1, the number of racks is 2, and the number of data blocks is 8. To distribute 8 data blocks (approximately) evenly to 3 racks, 3 even data block sets are required (data block set 1, data block set 2, data block set 3). The number of data blocks in data block set 1 is 3, the number of data blocks in data block set 2 is 2, and the number of data blocks in data block set 3 is 2. The difference in the number of data blocks between data block set 1 and data block set 2 is , the difference is less than the preset threshold. The difference in the number of data blocks between data block set 1 and data block set 3 is , the difference is less than the preset threshold. The difference in the number of data blocks between data block set 2 and data block set 3 is , the difference is less than the preset threshold.

[0051] However, for any group of connected components obtained in step S3, different connected components in the group (one connected component is a set of data blocks) may not be uniform, and the number of connected components in the group may not be equal to x. It is necessary to obtain uniform grouping while minimizing the changes to each group of connected components. In the process of changing each group of connected components to obtain uniform grouping, the connected components in each group will be split or combined. In this step, since step S3 may obtain multiple groups, step S4 will process these groups in turn to obtain multiple uniform groups.

[0052] Step S5 compares the single-block repair cross-rack traffic overheads of multiple uniform groups, and outputs the uniform group with the smallest overhead.

[0053] The step S5 is specifically as follows: the multiple uniform groups obtained in step S4 are calculated. Specifically, when uniform grouping is used to uniformly distribute multiple data blocks in the piggyback encoding result to each rack, the single-block repair cross-rack flow corresponding to each data block is counted; based on the single-block repair cross-rack flow corresponding to all data blocks, a sum calculation is performed to obtain the single-block repair cross-rack flow overhead corresponding to the uniform grouping. The cross-rack flow overhead of the single-block repair of each uniform grouping is compared, and the smallest one is the corresponding group (target group) output.

[0054] Since target grouping is the item with the smallest cross-rack traffic overhead for single-block repair, it can ensure that after multiple data blocks in the piggyback encoding result are evenly distributed to each rack according to the target grouping, the cross-rack traffic can be minimized when performing single-block repair.

[0055] Therefore, the method provided in the present application can reduce the cross-rack traffic overhead of a single-block repair during single-block repair, implement single-block repair for piggyback encoding, optimize cross-rack traffic, and improve repair performance.

[0056] The following is an example to illustrate the single-block repair cross-rack traffic optimization method for piggyback coding provided in this application.

[0057] by Figure 2 Taking the encoding shown as an example, each row is a data block, which are data blocks 1 to 6 from top to bottom, and each row has two sub-blocks.

[0058] First, obtain the relationship between other data blocks required for single-block repair of each data block. According to the functional relationship between the original data block and the check data block, the number of sub-blocks required from data blocks 1 to 6 to repair data block 1 is [0, 2, 1, 1, 1, 1]. Specifically, data block 1 needs to provide 0 sub-blocks, data block 2 needs to provide 2 sub-blocks, and so on, data block 6 needs to provide 1 sub-block. To repair data block 2, the number of sub-blocks required from data blocks 1 to 6 is [2, 0, 1, 1, 1, 1]; to repair data block 3, the number of sub-blocks required from data blocks 1 to 6 is [1, 1, 0, 2, 1, 1]; to repair data block 4, the number of sub-blocks required from data blocks 1 to 6 is [1, 1, 2, 0, 1, 1]; to repair data block 5, the number of sub-blocks required from data blocks 1 to 6 is [1.6, 1.6, 1.6, 1.6, 0, 1.6]; to repair data block 6, the number of sub-blocks required from data blocks 1 to 6 is [1.6, 1.6, 1.6, 1.6, 0]. When repairing data blocks 5 and 6, any 4 of the other 5 data blocks are needed. Each data block has 2 sub-blocks, so 1.6 sub-blocks of each of the remaining data blocks are needed.

[0059] like Figure 3 The figure shows the graph modeling of this embodiment, where the vertices represent the data block numbers and the edges represent the coupling degrees between the data blocks. It can be seen from the figure that there are three different coupling degrees, namely 1, 1.6 and 2.

[0060] Running the graph connectivity algorithm on a coupling degree of 1 results in a set of connected components (this group includes 1 connected component) of [12 3 4 5 6], indicating that data blocks 1 to 6 are in the same group; running the graph connectivity algorithm on a coupling degree of 1.6 results in a set of connected components (this group includes 3 connected components) of [3 4][5 6] [1 2]; running the graph connectivity algorithm on a coupling degree of 2 results in a set of connected components (this group includes 4 connected components) of [1 2] [3 4][5] [6].

[0061] Assuming that this embodiment needs to evenly distribute 6 data blocks to 2 racks, each rack stores 3 data blocks. The even grouping obtained by processing a set of connected components corresponding to coupling degree 1 is [1 2 3] [4 5 6], indicating that data blocks 1 to 3 are a set of data blocks, and data blocks 4 to 6 are a set of data blocks; the even grouping obtained by processing a set of connected components corresponding to coupling degree 1.6 is [3 4 1] [5 6 2]; the even grouping obtained by processing the corresponding connected components corresponding to coupling degree 2 is [1 2 5] [3 4 6].

[0062] The total cross-rack repair flow of the uniform grouping obtained with coupling degree 1 is 23.6 sub-blocks, the coupling degree 1.6 is 23.6 sub-blocks, and the coupling degree 2 is 21.6 sub-blocks. The cross-rack repair flow overhead of the uniform grouping with coupling degree 2 is the smallest, so the output grouping is [1 2 5] [3 4 6]. Data blocks 1, 2, and 5 can be placed in rack 1, and data blocks 3, 4, and 6 can be placed in rack 2.

[0063] Experimental verification shows that, compared with randomly placing data blocks evenly in two racks, the present application can reduce the cross-site overhead of single block repair by 8.61%. ,When the number of target racks is equal to 3, this application can reduce the cross-site overhead during single block repair by 11.19% compared to randomly and evenly placing data blocks.

[0064] Experiments show that compared with traditional solutions, this application can significantly reduce the cross-rack traffic overhead of single-block repair of piggyback encoding, thereby improving repair efficiency.

[0065] The following is a description of the piggyback-coded single-block repair cross-rack traffic optimization device provided in the present application. The piggyback-coded single-block repair cross-rack traffic optimization device described below and the piggyback-coded single-block repair cross-rack traffic optimization method described above can be referenced to each other.

[0066] Figure 4 is a schematic diagram of the structure of a single-block repair cross-rack traffic optimization device for piggyback coding provided in an embodiment of the present application, such as Figure 4 As shown, the device includes: a graph construction module 10, a coupling degree analysis module 20, a uniform grouping acquisition module 30 and a target grouping selection module 40. Among them: A graph construction module 10 is used to construct a target graph based on each data block in the piggyback encoding result, where the vertices in the target graph represent data blocks. Assuming that any two vertices in the target graph represent a first data block and a second data block, if the single-block repair of the first data block requires the second data block to provide x sub-blocks and the single-block repair of the second data block requires the first data block to provide y sub-blocks, then an edge is configured between the two vertices, and the weight of the edge is the coupling degree between the first data block and the second data block, and the coupling degree is the minimum value between x and y, where x and y are positive numbers; The coupling degree analysis module 20 is used to execute the graph connectivity algorithm for each coupling degree to obtain a set of connectivity components corresponding to each coupling degree. If the weight of an edge is less than the current coupling degree, the corresponding edge is an invalid edge, and the invalid edge does not participate in the execution process of the graph connectivity algorithm. A uniform grouping acquisition module 30 is used to adjust each group of connected components according to the number of racks to obtain a corresponding uniform grouping, wherein the uniform grouping includes a plurality of data block sets, the number of data block sets in the uniform grouping is equal to the number of racks, and the difference in the number of data blocks between different data block sets in the uniform grouping is less than a preset threshold; The target group selection module 40 is used to select one of the multiple uniform groups with the smallest cross-rack traffic overhead as the target group.

[0067] It can be understood that the detailed functional implementation of each of the above-mentioned units / modules can be found in the introduction of the aforementioned method embodiment, and will not be repeated here.

[0068] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method, which will not be repeated here.

[0069] Based on the method in the above embodiment, an embodiment of the present application provides an electronic device, Figure 5is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 5 As shown, the electronic device may include: a processor (Processor) 810, a communication interface (Communications Interface) 820, a memory (Memory) 830 and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the method in the above embodiment.

[0070] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.

[0071] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0072] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0073] It is understandable that the processor in the embodiment of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0074] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0075] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

[0076] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0077] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A piggyback-coded single-block repair cross-rack traffic optimization method, characterized in that: include: Based on each data block in the piggyback encoding result, a target graph is constructed, where the vertices in the target graph represent data blocks. Assuming that any two vertices in the target graph represent a first data block and a second data block, if the single-block repair of the first data block requires the second data block to provide x sub-blocks and the single-block repair of the second data block requires the first data block to provide y sub-blocks, then an edge is configured between the two vertices, and the weight of the edge is the coupling degree between the first data block and the second data block, and the coupling degree is the minimum value between x and y, where x and y are positive numbers; For each coupling degree, the graph connectivity algorithm is executed to obtain a set of connected components corresponding to each coupling degree. If the weight of the edge is less than the current coupling degree, the corresponding edge is an invalid edge, and the invalid edge does not participate in the execution process of the graph connectivity algorithm; According to the number of racks, each group of connected components is adjusted to obtain a corresponding uniform grouping, wherein the uniform grouping includes multiple data block sets, the number of data block sets in the uniform grouping is equal to the number of racks, and the difference in the number of data blocks between different data block sets in the uniform grouping is less than a preset threshold; Among multiple uniform groups, the one with the smallest cross-rack traffic overhead is selected as the target group.

2. The method for optimizing cross-rack traffic by single-block repair of piggyback coding according to claim 1, characterized in that: The step of constructing a target graph based on each data block in the piggyback encoding result includes: For each data block in the piggyback encoding result, the single-block repair data corresponding to each data block is counted, where the single-block repair data indicates the number of sub-blocks required to be provided by data blocks other than the corresponding data block in order to repair the corresponding data block; Based on the single-block repair data corresponding to each data block, graph modeling is performed to obtain the target graph.

3. The method for optimizing cross-rack traffic by single-block repair of piggyback coding according to claim 2, characterized in that: The counting of single-block repair data corresponding to each data block includes: Based on the functional relationship between the original data block and the check data block, the single block repair data corresponding to the data block in the piggyback encoding result is counted.

4. The method for optimizing cross-rack traffic by single-block repair of piggyback coding according to claim 1, characterized in that: The step of adjusting each group of connected components according to the number of racks to obtain corresponding uniform groupings includes: When the number of connected components in a group of connected components is greater than the number of racks, the connected components in the corresponding group are merged, and the difference in the number of data blocks between different connected components in the corresponding group is eliminated to obtain the corresponding uniform grouping; When the number of connected components in a group of connected components is less than the number of racks, the connected components in the corresponding group are divided, and the difference in the number of data blocks between different connected components in the corresponding group is eliminated to obtain the corresponding uniform grouping; When the number of connected components in a group of connected components is equal to the number of racks, the difference in the number of data blocks between different connected components in the corresponding group is eliminated to obtain the corresponding uniform grouping.

5. The method for optimizing cross-rack traffic by single-block repair of piggyback coding according to claim 1, characterized in that: The step of selecting a single-block repair cross-rack traffic overhead with the smallest amount from among the multiple uniform groups as the target group includes: Calculate the single-block repair cross-rack traffic overhead corresponding to each uniform grouping; Based on the single-block repair cross-rack traffic overhead corresponding to each uniform grouping, a uniform grouping with the minimum overhead is selected as the target grouping.

6. The method for optimizing cross-rack traffic by single-block repair of piggyback coding according to claim 5, characterized in that: The calculation of the single-block repair cross-rack traffic overhead corresponding to each uniform grouping includes: When uniform grouping is used to evenly distribute multiple data blocks in the piggyback encoding result to each rack, the single block repair cross-rack traffic corresponding to each data block is counted; Based on the single-block repair cross-rack traffic corresponding to all data blocks, a sum calculation is performed to obtain the single-block repair cross-rack traffic overhead corresponding to the uniform grouping.

7. A piggyback-coded single-block repair cross-rack traffic optimization device, characterized in that: include: A graph construction module is used to construct a target graph based on each data block in the piggyback encoding result, where the vertices in the target graph represent data blocks. Assuming that any two vertices in the target graph represent a first data block and a second data block, if single-block repair of the first data block requires the second data block to provide x sub-blocks and single-block repair of the second data block requires the first data block to provide y sub-blocks, then an edge is configured between the two vertices, and the weight of the edge is the coupling degree between the first data block and the second data block, where the coupling degree is the minimum value between x and y, and x and y are positive numbers; The coupling degree analysis module is used to execute the graph connectivity algorithm for each coupling degree and obtain a set of connectivity components corresponding to each coupling degree. If the weight of an edge is less than the current coupling degree, the corresponding edge is an invalid edge and does not participate in the execution process of the graph connectivity algorithm. A uniform grouping acquisition module is used to adjust each group of connected components according to the number of racks to obtain a corresponding uniform grouping, wherein the uniform grouping includes multiple data block sets, the number of data block sets in the uniform grouping is equal to the number of racks, and the difference in the number of data blocks between different data block sets in the uniform grouping is less than a preset threshold; The target group selection module is used to select the one with the smallest cross-rack traffic overhead from multiple uniform groups as the target group.

8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that When the computer program product runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Binary-based node repair method and system

    CN108628697A

  • Methods and devices for radio communications

    CN110301143A

  • Cross-cluster traffic optimization method for single-point failure repair of cluster storage system

    CN111614720A

  • Data recovery method, related device and equipment

    CN115934413A

  • Piggyback coding construction with local repair characteristic and fault node repair method

    CN119376997A

Cited By

  • Piggyback coding-based method for optimizing cross-rack traffic for single-block repair

    WO2026179086A1