Graph-based Cross-Rack Update Method, System, and Readable Storage Medium

By building relationship diagrams and parallel tree to optimize the placement and transmission path of data blocks, the problems of I/O operations and cross-cavity transmission in the existing erasure coded data update method are solved, and more efficient cloud storage system performance is achieved.

CN116009782BActive Publication Date: 2025-07-22SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211710238.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2025-07-22
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

The existing erasure coded data update method has unreasonable data placement strategies in the cloud storage system, resulting in additional I/O operations and cross-cavity data transmission, and the transmission strategy is not efficient enough, resulting in unbalanced cabinet load and network transmission delay.

Method used

By constructing a relationship diagram to evaluate the access popularity and correlation of data blocks, organize the data blocks into strips and place the most relevant set of data blocks in the same cabinet, and build a parallel tree to optimize network transmission paths and reduce cross-cluster transmission volume and I/O operations.

Benefits of technology

It effectively reduces the number of I/O operations and cross-club data transmission volume, optimizes network transmission performance, realizes load balancing and reduces network transmission delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116009782B_ABST
    Figure CN116009782B_ABST
Patent Text Reader

Abstract

The present invention provides a graph-based cross-rack update method, system and readable storage medium. The method includes the following steps: S1. Construct a relationship graph to assist data placement. In the relationship graph, points represent data blocks, edges represent the associations between data blocks, point weights represent the access heat of data blocks, and edge weights represent the degree of association between two data blocks; S2. Find a set of data blocks in the relationship graph such that the sum of the edge weights within this set of data blocks reaches the maximum and organize them into stripes; S3. Place the set of data blocks with the strongest correlation in the same rack and ensure that the difference between the point weights of each rack is minimized; S4. Construct a parallel tree. In the parallel tree, nodes represent racks and edges represent the network links between racks. By using the relationship graph to assist data placement, the present invention can reduce the number of I / O operations and the cross-rack transmission volume. Moreover, by constructing a parallel tree to select the nearest network transmission path, the network transmission delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud storage, and particularly to a graph-based cross-rack update method, system and readable storage medium. Background Art

[0002] Erasure code technology has been widely used in cloud storage systems. Compared with the replication technology, it can provide high reliability at a lower storage cost.

[0003] Data update methods based on erasure codes can generally be divided into two categories: in-place update and off-site update. Among them, in-place update directly rewrites at the original position of the data block, and can generally be divided into the following three categories:

[0004] (1) "Difference-based" update: PL (Parity Logging) accumulates several parity updates into one overall parity update, reducing the update load of parity nodes. PLR (Parity Logging with Reserved Space) is an optimization based on PL. PLR allocates additional continuous storage space for each parity block to store parity differences, so as to reduce disk addressing time.

[0005] (2) "Rack-aware" update: There is a bandwidth difference between inside and outside the rack. CAU (Cross-rack-Aware Updates) reduces the cross-rack transmission volume by swapping the positions of data blocks and placing the updated data blocks in the same rack. The key idea of RackCU (Rack-Coordinated Updates) is to select a rack containing the most updated data, and this rack collects all the remaining updated data and completes the calculation of parity. RackCU can minimize the cross-rack transmission volume.

[0006] (3) "Network-distance-aware" update. T-Update (Tree-structured Update scheme) defines the transmission distance between two nodes, and constructs a minimum spanning tree with the updated node as the root node, allowing data blocks to be transmitted along the edges of the tree to accelerate the transmission speed.

[0007] However, the existing methods still have some deficiencies. On the one hand, their data placement strategies are unreasonable. Most algorithms ignore the relationship between data blocks, which results in more stripes being involved in the data update process, causing additional I / O operations and cross-rack data transmission, which consume precious disk resources and cross-rack bandwidth. On the other hand, the transmission strategies of these methods are not efficient enough and do not consider the parallelism of network transmission. Therefore, only some racks may be frequently updated and bear most of the transceiver tasks. This leads to load imbalance between racks, thus causing tail latency in updates.

[0008] Therefore, in view of the above problems, the present application proposes a graph-based cross-rack update method for a cloud storage system, which can accelerate the cross-rack update process from two aspects of data placement and network parallel transmission. Summary of the Invention

[0009] The object of the present invention is to provide a graph-based cross-rack update method, system and readable storage medium, which solves the problems of additional I / O operations and cross-rack data transmission brought by the existing data update method based on erasure coding.

[0010] To achieve the above object, the present invention provides a graph-based cross-rack update method, including the following steps:

[0011] S1. Construct a relationship graph to assist data placement. In the relationship graph, points represent data blocks, edges represent the associations between the data blocks, point weights represent the access heat of the data blocks, and edge weights represent the degree of association between two data blocks;

[0012] S2. Find a group of data blocks in the relationship graph such that the sum of the edge weights within this group of data blocks reaches the maximum and organize them into a stripe;

[0013] S3. Place the group of data blocks with the strongest correlation in the stripe in the same rack and ensure that the difference between the point weights of each rack is minimized;

[0014] S4. Construct a parallel tree. In the parallel tree, nodes represent racks, edges represent network links between racks, and the data blocks can be transmitted between the racks along the network links.

[0015] Optionally, the S1 specifically includes:

[0016] S11. Set up a write buffer to save the data blocks written recently;

[0017] S12. Count the access times of each data block in the write buffer and the times that any two data blocks are accessed simultaneously;

[0018] S13. Construct the relationship graph with the data blocks as points and the connections between two data blocks as edges. The point weight is the access times of the data block, which is used to represent the access heat of the data block, and the edge weight is the times that two data blocks are accessed simultaneously, which is used to represent the degree of association between two data blocks.

[0019] Optionally, the S2 specifically includes:

[0020] S21. Initialize an empty stripe;

[0021] S22. Find the two data blocks corresponding to the maximum edge weight and add them to the stripe, and at the same time remove the two data blocks from the relationship graph;

[0022] S23. For the data blocks already added to the stripe, find the data block corresponding to the maximum edge weight in the relationship graph and fill it into the stripe;

[0023] S24. Repeat S21 - S23 until all the data blocks in the relationship graph are removed.

[0024] Optionally, the stripe is one or more.

[0025] Optionally, the specific steps of S3 are as follows:

[0026] S31. For each data block in the stripe, calculate the weight generated when the data block is placed in each cabinet. The weight consists of two parts: the first part represents the ratio of the total weight of the data blocks in the current cabinet to the average weight of the data blocks in the stripe, reflecting the average degree of access popularity of each cabinet; the second part represents the sum of the edge weights between the data blocks in the cabinet, reflecting the relationship between the data blocks in the cabinet;

[0027] S32. Place the data block in the cabinet that can make the weight maximum.

[0028] Optionally, the number of data blocks in each cabinet cannot exceed the number of check blocks in the corresponding stripe.

[0029] Optionally, the specific steps of S4 are as follows:

[0030] S41. Initialize an empty parallel tree;

[0031] S42. Select a cabinet that is closest to the cabinet storing the check block and add it to the parallel tree;

[0032] S43. For each cabinet in the parallel tree, select a cabinet with the closest network distance to it from the remaining cabinets and add it to the parallel tree. Repeat this step until all cabinets are added to the parallel tree.

[0033] Optionally, in the parallel tree, the degree of the nodes is increasing, and the transmission direction is from the cabinet with a smaller degree to the cabinet with a larger degree.

[0034] Based on the same inventive concept, the present application also proposes a cross - cabinet update system based on a graph, including:

[0035] A relationship graph construction module, which is configured to construct a relationship graph to assist data placement. In the relationship graph, points represent data blocks, edges represent the associations between the data blocks, the point weights represent the access hotness of the data blocks, and the edge weights represent the degree of association between two data blocks;

[0036] A stripe organization module, which is configured to find a group of data blocks in the relationship graph such that the sum of the edge weights within this group of data blocks reaches the maximum and organize them into stripes;

[0037] A stripe division module, which is configured to place the group of data blocks with the strongest correlation in the same cabinet and ensure that the difference between the point weights of each cabinet is minimized;

[0038] A parallel tree construction module, which is configured to construct a parallel tree. In the parallel tree, nodes represent cabinets, edges represent the network links between the cabinets, and the data blocks can be transmitted between the cabinets along the network links.

[0039] Based on the same inventive concept, the present application also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above-mentioned graph-based cross-cabinet update method.

[0040] In a graph-based cross-cabinet update method, system and readable storage medium provided by the present invention, by evaluating the access hotness of data blocks and the relationships between data blocks, using a relationship graph to assist data placement, the number of I / O operations and the cross-cabinet transmission volume can be reduced. And, by constructing a parallel tree to select the nearest network transmission path and transmitting data blocks in parallel to reduce network transmission latency, the cross-cabinet transmission performance is optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Those of ordinary skill in the art will understand that the provided drawings are used to better understand the present invention and do not constitute any limitation to the scope of the present invention. Among them:

[0042] Figure 1 is a flowchart of a graph-based cross-cabinet update method provided by an embodiment of the present invention;

[0043] Figure 2 is a schematic diagram of a graph-based cross-cabinet update device provided by an embodiment of the present invention.

[0044] In the drawings:

[0045] 100 - relationship graph construction module; 200 - stripe organization module; 300 - stripe division module; 400 - parallel tree construction module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] To make the objectives, advantages, and features of the present invention clearer, the following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments. It should be noted that the accompanying drawings are in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the objectives of the embodiments of the present invention. To make the objectives, features, and advantages of the present invention more obvious and understandable, please refer to the accompanying drawings. It should be known that the structures, scales, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Any modification of the structure, change in the proportional relationship, or adjustment of the size, in the case of being the same or similar to the effects that the present invention can produce and the objectives that can be achieved, should still fall within the scope covered by the technical content disclosed by the present invention.

[0047] As used in the present invention, the singular forms "a", "an", and "the" include plural objects unless the context clearly indicates otherwise. As used in the present invention, the term "or" is generally used in the sense of including "and / or" unless the context clearly indicates otherwise. As used in the present invention, the term "several" is generally used in the sense of including "at least one" unless the context clearly indicates otherwise. As used in the present invention, the term "at least two" is generally used in the sense of including "two or more" unless the context clearly indicates otherwise. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", "third" may explicitly or implicitly include one or at least two of such features.

[0048] In the description of the present invention, unless otherwise clearly specified and limited, the terms "install", "connect", "connection", "fix" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0049] Please refer to Figure 1 , Figure 1 , which is a flowchart of the graph-based cross-rack update method provided for an embodiment of the present invention. This embodiment provides a graph-based cross-rack update method, including the following steps:

[0050] S1. Construct a relationship graph to assist in data placement. In the relationship graph, points represent data blocks, edges represent the associations between the data blocks, the point weights represent the access heat of the data blocks, and the edge weights represent the degree of association between two data blocks;

[0051] S2. Find a set of data blocks in the relationship graph such that the sum of the edge weights within this set of data blocks reaches the maximum and organize them into a stripe;

[0052] S3. Place the set of data blocks with the strongest correlation in the stripe in the same cabinet and ensure that the difference between the point weights of each cabinet is minimized;

[0053] S4. Construct a parallel tree. In the parallel tree, nodes represent cabinets, edges represent the network links between cabinets, and the data blocks can be transmitted between the cabinets along the network links.

[0054] By evaluating the access heat of data blocks and the relationships between data blocks, and using a relationship graph to assist in data placement, the present invention can reduce the number of I / O operations and the cross-cabinet transmission volume. Moreover, by constructing a parallel tree to select the nearest network transmission path and transmitting data blocks in parallel to reduce network transmission latency, the cross-cabinet transmission performance is optimized.

[0055] First, execute step S1 to construct a relationship graph to assist in data placement. In the relationship graph, points represent data blocks, edges represent the associations between the data blocks, the point weights represent the access heat of the data blocks, and the edge weights represent the degree of association between two data blocks.

[0056] The specific content of S1 includes:

[0057] S11. Set up a write buffer to save the data blocks written recently;

[0058] S12. Count the access times of each data block in the write buffer and the times that any two data blocks are accessed simultaneously;

[0059] S13. Construct the relationship graph with the data blocks as points and the connections between two data blocks as edges. The point weight is the access times of the data block, which is used to represent the access heat of the data block, and the edge weight is the times that two data blocks are accessed simultaneously, which is used to represent the degree of association between two data blocks.

[0060] Then execute step S2. Before the data is written into the storage system, it needs to be organized into a stripe. According to the definition of the edge weight in the relationship graph, the larger the edge weight, the greater the probability that two data blocks are updated simultaneously. Therefore, to organize the most closely related data blocks into a stripe, a set of nodes needs to be found in the relationship graph such that the sum of the edge weights within this set of nodes reaches the maximum.

[0061] The specific steps of S2 are as follows:

[0062] S21: Initialize an empty stripe;

[0063] S22: Find the two data blocks corresponding to the maximum edge weight and add them to the stripe, and at the same time remove the two data blocks from the relationship graph;

[0064] S23: For the data blocks already added in the stripe, find the data block corresponding to the maximum edge weight in the relationship graph and fill it into the stripe;

[0065] S24: Repeat S21 - S23 until all the data blocks in the relationship graph are removed.

[0066] Optionally, there is one or more stripes.

[0067] Then, perform step S3, place the data blocks with the strongest correlation in the same cabinet, and ensure that the difference between the point weights of each cabinet is minimized. By placing the group of data blocks with the strongest correlation in the same cabinet, cross - cabinet transmission can be reduced, and at the same time, load balancing can be achieved by balancing the heat of the data blocks in each cabinet. There are three aspects to consider in this step. First, to ensure that the most relevant group of data blocks is placed in the same cabinet, it is necessary to ensure that the sum of the edge weights in the cabinet reaches the maximum. Second, to achieve load balancing between cabinets, it is also necessary to ensure that the difference between the point weights of each cabinet is minimized. Finally, the fault - tolerance ability at the cabinet level also needs to be ensured, so the number of data blocks in each cabinet cannot exceed the number of parity blocks in the corresponding stripe.

[0068] In this embodiment, the specific steps of S3 are as follows:

[0069] S31: For each data block in the stripe, calculate the weight generated when the data block is placed in each cabinet. The weight consists of two parts: the first part represents the ratio of the total weight of the data blocks in the current cabinet to the average weight of the data blocks in the stripe, reflecting the average degree of access heat of each cabinet; the second part represents the sum of the edge weights between the data blocks in the cabinet, reflecting the relationship between the data blocks in the cabinet;

[0070] S32: Place the data block in the cabinet that can make the weight maximum.

[0071] Finally, step S4 is executed to construct a parallel tree. In the parallel tree, nodes represent cabinets and edges represent network links between cabinets. The data blocks can be transmitted between the cabinets along the network links. A parallel tree is a weighted undirected graph used to determine the parallel transmission paths of data blocks. When constructing the parallel tree, first select a cabinet closest to the verification cabinet as the transmission end point and initialize the parallel tree with this cabinet. Then, expand the parallel tree periodically. In each round, select a cabinet with the closest distance for each node in the parallel tree to match it.

[0072] The specific steps of S4 include:

[0073] S41. Initialize an empty parallel tree;

[0074] S42. Select a cabinet with the closest distance to the cabinet storing the verification block and add it to the parallel tree;

[0075] S43. For each cabinet in the parallel tree, select a cabinet with the closest network distance to it from the remaining cabinets and add it to the parallel tree. Repeat this step until all cabinets are added to the parallel tree.

[0076] In the parallel tree, the degree of nodes is increasing, and the transmission direction is from the cabinet with a smaller degree to the cabinet with a larger degree.

[0077] Please refer to Figure 2 , based on the same inventive concept, an embodiment of the present invention further provides a cross-cabinet update system based on a graph, including:

[0078] A relationship graph construction module 100, which is configured to construct a relationship graph to assist data placement. In the relationship graph, points represent data blocks, edges represent the associations between the data blocks, the point weights represent the access heat of the data blocks, and the edge weights represent the degree of association between two data blocks;

[0079] A stripe organization module 200, which is configured to find a group of data blocks in the relationship graph so that the sum of the edge weights inside this group of data blocks reaches the maximum and organize them into stripes;

[0080] A stripe division module 300, which is configured to place a group of data blocks with the strongest correlation in the same cabinet and ensure that the difference between the point weights of each cabinet is the smallest;

[0081] A parallel tree construction module 400, which is configured to construct a parallel tree. In the parallel tree, nodes represent cabinets and edges represent network links between cabinets. The data blocks can be transmitted between the cabinets along the network links.

[0082] Based on the same inventive concept, the present application also provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it can implement the above-described graph-based cross-cabinet update method as described above.

[0083] The readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. For example, it may be, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer programs described herein can be downloaded from the readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer program from the network and forwards the computer program for storage in the readable storage medium in each computing / processing device. The computer program for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer program may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer program to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute computer-readable program instructions to implement various aspects of the present invention.

[0084] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, systems, and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer programs. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when these programs are executed by the processor of the computer or other programmable data processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. These computer programs can also be stored in a readable storage medium, and these computer programs cause a computer, a programmable data processing apparatus, and / or other devices to work in a specific manner. Thus, the readable storage medium storing the computer programs includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.

[0085] The computer programs can also be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operation steps are executed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the computer programs executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.

[0086] In summary, the embodiments of the present invention provide a graph-based cross-rack update method, system, and readable storage medium. By evaluating the access popularity of data blocks and the relationships between data blocks, and using a relationship graph to assist data placement, the number of I / O operations and the cross-rack transmission volume can be reduced. Moreover, by constructing a parallel tree to select the nearest network transmission path and transmitting data blocks in parallel to reduce network transmission latency, the cross-rack transmission performance is optimized.

[0087] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure belong to the protection scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A graph-based cross-cabinet update method, characterized in that, Including the following steps: S1. Construct a relationship graph to assist data placement. In the relationship graph, points represent data blocks, edges represent the associations between the data blocks, the point weight represents the access heat of the data block, and the edge weight represents the degree of association between two data blocks; S21. Initialize an empty stripe; S22. Find the two data blocks corresponding to the maximum edge weight and add them to the stripe, and at the same time remove the two data blocks from the relationship graph; S23. For the data blocks already added in the stripe, find the data block corresponding to the maximum edge weight in the relationship graph and fill it into the stripe; S24. Repeat S21 - S23 until all the data blocks in the relationship graph are removed; S3. Place the group of data blocks with the strongest correlation in the stripe in the same cabinet, and ensure that the difference between the point weights of each cabinet is the smallest; S4. Construct a parallel tree. In the parallel tree, nodes represent cabinets, edges represent the network links between cabinets, and the data blocks are transmitted between the cabinets along the network links; Among them, the specific content of S4 includes: S41. Initialize an empty parallel tree; S42. Select a cabinet that is closest to the cabinet storing the check block and add it to the parallel tree; S43. For each cabinet in the parallel tree, select a cabinet with the closest network distance to it from the remaining cabinets and add it to the parallel tree respectively. Repeat this step until all cabinets are added to the parallel tree.

2. The graph-based cross-cabinet update method according to claim 1, wherein The specific content of S1 includes: S11. Set a write buffer to save the data blocks written recently; S12. Count the access times of each data block in the write cache and the times when any two data blocks are accessed simultaneously; S13. Construct the relationship graph with the data blocks as points and the connections between two data blocks as edges. The point weight is the access times of the data block, which is used to represent the access heat of the data block, and the edge weight is the times when two data blocks are accessed simultaneously, which is used to represent the degree of association between two data blocks.

3. The graph-based cross-rack update method according to claim 1, wherein The stripe is one or more.

4. The graph-based cross-rack update method according to claim 1, wherein The specific content of S3 includes: S31. For each data block in the stripe, calculate the weight generated when the data block is placed in each cabinet. The weight consists of two parts: the first part represents the ratio of the total weight of the data blocks in the current cabinet to the average weight of the data blocks in the stripe, which reflects the average degree of access heat of each cabinet; the second part represents the sum of the edge weights between the data blocks in the cabinet, which reflects the relationship between the data blocks in the cabinet; S32. Place the data block in the cabinet that can make the weight the largest.

5. The graph-based cross-rack update method according to claim 4, wherein The number of data blocks in each cabinet cannot exceed the number of check blocks in the corresponding stripe.

6. The graph-based cross-rack update method according to claim 1, wherein In the parallel tree, the degree of the node is increasing, and the transmission direction is from the cabinet with a smaller degree to the cabinet with a larger degree.

7. A graph-based cross-cabinet update system, characterized in that, Including: A relationship graph construction module, which is configured to construct a relationship graph to assist data placement. In the relationship graph, points represent data blocks, edges represent the associations between the data blocks, the point weight represents the access heat of the data block, and the edge weight represents the degree of association between two data blocks; A stripe organization module, which is configured to execute the following steps: S21. Initialize an empty stripe; S22. Find the two data blocks corresponding to the maximum edge weight and add them to the stripe, and at the same time remove the two data blocks from the relationship graph; S23. For the data blocks that have been added to the stripe, find the data block corresponding to the maximum edge weight in the relationship graph and fill it into the stripe; S24. Repeat S21 - S23 until all the data blocks in the relationship graph are removed; A stripe partitioning module, which is configured to place the group of data blocks with the strongest correlation in the stripe in the same cabinet and ensure that the difference between the point weights of each cabinet is minimized; A parallel tree construction module, which is configured to construct a parallel tree. In the parallel tree, nodes represent cabinets and edges represent network links between cabinets, and the data blocks are transmitted between the cabinets along the network links; Among them, the parallel tree construction module is specifically configured to perform the following steps: S41. Initialize an empty parallel tree; S42. Select a cabinet that is closest to the cabinet storing the parity block and add it to the parallel tree; S43. For each cabinet in the parallel tree, select a cabinet with the closest network distance to it from the remaining cabinets and add it to the parallel tree. Repeat this step until all cabinets are added to the parallel tree.

8. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it can implement the graph-based cross-cabinet update method described in any one of claims 1 - 6.

Citation Information

Patent Citations

  • Mixed cloud storage method based on file access frequency

    CN103118133A

  • Data access method, device and equipment, and computer storage medium

    CN112632621A