Dynamic Node Selection Method and System for Data Repair in Distributed Storage
Through the adaptive node selection method, the node computing power and network distance are dynamically adjusted, which solves the problem of node load changes in erasure coded data repair, and optimizes the repair efficiency and resource utilization.
Patent Information
- Application Number
- CN202211511048.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-29
AI Technical Summary
When repairing existing erasure coded data, dynamic changes in node load are not considered, resulting in insufficient computing power of nodes, affecting repair efficiency, and excessive resource overhead.
Through the adaptive node selection method, the node computing power is dynamically adjusted, and combined with network distance and load balancing, the optimal node is selected to participate in the repair work.
Optimize data repair delay, reasonably plan node load, improve repair efficiency, and avoid insufficient computing power of high-load nodes.
Smart Images

Figure CN115883589B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cloud storage technology, specifically to the field of distributed storage data repair technology, and relates to an adaptive node selection technical solution, which dynamically adjusts the actual computing power of the node according to the change of the node load, and selects the optimal supply node to participate in data repair work by comprehensively considering the network distance and load balancing. Background Art
[0002] In recent years, with the rapid development of the Internet, data has shown a trend of massive growth. Distributed storage systems are widely used because they can flexibly and elastically store data in blocks to ensure data reliability. As storage demand continues to increase, the number of nodes in the system will increase with demand. Nodes may lose connection, causing stored data loss, and node failure is also inevitable. In order to deal with the problem of data loss caused by node failure or loss of connection, it is often necessary to introduce a certain amount of redundancy to ensure data reliability.
[0003] There are generally two ways to introduce data redundancy: one is the multi-copy method, which copies the original data into multiple copies and stores them on different nodes. When the data on a node is lost, the copies on other nodes are read. This method will increase storage costs and reduce resource utilization. Erasure coding technology is widely used in distributed storage systems due to its low storage overhead. It encodes k data blocks through certain encoding rules to generate m redundant blocks, which are stored on different nodes. When a data block on a node is lost, only the data of any k nodes needs to be collected to reconstruct the original file. However, due to the encoding nature of erasure codes, some additional overhead is required in the data repair process, such as network communication, computing, disk I / O and other resource overheads between nodes. These resources vary greatly in heterogeneous network environments. Designing and optimizing erasure code data repair methods based on the characteristics of heterogeneous environments has important practical value.
[0004] Existing methods for selecting nodes during erasure code repair include: based on network distance, available bandwidth, and fixed computing power of nodes, etc., but none of them consider the impact of dynamic changes in node load on data repair in actual storage systems. When the load of a node suddenly increases within a period of time, the processing resources of the node are tight. At this time, if the node is selected as a supply node, the node does not have enough processing margin to respond to data repair requests, but further increases the node load and consumes the node's computing power, which seriously affects the repair efficiency. Therefore, how to reduce the impact of the additional overhead generated during erasure code data repair on data regeneration is still a difficult problem. Summary of the invention
[0005] In view of the above problems existing in the prior art, the present invention conducts research on the node selection strategy involved in data repair, aiming to provide an adaptive node selection method during data repair, considering the impact of node load changes on the computing power of the node itself, and selecting the optimal supply node to participate in the repair work.
[0006] The present invention includes a method for dynamically adjusting the computing power of nodes according to the load and a node selection method based on network distance and load balancing.
[0007] The present invention first models the computing power of nodes, and then dynamically adjusts the computing power of nodes through the detection of node load, so as to adaptively select the optimal nodes to participate in the data repair work.
[0008] The present invention specifically adopts the following technical solutions:
[0009] A dynamic node selection method for data repair in distributed storage, which comprises the following steps:
[0010] (1) Model the computing power of nodes;
[0011] (2) Detect the change of node load;
[0012] (3) Dynamically adjust the computing power of nodes;
[0013] (4) Calculate the processing delay;
[0014] (5) Calculate the transmission delay;
[0015] (6) Select nodes based on network distance and load balancing.
[0016] Preferably, step (1) models the computing power of nodes as follows:
[0017] The main factors affecting the computing power of nodes include: memory storage capacity and cycle, disk I / O, number of CPU cores, CPU main frequency, number of bytes of data block, CPU operation speed, etc. In this step, k factors affecting the computing power of nodes are selected, and are respectively represented by x1, x2, x3,..., x k Considering that the influence effect of each factor on the computing power of nodes is different, different weights c1, c2, c3,..., c k (c1 + c2 + c3 +... + c k = 1) are assigned to the selected main factors. Larger weights are assigned to the factors with greater influence on the computing power of nodes, and smaller weights are assigned to the factors with smaller influence on the computing power of nodes. The computing power of each node can be expressed as:
[0018]
[0019] Preferably, the detection of node load change in step (2) is as follows:
[0020] Regularly send heartbeat packets to each node through the heartbeat detection mechanism to obtain the global load situation of all nodes at this moment. Let the load detection amount of a certain node in the current cycle be P t2 , and the load detection amount in the previous cycle is P t1 . Consider the impact of load change on the computing power of the node. The relative load change can be expressed as:
[0021]
[0022] Preferably, the dynamic adjustment of node computing power in step (3) is as follows:
[0023] Since the erasure code redundancy technology randomly selects storage nodes when storing data blocks, the storage loads of different nodes are not balanced, resulting in some nodes having a storage load greater than that of other nodes within a certain period. Nodes with a larger storage load have a higher probability of participating in data repair than nodes with a smaller load, and the corresponding repair processing load will also increase. When other nodes need to perform repairs, nodes with a larger load may not be the optimal candidate providers. Therefore, in this step, a storage load change threshold Δ preset and a load change lower limit value Δ lower will be preset, considering the impact of the load change amount on the computing power of the node.
[0024] When the relative load change Δ is greater than or equal to the load change threshold Δ preset , it indicates that the load of the node is relatively large in this cycle, and the change amount exceeds the load of the previous cycle. The increase in load will greatly affect the processing ability of the node itself. At this time, this node is no longer suitable as a candidate provider to participate in data repair work. If this node is still selected as the supply node to participate in data repair at this time, it will not only increase the load of this node but also affect the stability of the node. Update the computing power of the node. γ is the conversion coefficient set when the relative load change is greater than or equal to the load change threshold, as shown in the following formula:
[0025] A‘ = A × γ
[0026] When the relative load change Δ is between the load change threshold Δ preset and the load change lower limit value Δ lower , it indicates that the load of the node has increased by a certain amount in this cycle, but has not reached the preset load change proximity value. The change of the node load within the preset load change will have a certain impact on the computing power of the node itself, and the degree of influence depends on the change amount of the load. Update the computing power of the node. ω is the conversion coefficient set when Δ lower ≤Δ≤Δ preset , as shown in the following formula:
[0027] A' = A × (1 - Δ) × ω
[0028] When the relative load change Δ is less than the lower limit value of load change Δ lower , it indicates that the load of the node decreases within this period, and this node has spare computing power, which is more suitable as a candidate node for the supplier. Update the computing power of the node, For the conversion coefficient set when Δ ≤ Δ lower as follows:
[0029]
[0030] By detecting the load change amount, dynamically adjust the computing power of the node.
[0031] Preferably, calculate the processing delay in step (4); specifically as follows:
[0032] Assume that the amount of computation required for each node to repair data is fixed. According to the coding method of the regenerated code, let the storage capacity of each node be α. Then the processing delay of the node for data can be expressed as the following formula, where δ is the conversion coefficient,
[0033]
[0034] As can be seen from the above formula, the stronger the computing power of the node, the shorter the processing delay of the node for data. Conversely, the longer the processing delay of the node for data. Therefore, when repairing data, selecting a node with stronger computing power can optimize the data repair efficiency and reduce the repair delay. Selecting a node with high computing power as the supply node can improve the repair efficiency.
[0035] Preferably, calculate the transmission delay in step (5); specifically as follows:
[0036] The transmission delay refers to the time required for the electromagnetic wave carrying the transmission signal to propagate on a channel of a certain length. The transmission delay can be expressed by the following formula, T ij represents the transmission delay between nodes ij, k is the length of the channel, v is the transmission rate, and ξ is the conversion coefficient,
[0037]
[0038] As can be seen from the above formula, the transmission delay of the data is related to the length of the channel between nodes. The longer the channel, the longer the transmission delay of the data. Therefore, when repairing data, selecting a node close to the new node as the supply node to transmit data can reduce the repair delay and improve the repair efficiency.
[0039] Preferably, select nodes based on network distance and load balancing in step (6); specifically as follows:
[0040] When node i fails, first obtain the load P of the candidate supplier nodes in the current cycle t2 and the load P of the previous cycle t1 , complete the update of the computing power of the candidate supplier nodes through steps (2) and (3), and obtain the processing delay and transmission delay of the nodes for data according to steps (4) and (5).
[0041] When selecting nodes, consider the impact of bandwidth heterogeneity and node computing power on data repair delay, and select the node with the shortest sum of data processing delay and data transmission delay, that is, the node with the shortest repair time. The nodes selected by the above method participate in data regeneration work to recover the lost data.
[0042] The present invention also discloses a system based on the above dynamic node selection method for data repair in distributed storage, which includes the following modules:
[0043] Node computing power modeling module: used to model the computing power of nodes;
[0044] Node load change detection module: used to detect changes in node load;
[0045] Node computing power adjustment module: used to dynamically adjust the computing power of nodes;
[0046] Processing delay calculation module: used to calculate the processing delay;
[0047] Transmission delay calculation module: used to calculate the transmission delay;
[0048] Node selection module: selects nodes based on network distance and load balancing.
[0049] The node selection scheme based on network distance and load balancing proposed by the present invention considers the load changes of storage nodes and dynamically adjusts the computing power of nodes on the basis of heterogeneous computing capabilities of storage nodes. It can select different nodes to participate in data repair work according to the actual load changes, realizes adaptive node selection, can effectively solve the impact of sudden node load on data repair in actual storage systems, optimize data repair delay, and reasonably plan node load. Brief Description of the Drawings
[0050] Figure 1 is a schematic diagram of the network structure topology;
[0051] Figure 2 is a diagram of the node load change situation;
[0052] Figure 3 is a schematic diagram of the star repair model based on the present invention;
[0053] Figure 4 is a schematic diagram of the tree - type repair model based on the present invention;
[0054] Figure 5 is a comparison graph of repair time delays under different data block sizes based on the present invention;
[0055] Figure 6 is a flowchart of a dynamic node selection method for data repair in distributed storage;
[0056] Figure 7 is a block diagram of a dynamic node selection system for data repair in distributed storage. Detailed implementation manners
[0057] The present invention will be further described in detail in conjunction with the following specific embodiments and drawings.
[0058] As Figure 6 shown, the dynamic node selection method for data repair in distributed storage includes the following steps:
[0059] Step (1): Model the computing power of nodes; specifically as follows:
[0060] The main factors affecting the computing power of nodes include: memory storage capacity and cycle, disk I / O, number of CPU cores, CPU main frequency, data block bytes, CPU operation speed, etc. In this step, k factors affecting the computing power of nodes are selected, and are represented by x1, x2, x3,..., x k respectively. Considering that the influence effects of each factor on the computing power of nodes are different, different weights c1, c2, c3,..., c k (c1 + c2 + c3 +... + c k = 1) are assigned to the selected main factors. Larger weights are given to factors with greater influence on the computing power of nodes, and smaller weights are given to factors with smaller influence on the computing power of nodes. The computing power of each node can be expressed as:
[0061]
[0062] Step (2): Detect the change of node load; specifically as follows:
[0063] Heartbeat packets are sent to each node regularly through the heartbeat detection mechanism to obtain the global load situation of all nodes at this moment. Let the load detection value of a certain node in the current cycle be P t2 , and the load detection value in the previous cycle be P t1 . Consider the influence of load change on the computing power of nodes. The relative load change can be expressed as:
[0064]
[0065] Dynamic adjustment of node computing power in step (3); specifically as follows:
[0066] Since the erasure code redundancy technology randomly selects storage nodes when storing data blocks, the storage loads of different nodes are not balanced, resulting in some nodes having a storage load greater than that of other nodes within a certain period. Nodes with a larger storage load have a higher probability of participating in data repair than nodes with a smaller load, and the corresponding repair processing load will also increase. When other nodes need to perform repairs, nodes with a larger load may not be the optimal candidate providers. Therefore, in this step, a storage load change threshold Δ preset and a load change lower limit value Δ lower will be preset, considering the impact of the load change amount on the node computing power.
[0067] When the relative load change Δ is greater than or equal to the load change threshold Δ preset , it indicates that the load of the node is relatively large within this period, and the change amount exceeds the load of the previous period. The increase in load will greatly affect the processing ability of the node itself. At this time, this node is no longer suitable as a candidate provider node to participate in data repair work. If this node is still selected as the supply node to participate in data repair at this time, it will not only increase the load of this node but also affect the stability of the node. Update the computing power of the node. γ is the conversion coefficient set when the relative load change is greater than or equal to the load change threshold, as shown in the following formula:
[0068] A‘ = A × γ
[0069] When the relative load change Δ is between the load change threshold Δ preset and the load change lower limit value Δ lower , it indicates that the load of the node has increased by a certain amount within this period, but has not reached the preset load change proximity value. The change of the node load within the preset load change will have a certain impact on the computing power of the node itself, and the degree of influence depends on the change amount of the load. Update the computing power of the node. ω is the conversion coefficient set when Δ lower ≤Δ≤Δ preset , as shown in the following formula:
[0070] A‘ = A × (1 - Δ) × ω
[0071] When the relative load change Δ is less than the load change lower limit value Δ lower , it indicates that the load of the node becomes smaller within this period, and this node has spare computing power and is more suitable as a candidate provider node. Update the computing power of the node. is the conversion coefficient set when Δ≤Δ lower , as shown in the following formula:
[0072]
[0073] By detecting the change amount of the load, the computing power of the node is dynamically adjusted.
[0074] Calculation of the processing delay in step (4); specifically as follows:
[0075] Assume that the amount of computation required for each node to repair data is fixed. According to the coding method of the regenerating code, let the storage capacity of each node be α. Then the processing delay of the node for data can be expressed by the following formula, where δ is the conversion coefficient.
[0076]
[0077] It can be seen from the above formula that the stronger the computing power of the node, the shorter the processing delay of the node for data. Conversely, the longer the processing delay of the node for data. Therefore, when repairing data, selecting nodes with stronger computing power can optimize the data repair efficiency and reduce the repair delay. Selecting nodes with high computing power as supply nodes can improve the repair efficiency.
[0078] Calculation of the transmission delay in step (5); specifically as follows:
[0079] The transmission delay refers to the time required for the electromagnetic wave carrying the transmission signal to propagate on a channel of a certain length. The transmission delay can be expressed by the following formula, T ij represents the transmission delay of data between nodes ij, k is the length of the channel, v is the transmission rate, and ξ is the conversion coefficient.
[0080]
[0081] It can be seen from the above formula that the transmission delay of data is related to the length of the channel between nodes. The longer the channel, the longer the transmission delay of data. Therefore, when repairing data, selecting nodes close to the new node as supply nodes to transmit data can reduce the repair delay and improve the repair efficiency.
[0082] Step (6) Node selection based on network distance and load balancing; specifically as follows:
[0083] When node i fails, first obtain the load P of the candidate supply nodes in the current cycle t2 and the load P in the previous cycle t1 , complete the update of the computing power of the candidate supply nodes through steps (2) and (3), and obtain the processing delay and transmission delay of the nodes for data according to steps (4) and (5).
[0084] When selecting nodes, the impact of bandwidth heterogeneity and node computing power on data repair delay is taken into consideration, and the node with the shortest sum of data processing delay and data transmission delay is selected, that is, the node with the shortest repair time. The node selected by the above method participates in data regeneration to recover the failed data.
[0085] In this embodiment, (8, 5, 3) erasure codes are used, and the original file is encoded with a regeneration code to obtain 8 data blocks and stored on different network nodes. When Node6 data transmission is lost, the new node can recover the lost data by storing data in any 3 data blocks.
[0086] Figure 1 This is a schematic diagram of the network topology structure based on this embodiment. There are 9 limited nodes in the network, and each node can communicate with each other. Node8 is a new node responsible for data regeneration. The disk I / O, CPU, memory, and chip performance parameter values of each storage node are obtained through system monitoring, and corresponding weights are assigned. The weights are 30%, 25%, 25%, and 20% respectively, and the computing power of the node is calculated.
[0087]
[0088] At the same time, the available bandwidth between each storage node is simulated, see Table 1, and the transmission delay between nodes is calculated.
[0089] Table 1 Transmission delay between nodes
[0090] Node1 Node2 Node3 Node4 Node5 Node6 Node7 Node8 Node9 Node1 0 10.5071 10.4652 0.8485 1.1314 10.3846 10.3730 1.9799 0.5692 Node2 10.5071 0 0.2828 10.4038 10.3846 1.1314 1.4142 10.3730 9.2363 Node3 10.4652 0.2828 0 10.3846 10.3730 0.8485 1.1314 10.3846 1.5381 Node4 0.8485 10.4038 10.3846 0 0.2828 10.3730 10.3846 1.1314 8.9543 Node5 1.1314 10.3846 10.3730 0.2828 0 10.3846 10.4038 0.8485 1.6912 Node6 10.3846 1.1314 0.8485 10.3730 10.3846 0 0.2828 10.4652 10.9216 Node7 10.3730 1.4142 1.1314 10.3846 10.4038 0.2828 0 10.5071 10.1546 Node8 1.9799 10.3730 10.3846 1.1314 0.8485 10.4652 10.5071 0 6.5023 Node9 0.5692 9.2363 1.5381 8.9543 1.6912 10.9216 10.1546 6.5023 0
[0091] At time t1, the Node9 node in the system loses connection. The heartbeat detection mechanism sends heartbeat packets to each node to obtain the global load status of all candidate supply nodes at this moment, such as Figure 2 .
[0092] In this embodiment, a load change threshold Δ is preset. preset =1 and load change lower limit Δ lower =0, dynamically adjust the node computing capacity according to the relative changes in the load of the candidate supply nodes.
[0093] A1'=A1×γ=11.9*0.1=1.19
[0094]
[0095] A4'=A4×γ=6.8*0.1=0.68
[0096]
[0097] Assume that the computational effort required for each node to repair data is fixed. According to the encoding method of the regenerating code, the storage capacity of each node is α. Then, the processing delay of the node for the data can be expressed as:
[0098]
[0099] In one embodiment, α = 140 mb. In the star repair model, the supply node directly sends the local data to the new node. After the new node receives all the data, it performs encoding calculations on the data to restore the failed data, forming a star topology structure with the new node as the center and the supply nodes as the branches. Based on the star repair model of the present invention, the computing capabilities of each supply candidate node are obtained according to the above steps, the processing delay for the data is calculated, and at the same time, according to the network transmission distance from the supply candidate node to the new node, the transmission delay from the candidate supply node to the new node is calculated. Finally, the 3 nodes with the shortest time consumption are selected to participate in the repair work.
[0100] Figure 3 FIG. is a schematic diagram of the star repair model based on the present invention. The method for selecting the optimal supply node based on network distance and load balancing will preferentially select Node5, Node3, and Node2 as supply nodes, and try to avoid nodes Node1 and Node4 with large load changes, greatly improving the repair speed.
[0101] The data repair process of the tree repair model constructs an optimal regenerating tree based on the weights of each edge. The new node is the root node. The supply node transmits the data block it stores to its parent node. After the parent node receives the data of all its child nodes, it first performs encoding preprocessing, and then continues to transmit the encoded result to its parent node until the root node receives the required data and performs a linear combination on the received data to restore the lost data. The shape of this topology structure is like an inverted tree, with the tree root at the top and the branches below the tree root.
[0102] Figure 4 FIG. is a schematic diagram of the tree repair model based on the present invention. The goal is to minimize the repair time of the topology tree. The specific idea is to obtain the computing capabilities of each supply candidate node according to the previous steps, calculate the processing delay for the data, and combine the network transmission distance. Then, traverse the nodes with the shortest transmission and processing delays for the data in the repair path set in sequence, and add the selected nodes to the tree structure until 3 nodes are added to the tree structure. This method will preferentially select Node3 with a shorter network distance and remaining computing power as the child node of Node2. After Node2 receives the data block transmitted by Node3, it performs encoding preprocessing. After Node8 receives the data transmitted by Node2 and Node5, it performs data regeneration work, thereby forming a tree structure.
[0103] In one embodiment, α = 240 mb. In the star repair model, the method of this embodiment preferentially selects Node5, Node3, and Node2 as supply nodes. In the tree repair model, the method of this embodiment selects Node5, Node3, and Node2 as supply nodes. Node5 transmits data upward as a child node of Node8, Node3 transmits preprocessed data upward as a child node of Node2, and after Node2 receives all the data, it integrates and transmits it to the new node Node8 to complete the data regeneration work. Figure 5 It is a comparison chart of repair time delays under different data block sizes based on the present invention.
[0104] As Figure 7 shown, this embodiment discloses a system based on a dynamic node selection method for data repair in distributed storage, which includes the following modules:
[0105] Node computing power modeling module: used to model the computing power of nodes;
[0106] Node load change detection module: used to detect changes in node load;
[0107] Node computing power adjustment module: used to dynamically adjust the computing power of nodes;
[0108] Processing delay calculation module: used to calculate the processing delay;
[0109] Transmission delay calculation module: used to calculate the transmission delay;
[0110] Node selection module: selects nodes based on network distance and load balancing.
[0111] For a data repair request, the technical solution proposed by the present invention can adaptively select different nodes to participate in the data regeneration work according to the change of node load, effectively avoiding nodes without enough processing margin and avoiding overloading high-processing nodes. In the face of data loss, the present invention takes into account the load balance between nodes, realizes the adaptive selection of nodes, reasonably plans the node selection strategy, and improves the repair efficiency.
[0112] The specific embodiments of the present invention have been described above. It should be noted that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A dynamic node selection method for data repair in distributed storage, characterized in that According to the following steps: (1) Model the computing power of nodes; (2) Detect the changes in node load; (3) Dynamically adjust the computing power of nodes; (4) Calculate the processing delay; (5) Calculate the transmission delay; (6) Select nodes based on network distance and load balancing; Step (1) is specifically as follows: Select k factors that affect the computing power of nodes, and use x1, x2, x3, …, x k to represent. Different weights c1, c2, c3, …, c k are assigned to the selected main factors, where c1 + c2 + c3 + … + c k = 1. Larger weights are assigned to factors that have a greater impact on the computing power of nodes, and smaller weights are assigned to factors that have a smaller impact on the computing power of nodes. The computing power of each node is expressed as: Step (2) is specifically as follows: Send heartbeat packets to each node regularly through the heartbeat detection mechanism to obtain the global load conditions of all nodes at this moment; assume that the load detection volume of a certain node in the current cycle is p t2 , and the load detection volume in the previous cycle is p t1 , considering the impact of load changes on the computing power of the node; the relative load change is expressed as: Step (3) is specifically as follows: Preset a storage load change threshold Δ preset and a lower limit value of load change Δ lower , and consider the impact of the load change amount on the node computing power; when the relative load change Δ is greater than or equal to the load change threshold Δ preset , it indicates that the load of the node is relatively large during this period, and the change amount exceeds the load of the previous period. The increase in load will affect the processing capacity of the node itself. At this time, this node is no longer suitable to be a candidate node for the supplier to participate in the data repair work; update the computing power of the node. γ is the conversion coefficient set when the relative load change is greater than or equal to the load change threshold, as shown in the following formula: A' = A × γ When the relative change in load Δ is between the load change threshold Δ preset and the lower limit value of load change Δ lower it indicates that the load of the node has increased by a certain amount during this period, but has not reached the preset proximity value of load change; Update the computing power of the node, where ω is Δ lower ≤ Δ ≤ Δ preset is the conversion coefficient set when, as shown in the following formula: A' = A × (1 - Δ) × ω When the relative change in load Δ is less than the lower limit of load change Δ lower , it indicates that the load of the node decreases during this period, and this node has spare computing power; update the computing power of the node, where Δ is the conversion coefficient set when Δ ≤ Δ lower , as shown in the following formula: Dynamically adjust the computing power of nodes by detecting the amount of load change.
2. The dynamic node selection method for data repair in the distributed storage according to claim 1, characterized in that Step (4) is specifically as follows: Assume that the amount of computation required for each node to repair data is fixed. According to the encoding method of the regenerating code, let the storage capacity of each node be α. Then the processing delay of the node for data is expressed as where δ is the conversion coefficient and ms is milliseconds.
3. The dynamic node selection method for data repair in the distributed storage according to claim 2, characterized in that Step (5) is specifically as follows: The transmission delay is expressed by the following formula, T ij represents the transmission delay between nodes ij, k is the length of the channel, v is the transmission rate, and ξ is the conversion coefficient.
4. The dynamic node selection method for data repair in the distributed storage according to claim 3, wherein Step (6) is specifically as follows: When node i fails, first obtain the load P of the candidate supplier nodes in the current cycle t2 and the load P in the previous cycle t1 , complete the update of the computing power of the candidate supplier nodes through steps (2) and (3), and obtain the processing delay and transmission delay of the nodes for data according to steps (4) and (5); When selecting nodes, select the node with the shortest sum of the processing data delay and the transmission data delay, that is, the node with the shortest repair time.
5. A system for a dynamic node selection method for data repair in the distributed storage according to any one of claims 1-4, characterized in that It includes the following modules: Node computing power modeling module: used to model the computing power of nodes; Node load change detection module: used to detect the changes in node load; Node computing power adjustment module: used to dynamically adjust the computing power of nodes; Processing delay calculation module: used to calculate the processing delay; Transmission delay calculation module: used to calculate the transmission delay; Node selection module: select nodes based on network distance and load balancing.