A distributed storage heterogeneous data recovery method and system based on cloud-edge collaboration
By comprehensively characterizing the security, rate and anti-interference of the data recovery path, selecting the optimal recovery path for data recovery, the problem of lack of dynamicity and continuity in the existing technology is solved, and efficient and high-quality data recovery effect is achieved.
Patent Information
- Application Number
- CN202510175671.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The prior art lacks a comprehensive characterization of the transmission link security, anti-interference and data transmission rate between each node and the lost data node during data recovery, resulting in a lack of dynamicity and continuity in data recovery, which can easily lead to low data quality, long recovery time and high energy consumption.
Through interactive information, determine the working status of the node, calculate the transmission link security value, establish a link safety network diagram, calculate the node data rate and noise error distribution diagram based on the random walk mode, and select the optimal recovery path for data recovery using the A* path planning algorithm.
It improves the dynamic and continuity of data recovery, improves the quality and efficiency of data recovery, reduces recovery time and energy consumption, and enhances the availability and user experience of network services.
Smart Images

Figure CN119645739B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data recovery technology, and in particular to a distributed storage heterogeneous data recovery method and system based on cloud-edge collaboration. Background Art
[0002] Data is one of the most important core assets in business operations. Many business decisions and customer services rely on complete and effective data. Therefore, when data is damaged or lost, if decisions and operations are made based on damaged data, errors will occur in business decisions and operations. Data recovery can ensure data integrity and provide accurate information support for business decisions and operations. Therefore, the recovery and processing operations after data damage and loss are of great research significance.
[0003] At present, the Chinese invention patent with application number CN117667494A discloses a data recovery method, system, device and storage medium for a distributed storage system, including when there is at least one abnormal storage target in the distributed storage system, determining the current maximum amount of data to be recovered in the distributed storage system within a preset time length; determining the target storage node corresponding to the abnormal storage target, and the reference storage node corresponding to the target storage node; determining the data recovery speed limit value corresponding to each target storage node according to the current maximum amount of data to be recovered; and controlling each target storage node to read the copy of the data to be recovered from the corresponding reference storage node according to the data recovery speed limit. However, the relevant technology does not comprehensively characterize the security, anti-interference and data transmission rate of the transmission link between each node and the node where the data is lost during data recovery. The current recovery of lost data lacks dynamism and continuity, and there is no optimal path selection for data recovery based on the comprehensive characterization, which easily leads to low data quality, long recovery time and high recovery energy consumption in the data recovery process. Summary of the invention
[0004] The technical problem solved by the present invention is that: there is no comprehensive characterization of the security, anti-interference and data transmission rate of the transmission link between each node and the node that lost the data during data recovery, the current recovery of lost data lacks dynamism and continuity, and there is no optimal path selection for data recovery based on the comprehensive characterization, which easily leads to low data quality, long recovery time and high recovery energy consumption in the data recovery process.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: a distributed storage heterogeneous data recovery method based on cloud-edge collaboration, the specific steps include:
[0006] Step S1, determining the working state of the first node through interaction information between the current first node and other nodes, locating the node of the fault data of the first node according to the working state and calculating the capacity of the fault data to obtain a first data volume of the fault data;
[0007] Step S2, calculating the transmission link security value of the first node and other nodes, and establishing a link security network diagram of the first node and other nodes;
[0008] Step S3, calculating a node data rate distribution graph and a node noise error distribution graph of the first node and other nodes based on a random walk mode;
[0009] Step S4, using the A* path planning algorithm to calculate the final gain values of the link safety network diagram, the node data rate distribution diagram, and the node noise error distribution diagram, and selecting the transmission link with the largest final gain as the optimal recovery path;
[0010] Step S5, preprocessing the fault data to obtain the reorganization parameters of the first node, and selecting the optimal recovery path for data recovery based on the reorganization parameters of the preprocessed fault data to obtain recovery data;
[0011] Step S6, calculating hash values of the recovery data and the fault data based on a hash algorithm to obtain a first hash value and a second hash value, comparing the first hash value and the second hash value to obtain a comparison result, and sending a recovery process instruction to the other nodes according to the comparison result.
[0012] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, step S1 specifically includes:
[0013] Step S11, obtaining interaction information between the first node and other nodes, wherein the interaction information includes a node ID, a node timestamp, a node load size, and a remaining size of a node storage space;
[0014] Step S12, other nodes count the number of times the interaction information is received within a unit time, and compare the number of times with a preset threshold of the number of times of receiving. If the number of times of receiving is less than the threshold of the number of times of receiving, it is determined that the working state of the first node is abnormal; if the number of times of receiving is greater than the threshold of the number of times of receiving, it is determined that the working state of the first node is normal.
[0015] Step S13, when the working state of the first node is abnormal, the fault data of the first node is node located, the IDs of other nodes of the copy data corresponding to the first node are obtained, and the capacity of the fault data is calculated to obtain the first data volume of the fault data.
[0016] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, step S2 specifically includes:
[0017] Step S21, collecting network link status data between the first node and other nodes, wherein the network link status data includes four safety indicator data: throughput, channel utilization, delay time, and delay jitter rate;
[0018] Step S22, preprocessing the four safety index data, wherein the preprocessing operation includes standardization and discretization of missing filling data;
[0019] Step S23, using the AHP hierarchical analysis algorithm to calculate the weights of the four safety index data;
[0020] Step S24, calculating the network entropy value of each security indicator data, and the calculation expression is as follows:
[0021] ;
[0022] in, represents the broadband network entropy value, represents the weight of security index data i, represents the value of index i, and n represents the total number of safety indicators;
[0023] Step S25, for each safety indicator data Take the weighted average, denoted as , calculate the transmission link security value , and its calculation expression is as follows:
[0024] ;
[0025] in, Indicates the transmission link security value;
[0026] Step S26: the transmission link security value The levels are divided as follows:
[0027] 0.9 to 1 is the first safety level, corresponding to the lowest risk level;
[0028] 0.7~0.9 is the second safety level, corresponding to the second lowest risk level;
[0029] 0.6~0.7 is the third safety level, corresponding to the second highest risk level;
[0030] Below 0.6 is the fourth safety level, corresponding to the highest risk level;
[0031] Step S27: Obtain the transmission link security value between the first node and other nodes , establish a link security network diagram between the first node and other nodes.
[0032] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, step S3 specifically includes:
[0033] Step S31, according to the random walk mode, simulate the process of the first node transmitting data to other nodes, randomly select a target node, and record the transmission data between the first node and the target node, where the transmission data includes the transmission node, the amount of data transmitted and the first transmission time;
[0034] Step S32, calculating a node data rate between the first node and the target node according to the recorded transmission data, wherein the node data rate is equal to the data volume divided by the first transmission time;
[0035] Step S33, grouping the calculated node data rates according to the node pairs, and drawing a node data rate distribution graph, wherein the node data rate distribution graph shows the node data rate distribution between the first node and other nodes;
[0036] Step S34, when the first node and other nodes perform data transmission, simulate the random generation of noise; record the noise data of the first node and the target node, the noise data including the size of the noise and the second transmission time when the noise is generated;
[0037] Step S35, calculating a node noise error between the first node and the target node according to the recorded noise data, wherein the node noise error is equal to the noise magnitude divided by the second transmission time;
[0038] Step S36, grouping the calculated node noise errors according to node pairs, and drawing a node noise error distribution diagram, wherein the node noise error distribution diagram shows the node noise error distribution between the first node and other nodes.
[0039] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, step S4 specifically includes:
[0040] Step S41, selecting the transmission link security value on the link security network diagram The smallest transmission link is the first transmission path;
[0041] Step S42, selecting a transmission link with the largest node data rate in the node data rate distribution graph as a second transmission path;
[0042] Step S43, selecting a transmission link with the smallest node noise error in the node noise error distribution graph as a third transmission path;
[0043] Step S44, calculating the final gain values of the first transmission path, the second transmission path and the third transmission path based on the A* path planning algorithm, and selecting the transmission link with the largest final gain value as the optimal recovery path.
[0044] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, step S5 specifically includes:
[0045] Step S51, preprocessing the fault data, the preprocessing includes data integration, data cleaning and data conversion, the data conversion includes data format conversion, data type conversion and data unit conversion;
[0046] Step S52, obtaining the reorganization parameters of the first node, the reorganization parameters including the level of the redundant array of independent disks, the first data volume of the failed data, the block size and the verification mode and the data direction;
[0047] Step S53: The pre-processed fault data is restored by selecting the optimal recovery path based on the reorganization parameters to obtain restored data.
[0048] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, step S6 specifically includes:
[0049] Step S61, obtaining a preset hash function, inputting the restored data into the hash function, and obtaining a first hash value;
[0050] Step S62, inputting the fault data into the hash function to obtain a second hash value;
[0051] Step S63, comparing the first hash value and the second hash value, if the first hash value is equal to the second hash value, the first node sends a recovery process instruction with a recovery success flag 1 to other nodes;
[0052] If the first hash value is not equal to the second hash value, the first node sends a recovery process instruction with a recovery failure flag of 0 to other nodes, and restarts the node data recovery task based on different transmission links.
[0053] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, the specific steps of calculating the final gain values of the first transmission path, the second transmission path, and the third transmission path based on the A* path planning algorithm are as follows:
[0054] The first path weight R1, the second path weight R2 and the third path weight R3 of the first transmission path, the second transmission path and the third transmission path are obtained, and the calculation expression for calculating the final gain value is as follows:
[0055] Final gain value = (R1 × transmission link safety value )+(R2×node data rate)+(R3×node noise error).
[0056] As a preferred solution of the distributed storage heterogeneous data recovery method based on cloud-edge collaboration described in the present invention, calculating the capacity of the fault data specifically includes:
[0057] Counting the number of first redundant copies of the first node and other nodes in real time, and retrieving the preset number of second redundant copies of the first node and other nodes;
[0058] When the number of the first redundant copies is less than the number of the second redundant copies, the first node data block size, the total data volume and the preset number of the second redundant copies are retrieved;
[0059] The capacity of the fault data = (the size of the data block of the first node × the total amount of data on the first node) / the number of the second redundant copies.
[0060] A distributed storage heterogeneous data recovery system based on cloud-edge collaboration, including a node working status monitoring module, a node data calculation module, a transmission link planning module and a data integrity verification module;
[0061] The node working status monitoring module is used to calculate the working status of the first node and the size of the lost data capacity;
[0062] The node data calculation module is used to calculate the transmission link security value, node data rate and node noise error of the first node and other nodes;
[0063] The transmission link planning module is used to select a transmission link with the smallest final gain value of the transmission links of the first node and other nodes as the optimal recovery path;
[0064] The data integrity verification module is used to verify whether the restored data is consistent with the fault data lost by the first node.
[0065] The beneficial effects of the present invention are as follows: by establishing a link security network diagram, a node data rate distribution diagram and a node noise error distribution diagram between nodes, and calculating the gain value of data recovery between nodes based on a path planning algorithm, network operators can be helped to promptly detect and respond to network attacks, thereby ensuring the availability of network services and reducing the risk of the network being subjected to large-scale network attacks. The attack bandwidth of network attackers on network infrastructure can be limited, reducing the degree of damage. In a secure wide link, the user's network experience is improved, and the timeliness and accuracy of data transmission are ensured. By calculating the noise characteristics between nodes, it is helpful to optimize the data recovery algorithm, and a recovery algorithm that is more suitable for the current noise environment can be designed to improve the accuracy and efficiency of data recovery. By calculating the transmission rate between nodes, the performance of different transmission paths can be evaluated, thereby selecting the optimal path for data recovery, which helps to evaluate the degree of data damage, optimize the recovery algorithm, improve the recovery success rate, determine the recovery strategy, optimize resource allocation, and improve the recovery quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 A basic flow chart of a distributed storage heterogeneous data recovery method based on cloud-edge collaboration is provided for one embodiment of the present invention. DETAILED DESCRIPTION
[0067] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0068] Example, see Figure 1 , is an embodiment of the present invention, and provides a distributed storage heterogeneous data recovery method based on cloud-edge collaboration, and the specific steps include:
[0069] Step S1, determining the working state of the first node through interaction information between the current first node and other nodes, locating the node of the fault data of the first node according to the working state and calculating the capacity of the fault data to obtain a first data volume of the fault data;
[0070] Step S2, calculating the transmission link security value of the first node and other nodes, and establishing a link security network diagram of the first node and other nodes;
[0071] Step S3, calculating a node data rate distribution graph and a node noise error distribution graph of the first node and other nodes based on a random walk mode;
[0072] Step S4, using the A* path planning algorithm to calculate the final gain values of the link safety network diagram, the node data rate distribution diagram, and the node noise error distribution diagram, and selecting the transmission link with the largest final gain as the optimal recovery path;
[0073] Step S5, preprocessing the fault data to obtain the reorganization parameters of the first node, and selecting the optimal recovery path for data recovery based on the reorganization parameters of the preprocessed fault data to obtain recovery data;
[0074] Step S6, calculating hash values of the recovery data and the fault data based on a hash algorithm to obtain a first hash value and a second hash value, comparing the first hash value and the second hash value to obtain a comparison result, and sending a recovery process instruction to the other nodes according to the comparison result.
[0075] Step S1 specifically includes:
[0076] Step S11, obtaining interaction information between the first node and other nodes, wherein the interaction information includes a node ID, a node timestamp, a node load size, and a remaining size of a node storage space;
[0077] Step S12, other nodes count the number of times the interaction information is received within a unit time, and compare the number of times with a preset threshold of the number of times of receiving. If the number of times of receiving is less than the threshold of the number of times of receiving, it is determined that the working state of the first node is abnormal; if the number of times of receiving is greater than the threshold of the number of times of receiving, it is determined that the working state of the first node is normal;
[0078] Step S13, when the working state of the first node is abnormal, the fault data of the first node is node located, the IDs of other nodes of the copy data corresponding to the first node are obtained, and the capacity of the fault data is calculated to obtain the first data volume of the fault data.
[0079] In this embodiment, when a data node receives less than the reception threshold of interactive information within a unit time, which is 9s and 2, the master node will mark it as a faulty or offline state, which helps the system to respond quickly and take corresponding recovery measures. Once a data node failure is detected, the system can immediately start the data recovery process and the data block replication process to ensure data redundancy and availability.
[0080] Step S2 specifically includes:
[0081] Step S21, collecting network link status data between the first node and other nodes, wherein the network link status data includes four safety indicator data: throughput, channel utilization, delay time, and delay jitter rate;
[0082] Step S22, preprocessing the four safety index data, wherein the preprocessing operation includes standardization and discretization of missing filling data;
[0083] Step S23, using the AHP hierarchical analysis algorithm to calculate the weights of the four safety index data;
[0084] Step S24, calculating the network entropy value of each security indicator data, and the calculation expression is as follows:
[0085] ;
[0086] in, represents the broadband network entropy value, represents the weight of security index data i, represents the value of index i, and n represents the total number of safety indicators;
[0087] Step S25, for each safety indicator data Take the weighted average, denoted as , calculate the transmission link security value , and its calculation expression is as follows:
[0088] ;
[0089] in, Indicates the transmission link security value;
[0090] Step S26: the transmission link security value The levels are divided as follows:
[0091] 0.9 to 1 is the first safety level, corresponding to the lowest risk level;
[0092] 0.7~0.9 is the second safety level, corresponding to the second lowest risk level;
[0093] 0.6~0.7 is the third safety level, corresponding to the second highest risk level;
[0094] Below 0.6 is the fourth safety level, corresponding to the highest risk level;
[0095] Step S27: Obtain the transmission link security value between the first node and other nodes , establish a link security network diagram between the first node and other nodes.
[0096] In this embodiment, establishing a link security network map can help network operators to promptly detect and respond to network attacks, thereby ensuring the availability of network services and reducing the risk of the network being subjected to large-scale network attacks. It can limit the attack bandwidth of network attackers on network infrastructure and reduce the degree of damage. In a secure wide link, it can improve the user's network experience and ensure the timeliness and accuracy of data transmission.
[0097] Step S3 specifically includes:
[0098] Step S31, according to the random walk mode, simulate the process of the first node transmitting data to other nodes, randomly select a target node, and record the transmission data between the first node and the target node, where the transmission data includes the transmission node, the amount of data transmitted and the first transmission time;
[0099] Step S32, calculating a node data rate between the first node and the target node according to the recorded transmission data, wherein the node data rate is equal to the data volume divided by the first transmission time;
[0100] Step S33, grouping the calculated node data rates according to the node pairs, and drawing a node data rate distribution graph, wherein the node data rate distribution graph shows the node data rate distribution between the first node and other nodes;
[0101] Step S34, when the first node and other nodes perform data transmission, simulate the random generation of noise; record the noise data of the first node and the target node, the noise data including the size of the noise and the second transmission time when the noise is generated;
[0102] Step S35, calculating a node noise error between the first node and the target node according to the recorded noise data, wherein the node noise error is equal to the noise magnitude divided by the second transmission time;
[0103] Step S36, grouping the calculated node noise errors according to node pairs, and drawing a node noise error distribution diagram, wherein the node noise error distribution diagram shows the node noise error distribution between the first node and other nodes.
[0104] In this embodiment, by establishing a node noise error distribution map and a node data rate distribution map, calculating the noise characteristics between nodes, it is helpful to optimize the data recovery algorithm. A recovery algorithm that is more suitable for the current noise environment can be designed to improve the accuracy and efficiency of data recovery. By calculating the transmission rate between nodes, the performance of different transmission paths can be evaluated, so as to select the optimal path for data recovery, which is helpful to evaluate the degree of data damage, optimize the recovery algorithm, improve the recovery success rate, determine the recovery strategy, optimize resource allocation, and improve the recovery quality. Step S4 specifically includes:
[0105] Step S41, selecting the transmission link security value on the link security network diagram The smallest transmission link is the first transmission path;
[0106] Step S42, selecting a transmission link with the largest node data rate in the node data rate distribution graph as a second transmission path;
[0107] Step S43, selecting a transmission link with the smallest node noise error in the node noise error distribution graph as a third transmission path;
[0108] Step S44, calculating the final gain values of the first transmission path, the second transmission path and the third transmission path based on the A* path planning algorithm, and selecting the transmission link with the largest final gain value as the optimal recovery path.
[0109] In this embodiment, the link safety network diagram between nodes, the node data rate distribution diagram and the node noise error distribution diagram are integrated, and the transmission link with the largest final gain value is selected as the optimal recovery path. This can take into account both the recovery speed and recovery quality of data recovery between nodes, thereby improving the overall availability and efficiency of data recovery.
[0110] Step S5 specifically includes:
[0111] Step S51, preprocessing the fault data, the preprocessing includes data integration, data cleaning and data conversion, the data conversion includes data format conversion, data type conversion and data unit conversion;
[0112] Step S52, obtaining the reorganization parameters of the first node, the reorganization parameters including the level of the redundant array of independent disks, the first data volume of the failed data, the block size and the verification mode and the data direction;
[0113] Step S53: The pre-processed fault data is restored by selecting the optimal recovery path based on the reorganization parameters to obtain restored data.
[0114] In this embodiment, obtaining the reorganization parameters of the abnormal working state node helps to understand the storage structure and organization of the data before the failure, thereby improving the success rate of data recovery. According to the reorganization parameters of the failed node, the area where the damaged data is located can be quickly located, avoiding the traversal scanning of all data. It helps to shorten the processing time of data recovery, so that users can recover lost data faster and resume normal work.
[0115] Step S6 specifically includes:
[0116] Step S61, obtaining a preset hash function, inputting the restored data into the hash function, and obtaining a first hash value;
[0117] Step S62, inputting the fault data into the hash function to obtain a second hash value;
[0118] Step S63, comparing the first hash value and the second hash value, if the first hash value is equal to the second hash value, the first node sends a recovery process instruction with a recovery success flag 1 to other nodes;
[0119] If the first hash value is not equal to the second hash value, the first node sends a recovery process instruction with a recovery failure flag of 0 to other nodes, and restarts the node data recovery task based on different transmission links.
[0120] In this embodiment, the preset hash function is SHA-256, which can generate a hash value of a fixed length, making it very fast to verify the consistency of data. It is only necessary to compare the hash values of two data to determine whether they are the same, without having to compare the entire data set byte by byte. The storage space required to store the hash value is much smaller than the storage space required to store the original data, and it can be applied to distributed systems that need to frequently verify data consistency.
[0121] The specific steps of calculating the final gain values of the first transmission path, the second transmission path, and the third transmission path based on the A* path planning algorithm are as follows:
[0122] The first path weight R1, the second path weight R2 and the third path weight R3 of the first transmission path, the second transmission path and the third transmission path are obtained, and the calculation expression for calculating the final gain value is as follows:
[0123] Final gain value = (R1 × transmission link safety value )+(R2×node data rate)+(R3×node noise error).
[0124] In this embodiment, the first path weight R1 is 0.6, the second path weight R2 is 0.2, and the third path weight R3 is 0.2. The priority of the first transmission path is greater than that of the second transmission path, and the priority of the second transmission path is greater than that of the third transmission path, which can ensure that the data loss rate is always low and improve the overall data recovery performance.
[0125] Calculating the capacity of the fault data specifically includes:
[0126] Counting the number of first redundant copies of the first node and other nodes in real time, and retrieving the preset number of second redundant copies of the first node and other nodes;
[0127] When the number of the first redundant copies is less than the number of the second redundant copies, the first node data block size, the total data volume and the preset number of the second redundant copies are retrieved;
[0128] The capacity of the fault data = (the size of the data block of the first node × the total amount of data on the first node) / the number of the second redundant copies.
[0129] A distributed storage heterogeneous data recovery system based on cloud-edge collaboration, including a node working status monitoring module, a node data calculation module, a transmission link planning module and a data integrity verification module;
[0130] The node working status monitoring module is used to calculate the working status of the first node and the size of the lost data capacity;
[0131] The node data calculation module is used to calculate the transmission link security value, node data rate and node noise error of the first node and other nodes;
[0132] The transmission link planning module is used to select a transmission link with the smallest final gain value of the transmission links of the first node and other nodes as the optimal recovery path;
[0133] The data integrity verification module is used to verify whether the restored data is consistent with the fault data lost by the first node.
[0134] By establishing a link safety network diagram, a node data rate distribution diagram, and a node noise error distribution diagram between nodes, and calculating the gain value of data recovery between nodes based on a path planning algorithm, network operators can be helped to discover and respond to network attacks in a timely manner, thereby ensuring the availability of network services, reducing the risk of the network being subjected to large-scale network attacks, and limiting the attack bandwidth of network attackers on network infrastructure, reducing the degree of damage, and improving the user's network experience in a secure wide link, ensuring the timeliness and accuracy of data transmission, and by calculating the noise characteristics between nodes, it is helpful to optimize the data recovery algorithm. A recovery algorithm that is more suitable for the current noise environment can be designed to improve the accuracy and efficiency of data recovery. By calculating the transmission rate between nodes, the performance of different transmission paths can be evaluated, thereby selecting the optimal path for data recovery, which helps to evaluate the degree of data damage, optimize the recovery algorithm, improve the recovery success rate, determine the recovery strategy, optimize resource allocation, and improve the quality of recovery. It should be understood by those skilled in the art that the embodiments of the present invention can be provided as a method, system, or computer program product. Therefore, the present invention can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. The storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable red-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which is implemented in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0135] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A distributed storage heterogeneous data recovery method based on cloud-edge collaboration, characterized in that: include: Step S1, determining the working state of the first node through the interaction information between the current first node and other nodes, performing node location on the fault data of the first node and calculating a first data volume of the fault data; Step S2, calculating the transmission link security value of the first node and other nodes, and establishing a link security network diagram of the first node and other nodes; Step S3, calculating a node data rate distribution graph and a node noise error distribution graph of the first node and other nodes based on a random walk mode; Step S4, using the A* path planning algorithm to calculate the final gain values of the link safety network diagram, the node data rate distribution diagram, and the node noise error distribution diagram, and selecting the transmission link with the largest final gain as the optimal recovery path; Step S5, preprocessing the fault data to obtain the reorganization parameters of the first node, and selecting the optimal recovery path for data recovery based on the reorganization parameters of the preprocessed fault data to obtain recovery data; Step S6, calculating hash values of the recovery data and the fault data based on a hash algorithm to obtain a first hash value and a second hash value, comparing the first hash value and the second hash value to obtain a comparison result, and sending a recovery process instruction to the other nodes according to the comparison result; Step S3 specifically includes: Step S31, according to the random walk mode, simulate the process of the first node transmitting data to other nodes, randomly select a target node, and record the transmission data between the first node and the target node, where the transmission data includes the transmission node, the amount of data transmitted and the first transmission time; Step S32, calculating a node data rate between the first node and the target node according to the recorded transmission data, wherein the node data rate is equal to the data volume divided by the first transmission time; Step S33, grouping the calculated node data rates according to the node pairs, and drawing a node data rate distribution graph, wherein the node data rate distribution graph shows the node data rate distribution between the first node and other nodes; Step S34, when the first node and other nodes perform data transmission, simulate the random generation of noise; record the noise data of the first node and the target node, the noise data including the size of the noise and the second transmission time when the noise is generated; Step S35, calculating a node noise error between the first node and the target node according to the recorded noise data, wherein the node noise error is equal to the noise magnitude divided by the second transmission time; Step S36, grouping the calculated node noise errors according to node pairs, and drawing a node noise error distribution diagram, wherein the node noise error distribution diagram shows the node noise error distribution between the first node and other nodes.
2. The distributed storage heterogeneous data recovery method based on cloud-edge collaboration as described in claim 1 is characterized by: Step S1 specifically includes: Step S11, obtaining interaction information between the first node and other nodes, wherein the interaction information includes a node ID, a node timestamp, a node load size, and a remaining size of a node storage space; Step S12, other nodes count the number of times the interaction information is received within a unit time, and compare the number of times with a preset threshold of the number of times of receiving. If the number of times of receiving is less than the threshold of the number of times of receiving, it is determined that the working state of the first node is abnormal; if the number of times of receiving is greater than the threshold of the number of times of receiving, it is determined that the working state of the first node is normal. Step S13, when the working state of the first node is abnormal, the fault data of the first node is node located, the IDs of other nodes of the copy data corresponding to the first node are obtained, and the capacity of the fault data is calculated to obtain the first data volume of the fault data.
3. The distributed storage heterogeneous data recovery method based on cloud-edge collaboration as claimed in claim 1, characterized in that: Step S2 specifically includes: Step S21, collecting network link status data between the first node and other nodes, wherein the network link status data includes four safety indicator data: throughput, channel utilization, delay time, and delay jitter rate; Step S22, preprocessing the four safety index data, wherein the preprocessing operation includes standardization and discretization of missing filling data; Step S23, using the AHP hierarchical analysis algorithm to calculate the weights of the four safety index data; Step S24, calculating the network entropy value of each security indicator data, and the calculation expression is as follows: ; in, represents the broadband network entropy value, represents the weight of the safety index data i, represents the value of index i, and n represents the total number of safety indicators; Step S25, for each safety indicator data Take the weighted average, denoted as 1. Calculate the transmission link security value 2, its calculation expression is as follows: ; in, 2 indicates the transmission link security value; Step S26: the transmission link security value 2. The logical levels of division are as follows: 0.9 to 1 is the first safety level, corresponding to the lowest risk level; 0.7~0.9 is the second safety level, corresponding to the second lowest risk level; 0.6~0.7 is the third safety level, corresponding to the second highest risk level; Below 0.6 is the fourth safety level, corresponding to the highest risk level; Step S27: Obtain the transmission link security value between the first node and other nodes 2. Establish a link security network diagram between the first node and other nodes.
4. The distributed storage heterogeneous data recovery method based on cloud-edge collaboration as claimed in claim 1, characterized in that: Step S4 specifically includes: Step S41, selecting the transmission link security value on the link security network diagram 2 The smallest transmission link is the first transmission path; Step S42, selecting a transmission link with the largest node data rate in the node data rate distribution graph as a second transmission path; Step S43, selecting a transmission link with the smallest node noise error in the node noise error distribution graph as a third transmission path; Step S44, calculating the final gain values of the first transmission path, the second transmission path and the third transmission path based on the A* path planning algorithm, and selecting the transmission link with the largest final gain value as the optimal recovery path.
5. The distributed storage heterogeneous data recovery method based on cloud-edge collaboration as claimed in claim 1, characterized in that: Step S5 specifically includes: Step S51, preprocessing the fault data, the preprocessing includes data integration, data cleaning and data conversion, the data conversion includes data format conversion, data type conversion and data unit conversion; Step S52, obtaining the reorganization parameters of the first node, the reorganization parameters including the level of the redundant array of independent disks, the first data volume of the failed data, the block size and the verification mode and the data direction; Step S53: The pre-processed fault data is restored by selecting the optimal recovery path based on the reorganization parameters to obtain restored data.
6. The distributed storage heterogeneous data recovery method based on cloud-edge collaboration as claimed in claim 1, characterized in that: Step S6 specifically includes: Step S61, obtaining a preset hash function, inputting the restored data into the hash function, and obtaining a first hash value; Step S62, inputting the fault data into the hash function to obtain a second hash value; Step S63, comparing the first hash value and the second hash value, if the first hash value is equal to the second hash value, the first node sends a recovery process instruction with a recovery success flag 1 to other nodes; If the first hash value is not equal to the second hash value, the first node sends a recovery process instruction with a recovery failure flag of 0 to other nodes, and restarts the node data recovery task based on different transmission links.
7. The distributed storage heterogeneous data recovery method based on cloud-edge collaboration as claimed in claim 4, characterized in that: The specific steps of calculating the final gain values of the first transmission path, the second transmission path, and the third transmission path based on the A* path planning algorithm are as follows: The first path weight R1, the second path weight R2 and the third path weight R3 of the first transmission path, the second transmission path and the third transmission path are obtained, and the calculation expression for calculating the final gain value is as follows: Final gain value = (R1 × transmission link safety value 2)+(R2×node data rate)+(R3×node noise error).
8. The distributed storage heterogeneous data recovery method based on cloud-edge collaboration as claimed in claim 2, characterized in that: Calculating the capacity of the fault data specifically includes: Counting the number of first redundant copies of the first node and other nodes in real time, and retrieving the preset number of second redundant copies of the first node and other nodes; When the number of the first redundant copies is less than the number of the second redundant copies, retrieving the first node data block size, the total data volume and the preset number of the second redundant copies; The capacity of the fault data = (the size of the data block of the first node × the total amount of data on the first node) / the number of the second redundant copies.
9. A distributed storage heterogeneous data recovery system based on cloud-edge collaboration, comprising a distributed storage heterogeneous data recovery method based on cloud-edge collaboration as described in any one of claims 1 to 8, characterized in that: It includes a node working status monitoring module, a node data calculation module, a transmission link planning module and a data integrity verification module; The node working status monitoring module is used to calculate the working status of the first node and the size of the lost data capacity; The node data calculation module is used to calculate the transmission link security value, node data rate and node noise error of the first node and other nodes; The transmission link planning module is used to select a transmission link with the smallest final gain value of the transmission links of the first node and other nodes as the optimal recovery path; The data integrity verification module is used to verify whether the restored data is consistent with the fault data lost by the first node.
Citation Information
Patent Citations
Data recovery method, system and equipment of distributed storage system and storage medium
CN117667494A
Network information security comprehensive analysis and monitoring system and method
CN118413359A
Channel Estimation for Optical Orthogonal Frequency Division Multiplexing Systems
US20140072307A1