A Distributed Data Security Storage Method, Device and System
By analyzing the differences and interactions of nodes in the distributed storage system and selecting appropriate backup nodes, the problems of data security and redundancy balance of distributed storage systems are solved, and more efficient data backup and storage are achieved.
Patent Information
- Application Number
- CN202510361222.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Existing distributed storage systems are difficult to achieve a good balance between ensuring data security and redundancy, resulting in load imbalance and redundancy in backup quantity.
By obtaining node information in the distributed data storage system, analyzing the functional differences and interaction between nodes, selecting the final backup node that meets the security requirements of the data to be backed up, ensuring that the data is stored evenly among multiple nodes.
It realizes that while ensuring the security of distributed data storage, redundant backups can be reduced as much as possible, and the backup effect can be improved, so as to achieve a good balance between the security and redundancy of distributed storage systems.
Smart Images

Figure CN119883139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a distributed data security storage method, device and system. Background Art
[0002] A distributed storage system stores data dispersedly on multiple independent devices. Traditional network storage systems use a centralized storage server to store all data. The storage server becomes the bottleneck of system performance and the focus of reliability and security, and cannot meet the needs of large-scale storage applications. The distributed network storage system adopts an extensible system structure, uses multiple storage servers to share the storage load, and uses a location server to locate storage information. It not only improves the reliability, availability and access efficiency of the system, but also is easy to expand.
[0003] There are also some security risks in the distributed storage system. Due to the characteristics of numerous nodes, unfixed structure, wide distribution, etc., it is difficult to effectively guarantee the security of the actual data in the memory. When a single node is damaged, data loss may occur. Therefore, it is often necessary to synchronously back up the data to other nodes during data storage, which can reduce the impact caused by a single node failure and also ensure the recoverability of the data and the security of the data.
[0004] In the prior art, the security of data in the distributed storage system is improved by combining node heterogeneity and the multi-copy mechanism. However, the existing methods only quantify the heterogeneity from the similarities and differences in functions, lacking the analysis of the data itself, resulting in load imbalance and redundancy in the number of backups, so that the security and redundancy of the distributed storage system cannot reach a good balance. Summary of the Invention
[0005] In order to solve the technical problem that the security and redundancy of the distributed storage system in the prior art cannot be well balanced, the purpose of the present invention is to provide a distributed data security storage method, device and system, and the specific technical solutions adopted are as follows:
[0006] In the first aspect, a distributed data security storage method is provided, and the method includes:
[0007] Step S1: Obtain node information in the distributed data storage system;
[0008] Step S2: Obtain the difference degree between any two nodes according to the functional difference degree and interaction situation between the nodes;
[0009] Step S3: Obtain the candidate backup nodes of the data to be backed up according to the security requirements of the data to be backed up and the difference degree aggregation situation between the node to which the data belongs and the selectable nodes;
[0010] Step S4: Screen out the final backup nodes from the candidate backup nodes according to the activity and storage conditions of the candidate backup nodes;
[0011] Step S5: Synchronously store the data to be backed up into all the final backup nodes to complete the distributed data secure storage.
[0012] Further, the obtaining of the node information in the distributed data storage system in the step S1 specifically includes:
[0013] For any distributed data storage system, regard any computer as a node, split any node into different components with independent functions, and obtain the component set of each node;
[0014] Obtain the vulnerability set of the software or hardware of any node.
[0015] Further, the step S2 specifically includes:
[0016] According to the number of components in the intersection and union of any two nodes, and the number of vulnerabilities in the intersection and union of these two nodes, obtain the functional difference degree of these two nodes;
[0017] According to the data volume transmitted between these two nodes, the number of intersection nodes and union nodes in the set of all nodes with data transmission between these two nodes, obtain the interaction situation of these two nodes;
[0018] According to the functional difference degree and interaction situation of these two nodes, obtain the difference degree of these two nodes.
[0019] Further, the data volume transmitted between these two nodes is negatively correlated with the difference degree of these two nodes, and the ratio of the number of intersection nodes to the number of union nodes is negatively correlated with the difference degree of these two nodes.
[0020] Further, the step S3 specifically includes:
[0021] According to the number of node classes, the number of steps for encrypting and verifying the data to be backed up, and the number of steps with the most encryption and verification in all the data, obtain the backup number of the data to be backed up;
[0022] According to the backup number of the data to be backed up, the average difference degree between the node to which the data to be backed up belongs and all the nodes in each node class, and the average difference degree between all the nodes of any two node classes in the candidate backup scheme, obtain the comprehensive difference degree between the data to be backed up and any candidate backup scheme;
[0023] Regard all the nodes of the candidate backup scheme with the largest comprehensive difference degree as the candidate backup nodes.
[0024] Further, the node class is specifically: clustering the undirected weighted mesh graph of nodes until all nodes are added to the clustering clusters or marked as noise points, and any noise point and any clustering cluster are recorded as a node class; the candidate backup scheme is composed of a combination of multiple node classes.
[0025] Further, step S4 is specifically: according to the data storage amount, the already stored amount, and the total storage space of the candidate backup nodes on the backup day and the previous day, obtain the appropriate backup degree of each candidate backup node, and record the candidate backup node with the largest appropriate backup degree as the final backup node.
[0026] Further, the growth rate of the data storage amount of the candidate backup nodes on the backup day compared to the previous day of backup, and the ratio of the already stored amount of the candidate backup nodes to the total storage space of the candidate backup nodes are both negatively correlated with the appropriate backup degree of the candidate backup nodes.
[0027] In a second aspect, the present invention provides a distributed data security storage device, which includes:
[0028] A node information acquisition module, configured to acquire node information in a distributed data storage system;
[0029] A difference degree acquisition module, configured to acquire the difference degree between any two nodes according to the functional difference degree and interaction situation between nodes;
[0030] A candidate backup node acquisition module, configured to acquire candidate backup nodes for data to be backed up according to the security requirements of the data to be backed up and the difference degree aggregation situation between the node to which the data belongs and the selectable nodes;
[0031] A final backup node acquisition module, configured to screen out the final backup nodes from the candidate backup nodes according to the activity and storage situation of the candidate backup nodes;
[0032] A storage execution module, configured to synchronously store the data to be backed up in all the final backup nodes.
[0033] In a third aspect, the present invention provides a distributed data security storage system, the system includes the above-mentioned distributed data security storage device, and further includes multiple nodes composed of multiple computers.
[0034] The present invention has the following beneficial effects: The present invention analyzes the differences among all nodes in a distributed data storage system to obtain an undirected weighted mesh graph of the nodes. Combining the characteristics that different data have different confidentiality requirements, it determines the backup quantity of any data, and further obtains backup nodes according to the storage ratio and activity of the data, so that each node in the storage system can store data evenly and complete multi-node backup of the data. The present invention is conducive to reducing redundant backups as much as possible on the premise of ensuring the security of distributed data storage, maximizing the backup effect, and achieving a good balance between the security and redundancy of the distributed storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a flowchart of a distributed data security storage method provided by an embodiment of the present invention.
[0037] Figure 2 It is an undirected weighted mesh graph of a distributed data storage system provided by an embodiment of the present invention.
[0038] Figure 3 It is a block diagram of a distributed data security storage device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects of a distributed data security storage method, device, and system proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0040] The following specifically describes the specific solutions of a distributed data security storage method, device, and system provided by the present invention with reference to the drawings.
[0041] In a first aspect, please refer to Figure 1 , which shows a flowchart of a distributed data security storage method provided by an embodiment of the present invention. The method includes the following steps:
[0042] Step S1: Obtain the node information in the distributed data storage system.
[0043] Among them, obtaining the node information in the distributed data storage system in step S1 specifically includes: for any distributed data storage system, taking any computer as a node, splitting any node into different components with independent functions, and obtaining the component set of each node; obtaining the set of vulnerabilities existing in the software or hardware of any node.
[0044] More specifically, for any distributed data storage system, taking any computer as a node, splitting any node into several components with independent functions, and different components with the same function are denoted as heterogeneous components. For example, the CPU type or the operating system type is a kind of function, X86 and Arm are two CPU types, that is, two components of one function, and they are heterogeneous components to each other; another example is that Windows10, Centos8, and RedHat7 are three operating system types, that is, three components of one function, and they are heterogeneous components to each other. The distributed data storage system can be denoted as a set of several components of several functions, denoted as , , , where represents the distributed data storage system, represents the i-th function, represents the j-th component of the i-th function, n represents the number of functions, and m represents the number of components corresponding to any function. Thus, any node is also a set of several components, and in this way, the component set of each node is obtained. In addition, obtain the set of vulnerabilities existing in the software or hardware of any node, denoted as the vulnerability set of any node, denoted as .
[0045] In the distributed data storage system, the functional compositions and interaction situations between different nodes are different. If there are two backup nodes with relatively high similarity, then the necessity of backing up both of these two backup nodes is relatively small. If both of these two backup nodes are backed up, the backup redundancy will be relatively high. Therefore, it is necessary to understand the differences between nodes. Accordingly, the following steps are set in this implementation.
[0046] Step S2: Obtain the difference degree between any two nodes according to the functional difference degree and interaction situation between nodes.
[0047] Among them, step S2 specifically includes: obtaining the functional difference degree between two nodes according to the number of components in the intersection and union of any two nodes and the number of vulnerabilities in the intersection and union of these two nodes; obtaining the interaction situation between the two nodes according to the amount of data transmitted between the two nodes, the number of intersection nodes and union nodes in the set of all nodes with data transmission between the two nodes; and obtaining the difference degree between the two nodes according to the functional difference degree and interaction situation between the two nodes.
[0048] More specifically, for the component sets of any two nodes, the difference between the two nodes is first manifested in the difference of components. When all the components of the functions of the nodes are heterogeneous components, the difference degree of the nodes is the largest. Among them, the intersection of the component sets of the two nodes is the same components of the two nodes, and the union is all the components that make up the two nodes. The ratio of the intersection to the union can reflect the proportion of the same components of the two nodes and reflect the similarity. At the same time, considering that the vulnerabilities of the nodes are the main reasons for attacks, if the vulnerabilities of the two nodes are the same, then the difference degree of the nodes is relatively small. Among them, the intersection of the vulnerability sets of the two nodes is the same vulnerabilities, and the union is all vulnerabilities. The ratio of the intersection and union of the vulnerability sets also reflects the similarity. The smaller the similarity of the component sets and the smaller the similarity of the vulnerability sets, the greater the difference degree between the nodes. This difference degree is mainly reflected as a functional difference. Therefore, the mathematical calculation formula for the functional difference degree of any two nodes constructed first in this embodiment is as follows:
[0049]
[0050] In the formula, and represent the number of intersection components and the number of union components of any two nodes. Obviously, each node has functions and components corresponding to the functions. Therefore, the component sets of any two nodes are not zero. Further, the number of union components of any two nodes ; and represent the number of intersection vulnerabilities and the number of union vulnerabilities of the two nodes. Obviously, each component is not 100% perfect and there are vulnerabilities. Therefore, the number of union vulnerabilities of the two nodes ; exp() represents the exponential function with the natural constant e as the base; represents the functional difference degree between the two arbitrary nodes.
[0051] In the mathematical formula for calculating the functional difference degree between any two nodes constructed above, the ratio of the number of intersection components to the number of union components of any two nodes can reflect the proportion of the same components of the two nodes and reflect the similarity between the two nodes; the ratio of the number of intersection vulnerabilities to the number of union vulnerabilities of the two nodes reflects the similarity in the functional vulnerabilities of the two components. The smaller the similarity of the component set and the smaller the similarity of the vulnerability set, the greater the functional difference degree between the nodes. That is, the ratio of the number of intersection components to the number of union components of the two nodes is negatively correlated with the functional difference degree between the two nodes, and the ratio of the number of intersection vulnerabilities to the number of union vulnerabilities of the two nodes is also negatively correlated with the functional difference degree between the two nodes.
[0052] Based on the functional difference, in actual operation, if there is frequent data transmission between two nodes, then there is a large amount of the same data between the two nodes. When a certain node is attacked, it is more likely to choose to attack the other node. Therefore, for data backup, this node should not be selected for backup. On the contrary, when the data transmission between the two nodes is less and the number of common nodes with data transmission is less, the difference degree between the two nodes is greater. Thus, based on the interaction situation of the nodes, the mathematical formula for calculating the difference degree between any two nodes constructed in this embodiment is as follows:
[0053]
[0054] In the formula, represents the difference degree between any two nodes; represents the functional difference degree between the two nodes; represents the amount of data transmitted between the two nodes; 、 represent the number of intersection nodes and the number of union nodes of all the node sets with data transmission between the two nodes. Among them, since each node in the distributed system has a data transmission situation, obviously, the number of union nodes of all the node sets with data transmission between the two nodes ; exp() represents the exponential function with the natural constant e as the base.
[0055] In the mathematical formula for calculating the difference degree between any two nodes constructed above, the functional difference degree between the two nodes is obviously positively correlated with the difference degree between any two nodes. The amount of data transmitted between the two nodes is negatively correlated with the difference degree between the two nodes. The ratio of the number of intersection nodes to the number of union nodes of all the node sets with data transmission between the two nodes is negatively correlated with the difference degree between the two nodes. When the data transmission between the two nodes is less and the number of common nodes with data transmission is less, the difference degree between the two nodes is greater.
[0056] After obtaining the difference degrees between any two nodes and regarding the difference degrees as distance weights, all nodes of the distributed storage system can be drawn into an undirected weighted mesh graph. When selecting backup nodes for the data to be backed up, in order to ensure security while minimizing the number of backups, the backup nodes should have the largest possible difference degree from the node to which the data to be backed up belongs, and at the same time, the difference degrees between the backup nodes should also be ensured to be as large as possible. Due to the large number of nodes and the existence of some similar nodes, the undirected weighted mesh graph can be clustered according to the difference degrees first, and the number of backups can be selected with reference to the clustering situation. On the other hand, the security requirements of different data are different. Referring to the encryption and verification steps of the data to be backed up can further determine the number of backups, and then the candidate backup nodes can be preliminarily determined according to the clustering situation. Therefore, the present embodiment further sets the following steps.
[0057] Step S3: Obtain the candidate backup nodes for the data to be backed up according to the security requirements of the data to be backed up, the clustering situation of the difference degrees between the node to which it belongs and the selectable nodes.
[0058] Among them, step S3 specifically includes: obtaining the number of backups of the data to be backed up according to the number of node classes, the number of steps of encrypting and verifying the data to be backed up, and the number of steps of encrypting and verifying the most among all data; obtaining the comprehensive difference degree between the data to be backed up and any candidate backup scheme according to the number of backups of the data to be backed up, the average difference degree between the node to which the data to be backed up belongs and all nodes in each node class, and the average difference degree between all nodes of any two node classes of the candidate backup scheme; recording all nodes of the candidate backup scheme with the largest comprehensive difference degree as the candidate backup nodes. It should be noted that the node class specifically is: clustering the undirected weighted mesh graph of nodes until all nodes are added to the clustering clusters or marked as noise points, and any noise point and any clustering cluster are recorded as a node class; in addition, the candidate backup scheme is composed of a combination of multiple node classes.
[0059] More specifically, first, all nodes of the distributed storage system can be drawn into an undirected weighted mesh graph, as shown in the appendix Figure 2 shown, in the figure , respectively represent the weights between nodes , and the weights between nodes , . Cluster the undirected weighted mesh graph of nodes, use the DBSCAN algorithm, take the first node as the starting point, and set the minimum neighbor distance to (Half of the average of the difference degrees between all nodes), set the minimum number of clustering neighboring points to 2; starting from the starting point, find all its neighboring points within the minimum neighboring distance range. If the number is greater than or equal to the minimum number of clustering neighboring points, add these points to the current clustering cluster, mark these points as visited, and then repeatedly find their neighboring points from the visited points and add them to the clustering until no more points can be added; continue to select the starting point from the unvisited points and repeat the operation until all points are added to the clustering cluster or are marked as noise points because they do not meet the addition conditions. Denote any noise point and any clustering cluster as a node class, and obtain several node classes of the weighted mesh graph of the nodes.
[0060] Secondly, since each node class consists of nodes with similar difference degrees, when selecting backup nodes for the data to be backed up, one node can be selected from each node class. Then, the necessity of the backup quantity needs to be considered, that is, from how many node classes the data to be backed up needs to select nodes for backup. The more steps the data goes through for encryption and verification, the greater the security requirement for the data to be backed up, and the more node classes should be selected to choose nodes. In this embodiment, the mathematical calculation formula for the backup quantity of the data to be backed up is constructed as follows:
[0061]
[0062] In the formula, represents the backup quantity of the data to be backed up; represents the number of node classes; represents the maximum number of steps for encryption and verification among all data; represents the number of steps for encryption and verification of the data to be backed up.
[0063] In the above - constructed mathematical calculation formula for the backup quantity of the data to be backed up, the number of node classes serves as the base value for the backup quantity of the data to be backed up. Each node class consists of nodes with similar difference degrees. When selecting backup nodes for the data to be backed up, generally one node is selected from each node class. The greater the difference between the number of steps for encryption and verification of the data to be backed up and the maximum number of steps for encryption and verification among all data, the greater the security requirement for the data to be backed up, the higher the necessity of considering backup, and then it is necessary to increase on the basis of the backup quantity being the number of node classes ; the smaller the difference, the smaller the security requirement for the data to be backed up, the lower the necessity of considering backup, and then it is necessary to decrease on the basis of the backup quantity being the number of node classes 。
[0064] After obtaining the backup quantity The node class with the largest difference degree from the data to be backed up is the candidate backup node. At the same time, the difference degree between node classes also needs to be considered. All node classes are combined into several candidate backup schemes composed of N node classes, and the comprehensive difference degree between the data to be backed up and any candidate backup scheme is calculated. In this embodiment, the mathematical calculation formula for the comprehensive difference degree between the data to be backed up and any candidate backup scheme is as follows:
[0065]
[0066] In the formula, represents the comprehensive difference degree between the data to be backed up and any candidate backup scheme; represents the number of node classes in the candidate backup scheme; represents the average difference degree between the node to which the data to be backed up belongs and all nodes in the h-th node class; represents the average difference degree between all nodes of any two node classes in the candidate backup scheme.
[0067] In the above mathematical calculation formula for the comprehensive difference degree between the data to be backed up and any candidate backup scheme, represents the comprehensive difference degree between all node classes of the data to be backed up and the candidate backup scheme, represents the comprehensive difference degree between all node classes of the candidate backup scheme. The two together reflect the comprehensive difference degree of the candidate backup scheme.
[0068] In this way, the comprehensive difference degree between the data to be backed up and any candidate backup scheme is obtained. The N node classes of the candidate backup scheme with the largest comprehensive difference degree are used as the backup scheme for the data to be backed up, and all nodes are used as candidate backup nodes.
[0069] Since there are differences in data storage capacity and storage growth rate among candidate backup nodes, not all of them are suitable as backup nodes for the data to be backed up. For example, a certain node has been stored with a large amount of data by users recently and may also be used for data storage in the future. At the same time, the remaining storage capacity of the node itself is relatively small. Then, in order to make the storage of all nodes more balanced, it should be avoided for backup. Therefore, among the candidate backup nodes, the appropriate backup degree of all nodes should be calculated, and the nodes with a greater appropriate backup degree are used as backup nodes, ensuring the load balance of the storage system. Therefore, this embodiment further sets the following steps:
[0070] Step S4: Screen out the final backup nodes from the candidate backup nodes according to the activity and storage conditions of the candidate backup nodes.
[0071] Among them, step S4 is specifically as follows: According to the data storage amounts of the candidate backup nodes on the current day and the previous day of backup, the stored amounts of the candidate backup nodes, and the total storage spaces of the candidate backup nodes, obtain the appropriate backup degrees of each candidate backup node, and record the candidate backup node with the highest appropriate backup degree as the final backup node. Since the slower the data growth rate of a node and the smaller the proportion of the stored amount, the higher the appropriate backup degree of the node. Therefore, the mathematical calculation formula for the appropriate backup degree of any candidate backup node constructed in this embodiment can be expressed as:
[0072]
[0073] In the formula, represents the appropriate backup degree of any candidate backup node; and represent the data storage amounts of the candidate backup nodes on the current day and the previous day of backup; represents the stored amount of the candidate backup node; represents the total storage space of the candidate backup node. Since each candidate backup node has a total storage space, obviously, ;
[0074] In the above-mentioned mathematical calculation formula for the appropriate backup degree of any candidate backup node constructed, represents the ratio between the data storage amounts of the candidate backup nodes on the current day and the previous day of backup, represents the growth rate of the data storage amounts of the candidate backup nodes on the current day and the previous day of backup, which is negatively correlated with the appropriate backup degree of any candidate backup node. The slower the growth rate of the data storage amounts of the candidate backup nodes on the current day and the previous day of backup, the higher the appropriate backup degree of any candidate backup node; represents the proportion of the stored amount of the node, which is negatively correlated with the appropriate backup degree of any candidate backup node. The smaller the proportion of the stored amount, the higher the appropriate backup degree of any candidate backup node. That is, both the growth rate of the data storage amount of the candidate backup node on the current day compared to the previous day of backup and the ratio of the stored amount of the candidate backup node to the total storage space of the candidate backup node are negatively correlated with the appropriate backup degree of the candidate backup node.
[0075] Furthermore, obtain the node with the highest appropriate backup degree in each node class in the backup plan, that is, the final backup node.
[0076] Step S5: Synchronously store the data to be backed up into all the final backup nodes to complete the distributed data secure storage.
[0077] Specifically, the data to be backed up is synchronously stored in all the final backup nodes to complete the data backup. It should be noted that the weighted mesh graph of the storage system nodes is updated in real time to ensure the classification effect of the node classes, so as to ensure the best selection of data backup nodes and balance the data security and redundancy.
[0078] Thus, the secure storage of the data in the arbitrary distributed data storage system is completed. In this way, the redundant backup can be reduced as much as possible, and the backup effect can be maximally improved, so that the security and redundancy of the distributed storage system reach a good balance.
[0079] In a second aspect, please refer to Figure 3 , which shows a distributed data secure storage device provided by an embodiment of the present invention. The device includes:
[0080] A node information acquisition module 101, configured to acquire node information in a distributed data storage system;
[0081] A difference degree acquisition module 102, configured to acquire the difference degree between any two nodes according to the functional difference degree and interaction situation between the nodes;
[0082] A candidate backup node acquisition module 103, configured to acquire candidate backup nodes for the data to be backed up according to the security requirements of the data to be backed up and the difference degree aggregation situation between the node to which the data belongs and the selectable nodes;
[0083] A final backup node acquisition module 104, configured to screen out final backup nodes from the candidate backup nodes according to the activity and storage situation of the candidate backup nodes;
[0084] A storage execution module 105, configured to synchronously store the data to be backed up in all the final backup nodes.
[0085] Further, the device further includes:
[0086] A functional difference degree acquisition module, configured to acquire the functional difference degree between the two nodes according to the number of components in the intersection and union of any two nodes and the number of vulnerabilities in the intersection and union of the two nodes;
[0087] An interaction situation acquisition module, configured to acquire the interaction situation between the two nodes according to the amount of data transmitted between the two nodes, the number of intersection nodes and union nodes in the set of all nodes where data transmission exists between the two nodes;
[0088] A backup quantity acquisition module, configured to acquire the backup quantity of the data to be backed up according to the number of nodes, the number of steps for encrypting and verifying the data to be backed up, and the number of steps with the most encryption and verification in all the data;
[0089] The comprehensive difference degree acquisition module is used to obtain the comprehensive difference degree between the data to be backed up and any candidate backup scheme according to the backup quantity of the data to be backed up, the average difference degree between the node to which the data to be backed up belongs and all nodes of each node class, and the average difference degree between all nodes of any two node classes of the candidate backup scheme.
[0090] In a third aspect, this embodiment provides a distributed data security storage system, which includes any one of the above-mentioned distributed data security storage devices, and also includes a plurality of nodes composed of a plurality of computers.
[0091] A distributed data security storage method, device and system provided in this embodiment first obtains node information in the distributed data storage system; then, according to the functional difference degree and interaction situation between nodes, obtains the difference degree between any two nodes; according to the security requirements of the data to be backed up, the difference degree aggregation situation between the node to which the data belongs and the selectable nodes, obtains the candidate backup nodes of the data to be backed up; according to the activity and storage situation of the candidate backup nodes, screens out the final backup nodes from the candidate backup nodes; finally, synchronously stores the data to be backed up into all the final backup nodes to complete the distributed data security storage. In this way, each node of the distributed data security system can store evenly and complete the multi-node backup of the data, which is beneficial to reducing redundant backup as much as possible on the premise of ensuring the security of distributed data storage.
[0092] It should be noted that the above-mentioned sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0093] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A distributed data security storage method, characterized in that: The method comprises: Step S1: Obtain node information in a distributed data storage system; Step S2: Obtain the difference between any two nodes according to the functional difference and interaction between the nodes; Step S3: obtaining a candidate backup node for the data to be backed up according to the security requirements of the data to be backed up and the aggregation of differences between the node to which the data belongs and the selectable nodes; Step S4: Select the final backup node from the candidate backup nodes according to the activity and storage conditions of the candidate backup nodes; Step S5: Synchronously store the data to be backed up in all final backup nodes to complete the distributed data security storage; The step S1 of acquiring the node information in the distributed data storage system specifically includes: For any distributed data storage system, any computer is regarded as a node, any node is split into different components with independent functions, and the component set of each node is obtained; Get the set of vulnerabilities of the software or hardware of any node; The step S2 specifically includes: According to the number of components of the intersection and union of any two nodes and the number of vulnerabilities of the intersection and union of the two nodes, the functional difference of the two nodes is obtained; Obtain the interaction status of the two nodes according to the amount of data transmitted between the two nodes, the number of intersection nodes and the number of union nodes of all node sets with data transmission between the two nodes; According to the functional difference and interaction status of the two nodes, the difference between the two nodes is obtained; The mathematical calculation formula for the functional difference between any two nodes is as follows: In the formula, , Indicates the number of intersection components and the number of union components of any two nodes; , Indicates the number of intersection vulnerabilities and the number of union vulnerabilities of the two nodes; exp() represents an exponential function with the natural constant e as the base; Indicates the functional difference between any two nodes; The mathematical formula for calculating the difference between any two nodes is as follows: In the formula, Represents the difference between any two nodes; Indicates the amount of data transmitted between the two nodes; , Indicates the number of intersection nodes and union nodes of all node sets with which data is transmitted between the two nodes.
2. A distributed data security storage method according to claim 1, characterized in that: The amount of data transmitted between the two nodes is negatively correlated with the difference between the two nodes, and the ratio of the number of intersection nodes to the number of union nodes is negatively correlated with the difference between the two nodes.
3. A distributed data security storage method according to claim 1, characterized in that: The step S3 specifically includes: Obtain the number of backups of the data to be backed up according to the number of node classes, the number of encryption and verification steps of the data to be backed up, and the number of the most encryption and verification steps of all data; Obtain the comprehensive difference between the data to be backed up and any backup scheme to be selected based on the number of backups of the data to be backed up, the average difference between the node to which the data to be backed up belongs and all nodes in each node class, and the average difference between all nodes of any two node classes of the backup scheme to be selected; All nodes of the candidate backup solution with the largest comprehensive difference are recorded as candidate backup nodes.
4. A distributed data security storage method according to claim 3, characterized in that: The node class is specifically: clustering the undirected weighted network graph of nodes until all nodes join the cluster or are marked as noise points, and any noise point and any cluster are recorded as a node class; the backup plan to be selected is composed of a combination of multiple node classes.
5. A distributed data security storage method according to claim 1, characterized in that: The step S4 is specifically as follows: according to the data storage capacity of the backup node on the backup day and the previous day, the storage capacity of the backup node, and the total storage space of the backup node, the suitable backup degree of each backup node is obtained, and the backup node with the largest suitable backup degree is recorded as the final backup node.
6. A distributed data secure storage method according to claim 5, characterized in that: The growth rate of the data storage capacity of the candidate backup node on the backup day compared with the day before the backup, and the ratio of the storage capacity of the candidate backup node to the total storage space of the candidate backup node are negatively correlated with the suitability of the candidate backup node for backup.
7. A distributed data security storage device, characterized in that: The device includes: A node information acquisition module is used to acquire node information in a distributed data storage system; The difference acquisition module is used to obtain the difference between any two nodes according to the functional difference and interaction between the nodes; A candidate backup node acquisition module is used to acquire the candidate backup node of the data to be backed up according to the security requirements of the data to be backed up and the aggregation of differences between the node to which it belongs and the selectable nodes; The final backup node acquisition module is used to select the final backup node from the candidate backup nodes according to the activity and storage status of the candidate backup nodes; A storage execution module is used to synchronously store the data to be backed up in all final backup nodes; Obtaining node information in a distributed data storage system specifically includes: For any distributed data storage system, any computer is regarded as a node, any node is split into different components with independent functions, and the component set of each node is obtained; Get the set of vulnerabilities of the software or hardware of any node; Getting the difference between any two nodes specifically includes: According to the number of components of the intersection and union of any two nodes and the number of vulnerabilities of the intersection and union of the two nodes, the functional difference of the two nodes is obtained; Obtain the interaction status of the two nodes according to the amount of data transmitted between the two nodes, the number of intersection nodes and the number of union nodes of all node sets with data transmission between the two nodes; According to the functional difference and interaction status of the two nodes, the difference between the two nodes is obtained; The mathematical calculation formula for the functional difference between any two nodes is as follows: In the formula, , Indicates the number of intersection components and the number of union components of any two nodes; , Indicates the number of intersection vulnerabilities and the number of union vulnerabilities of the two nodes; exp() represents an exponential function with the natural constant e as the base; Indicates the functional difference between any two nodes; The mathematical formula for calculating the difference between any two nodes is as follows: In the formula, Represents the difference between any two nodes; Indicates the amount of data transmitted between the two nodes; , Indicates the number of intersection nodes and union nodes of all node sets with which data is transmitted between the two nodes.
8. A distributed data security storage system, characterized in that: The system includes a distributed data security storage device as described in claim 7, and also includes multiple nodes composed of multiple computers.
Citation Information
Patent Citations
Self-adaptive feedback resource scheduling method for improving cloud reliability
CN111580950A
Load weighing method based on systematic grade diagnosis information
CN1512380A