A Method for Cluster Adaptive Data Backup and Recovery Based on CopySet
By calculating the correlation degree and scatter width functions of disk nodes, the number of CopySets is determined, data replication and fault repair are performed, and the problem of inaccurate determination of the number of replication sets is solved, which reduces the probability of data loss and improves recovery efficiency.
Patent Information
- Application Number
- CN202211741445.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In the prior art, the number of replica sets is determined inaccurately, resulting in insufficient number of replicas or slow data recovery, low recovery efficiency of replica set backup storage unit, and the probability of data loss cannot be effectively reduced.
By obtaining the configuration information of the distributed database, the correlation degree and scatter width functions between disk nodes are calculated, the number of CopySets is determined based on the bandwidth and recovery rate, the data is copied and stored in the replica node, and the polling module is used to detect the faulty node and repair it to generate early warning information.
The correlation between replication sets and cluster size is optimized, the probability of data loss is reduced, data recovery efficiency is improved, and operation and maintenance costs are simplified.
Smart Images

Figure CN115809166B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method for cluster adaptive data backup and recovery based on CopySet. Background Art
[0002] With the advent of the cloud computing and big data era, the data volume has increased sharply, and the importance of storage systems has become more prominent. The Internet plays a great role in social management, and a large amount of data information will be generated in social operations. Sufficient data information is conducive to the most objective and accurate analysis by various government agencies. Therefore, data backup technologies and models are particularly important.
[0003] Copyset Replication is a new general replication technology that can significantly reduce the frequency of data loss events. When conducting experiments on the facebook HDFS cluster, after changing from the original RandomReplicaiton to the Copyset Replication algorithm, the probability of data loss when 1% of the nodes in the FaceBook HDFS cluster fail can be reduced from 22.8% to 0.78%, effectively reducing the probability of data loss.
[0004] However, currently, it has been rigorously mathematically proven that the number of copy sets determines the probability that multiple copies of the same data are simultaneously hit. Therefore, the number of copy sets should be as small as possible. However, too few copy sets will result in insufficient number of replicas and slow data recovery. Therefore, the current number of corresponding copy sets cannot be determined, and in terms of data recovery, the efficiency of data recovery through the backup storage unit of the copy set is not high. Summary of the Invention
[0005] In view of the problems existing in the prior art, an embodiment of the present invention provides a method and device for cluster adaptive data backup and recovery based on CopySet.
[0006] An embodiment of the present invention provides a method for cluster adaptive data backup and recovery based on CopySet, including:
[0007] Obtaining configuration information of a distributed database, where the distributed database is respectively stored on different data clusters, the data clusters include corresponding disk nodes, and the configuration information includes: the degree of association between disk nodes and the size of the data cluster;
[0008] Determine the association degree function between disk nodes based on the degree of association, calculate the scattering width function of the data cluster based on the degree of association and the size of the data cluster, and determine the CopySet function by combining the association degree function and the scattering width function;
[0009] Obtain the bandwidth and data recovery rate of the disk node, determine the corresponding data recovery function according to the bandwidth and data recovery rate, and solve based on the CopySet function and the data recovery function to obtain the number of CopySets of the CopySet function;
[0010] Perform data replication on the data to be stored based on the number of CopySets to obtain the corresponding replicated data, store the replicated data on the corresponding N replica nodes, generate the corresponding N-tuple and store it in the central table of the distributed database, and store the N-tuple in the unit tables of the N replica nodes;
[0011] Poll the disk nodes through the polling module. When a disk node failure is detected, query the unit table of the faulty disk node, determine the associated repair replica node through the query result, and perform data repair on the faulty disk node through the repair replica node.
[0012] In one embodiment, the method further includes:
[0013] When the query operation of querying the unit table of the faulty disk node fails, obtain the N-tuple of the faulty disk node through the central table, generate a warning message through the N-tuple, and send the warning message to the bound terminal.
[0014] In one embodiment, the method further includes:
[0015] Determine the corresponding minimum value function of the CopySet function based on the CopySet function, determine the corresponding maximum value function of the data recovery function based on the data recovery function, and perform fitting through the system of equations of the minimum value function of the CopySet function and the maximum value function of the data recovery function to obtain the number of CopySets of the CopySet function.
[0016] In one embodiment, the system of equations includes:
[0017] max(v)=max(fun(b,m))
[0018] min(m)=min(fun(S,r))
[0019] Where v is the data recovery speed, b is the bandwidth, m is the number of CopySets, S is the scattering width, and r is the association degree function.
[0020] In one embodiment, the method further includes:
[0021] Generating a log file of the disk node with a fault and storing the log file in the central table.
[0022] An embodiment of the present invention provides an apparatus for cluster adaptive data backup and recovery based on CopySet, including:
[0023] A first acquisition module, configured to acquire configuration information of a distributed database, where the distributed database is respectively stored on different data clusters, the data clusters include corresponding disk nodes, and the configuration information includes: the degree of association between disk nodes and the size of the data cluster;
[0024] A calculation module, configured to determine an association degree function between disk nodes based on the degree of association, calculate a scattering width function of the data cluster based on the degree of association and the size of the data cluster, and determine a CopySet function by combining the association degree function and the scattering width function;
[0025] A second acquisition module, configured to acquire the bandwidth and data recovery rate of the disk node, determine a corresponding data recovery function according to the bandwidth and the data recovery rate, and solve the number of CopySets of the CopySet function based on the CopySet function and the data recovery function;
[0026] A replication module, configured to perform data replication on the data to be stored based on the number of CopySets to obtain corresponding replicated data, store the replicated data on corresponding N replica nodes, generate corresponding N-tuples and store them in the central table of the distributed database, and store the N-tuples in the unit tables of the N replica nodes;
[0027] A polling module, configured to poll the disk nodes through the polling module. When a disk node with a fault is detected, query the unit table of the disk node with a fault, determine associated repair replica nodes through the query result, and perform data repair on the disk node with a fault through the repair replica nodes.
[0028] In one embodiment, the apparatus further includes:
[0029] An early warning module, configured to, when the query operation of querying the unit table of the disk node with a fault fails, obtain the N-tuples of the disk node with a fault through the central table, generate an early warning information through the N-tuples, and send the early warning information to a bound terminal.
[0030] In one embodiment, the apparatus further includes:
[0031] A fitting module, configured to determine a minimum function of the corresponding CopySet function based on the CopySet function, determine a maximum function of the corresponding data recovery function based on the data recovery function, and perform fitting through a system of equations of the minimum function of the CopySet function and the maximum function of the data recovery function, so as to obtain the number of CopySets of the CopySet function.
[0032] An embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method for cluster adaptive data backup and recovery based on CopySet are implemented.
[0033] An embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method for cluster adaptive data backup and recovery based on CopySet are implemented.
[0034] The embodiment of the present invention provides a cluster adaptive data backup and recovery method and device based on CopySet, which obtains configuration information of a distributed database, wherein the distributed database is stored in different data clusters respectively, and the data cluster includes corresponding disk nodes, and the configuration information includes: the degree of association between the disk nodes and the size of the data cluster; the association function between the disk nodes is determined based on the association degree, the scattering width function of the data cluster is calculated based on the association degree and the size of the data cluster, and the CopySet function is determined in combination with the association function and the scattering width function; the bandwidth and data recovery rate of the disk node are obtained, and the corresponding data recovery function is determined according to the bandwidth and the data recovery rate. The number of CopySets is obtained based on the CopySet function and the data recovery function. Based on the number of CopySets, the data to be stored is copied to obtain the corresponding replica data, and the replica data is stored on the corresponding N replica nodes. The corresponding N tuples are generated and stored in the central table of the distributed database, and the N tuples are stored in the unit table of the N replica nodes. The disk nodes are polled by the polling module. When a disk node failure is detected, the unit table of the failed disk node is queried, and the associated repair replica node is determined by the query result. The data of the failed disk node is repaired by the repair replica node. In this way, an adaptive selection scheme of replica sets based on cluster scale can be designed, which can greatly enhance the correlation between replica sets and cluster scale, and increase parameters such as cluster scale and correlation. The advantage is that the correlation in the distribution of cluster storage units is taken into account, and there may be a joint failure of storage units. Through this parameter, the distribution strategy with the smallest replica set can be selected, thereby effectively reducing the probability of data loss. In addition, when data needs to be recovered, more disks will participate in data recovery, thereby effectively improving data recovery efficiency. In addition, once a failure occurs, the operation and maintenance personnel can query the table to obtain the storage units where the backup data associated with the damaged / lost data is located, and check these storage units in time to reduce the risk of data loss, greatly simplifying the operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0036] Figure 1 It is a flow chart of a cluster adaptive data backup and recovery method based on CopySet in an embodiment of the present invention;
[0037] Figure 2Schematic diagram of a scattering width in an embodiment of the present invention;
[0038] Figure 3 Architecture diagram of a method for cluster adaptive data backup and recovery based on CopySet in an embodiment of the present invention;
[0039] Figure 4 Structure diagram of a device for cluster adaptive data backup and recovery based on CopySet in an embodiment of the present invention;
[0040] Figure 5 Schematic diagram of the structure of an electronic device in an embodiment of the present invention. Detailed implementation manners
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] Figure 1 Flow schematic diagram of a method for cluster adaptive data backup and recovery based on CopySet provided by an embodiment of the present invention. As Figure 1 shown, an embodiment of the present invention provides a method for cluster adaptive data backup and recovery based on CopySet, including:
[0043] Step S101, obtain configuration information of a distributed database, where the distributed database is respectively stored on different data clusters, the data clusters include corresponding disk nodes, and the configuration information includes: the degree of association between disk nodes, and the size of the data clusters.
[0044] Specifically, the data in the distributed storage system is divided into data clusters for cluster storage. Cluster storage takes each storage device as a storage node, connects them through a high-speed interconnection network, and disperses and stores the data on multiple independent devices according to the size of the disk nodes or manual settings. These devices can operate independently and cooperate with each other. The data cluster includes each disk node for storing corresponding data. Among them, the disk node is the unit node for storing data. In addition, relevant configuration information in the distributed database is obtained, such as: the degree of association between disk nodes, the size of the data cluster, such as series and distributed clusters, and the degree of association between disks is different, as well as the distribution of the cluster for storage units. For example, the correlation between disk nodes with the same power supply and different power supplies is also inconsistent. According to the relevant configuration information of the distributed database, the degree of association between disk nodes is determined.
[0045] Step S102, determine the association degree function between disk nodes based on the degree of association, calculate the scattering width function of the data cluster based on the degree of association and the size of the data cluster, and determine the CopySet function by combining the association degree function and the scattering width function.
[0046] Specifically, determine the association degree function between disk nodes based on the degree of association. Among them, in this embodiment, the association degree function can be expressed as r = fun(x1, x2,..., xn), where xi (i = 1, 2,..., n) represents various correlation influencing factors, such as the cluster distribution type, power supply type, etc. In addition, when performing Copyset Replication (CopySet Replication), the scattering width is involved, that is, the maximum interval between the original data and its copy, which limits the distribution range between two data. In this embodiment, the scattering width function of the data cluster is calculated based on the degree of association and the size of the data cluster. The function corresponding to the scattering width is S = fun(n, r). The schematic diagram of the scattering width is as Figure 2 shown. Secondly, the number of CopySets is a function based on the above two (association degree function and scattering width function). If m is used to represent the number of CopySets, then the CopySet function is m = fun(S, r), and min(m) is solved, that is, min(m) = min(fun(S, r)).
[0047] Step S103, obtain the bandwidth and data recovery rate of the disk node, determine the corresponding data recovery function according to the bandwidth and data recovery rate, and fit the number of CopySets of the CopySet function based on the CopySet function and the data recovery function.
[0048] Specifically, when determining the number of CopySets, it is not feasible to only determine the number of CopySets through the CopySet function, which will affect the efficiency of data recovery. Therefore, the data recovery ability needs to be considered simultaneously. Then, the bandwidth and data recovery rate of the disk node are determined according to various indicators of the disk node. That is, the data recovery speed depends on the bandwidth b and the number m of Copysets participating in the data recovery work. Here, let the data recovery speed v = fun(b, m). Then, for the number of CopySets, it is necessary to satisfy that v is as large as possible when m is the smallest. Then, the following equations need to be solved:
[0049] max(v)=max(fun(b,m))
[0050] min(m)=min(fun(S,r))
[0051] The above equations are optimization problems and can be solved by fitting or matrix theory methods to obtain the optimal m value.
[0052] Step S104: Based on the number of CopySets, perform data replication on the data to be stored to obtain corresponding replicated data. Store the replicated data on corresponding N replica nodes, generate corresponding N-tuples and store them in the central table of the distributed database, and store the N-tuples in the unit tables of the N replica nodes.
[0053] Specifically, after determining the number of CopySets, perform data replication on the data to be stored based on the number of CopySets to obtain corresponding replicated data. When storing the replicated data on corresponding disk nodes, design a corresponding Copyset storage table. The Copyset storage table is divided into a central table and a unit table, including: each time when storing, that is, storing the replicated data on corresponding N replica nodes, generate a corresponding N-tuple and store it in the central table of the distributed database, and then store the N-tuples in the unit tables of the N replica nodes. Taking the case where the number of replica nodes is 3 as an example, each data storage will generate a triple and store it in the central table. The triple respectively represents three storage units with different numbers, and then store the triple in the unit tables of the three corresponding storage units. By adopting this central-local dual-replica mechanism, not only can the risk of data loss or problems in the storage table be greatly reduced, but also the retrieval events of storage units in the data recovery process can be reduced, improving data recovery.
[0054] Step S105: Poll the disk nodes through a polling module. When it is detected that a disk node fails, query the unit table of the failed disk node, determine the associated repair replica node through the query result, and perform data repair on the failed disk node through the repair replica node.
[0055] Specifically, the polling module periodically polls the disk nodes in the distributed storage system for faults. When a disk node is detected to have a fault, data loss may occur in the disk node. Then, the unit table of the faulty disk node is queried, and the associated repair replica nodes are determined based on the query results. At the same time, it is also possible for relevant operation and maintenance personnel to repair the data of the faulty disk node based on the associated repair replica nodes. In this embodiment, the architecture diagram of the method for cluster adaptive data backup and recovery based on CopySet can be as shown in Figure 3 shown.
[0056] In addition, as shown in the architecture diagram of Figure 3 , when the query operation for the unit table of the faulty disk node fails, the N-tuple of the faulty disk node is also obtained through the central table, a warning message is generated based on the N-tuple, and the warning message is sent to the bound terminal for the operation and maintenance personnel of the bound terminal to perform repairs according to the N-tuple in the central table.
[0057] In addition, after repairing the data of the faulty disk node, a log file of the faulty disk node can be generated and stored in the central table for the operation and maintenance personnel to perform statistics on the faulty disk nodes, facilitating the operation and maintenance personnel to adjust the disk nodes prone to faults.
[0058] The embodiment of the present invention provides a cluster adaptive data backup and recovery method based on CopySet, which obtains configuration information of a distributed database, wherein the distributed database is stored in different data clusters respectively, and the data cluster includes corresponding disk nodes, and the configuration information includes: the degree of association between the disk nodes and the size of the data cluster; the association function between the disk nodes is determined based on the association degree, the scattering width function of the data cluster is calculated based on the association degree and the size of the data cluster, and the CopySet function is determined in combination with the association function and the scattering width function; the bandwidth and data recovery rate of the disk node are obtained, and the corresponding data recovery function is determined according to the bandwidth and the data recovery rate. , and based on the CopySet function and the data recovery function, the CopySet number of the CopySet function is obtained; based on the CopySet number, the data to be stored is copied to obtain the corresponding copy data, the copy data is stored on the corresponding N copy nodes, the corresponding N tuples are generated and stored in the central table of the distributed database, and the N tuples are stored in the unit table of the N copy nodes; the disk nodes are polled through the polling module, and when a disk node failure is detected, the unit table of the failed disk node is queried, and the associated repair copy node is determined through the query result, and the data of the failed disk node is repaired through the repair copy node. In this way, a replica set adaptive selection scheme based on cluster scale can be designed, which can greatly enhance the correlation between the replica set and the cluster scale, and increase parameters such as cluster scale and correlation. The advantage is that the correlation in the distribution of cluster storage units is taken into account, and there may be a joint failure of storage units. Through this parameter, the distribution strategy with the smallest replica set can be selected, thereby effectively reducing the probability of data loss. In addition, when data needs to be recovered, more disks will participate in data recovery, thereby effectively improving data recovery efficiency. In addition, once a failure occurs, the operation and maintenance personnel can query the table to obtain the storage units where the backup data associated with the damaged / lost data is located, and check these storage units in time to reduce the risk of data loss, greatly simplifying the operation and maintenance costs.
[0059] Figure 4 A cluster adaptive data backup and recovery device based on CopySet provided by an embodiment of the present invention includes: a first acquisition module S201, a calculation module S202, a second acquisition module S203, a replication module S204, and a polling module S205, wherein:
[0060] The first acquisition module S201 is used to acquire configuration information of a distributed database, wherein the distributed database is stored in different data clusters, wherein the data clusters include corresponding disk nodes, and the configuration information includes: the degree of association between the disk nodes and the size of the data clusters.
[0061] A calculation module S202, configured to determine an association degree function between disk nodes based on the association degree, calculate a scattering width function of the data cluster based on the association degree and the size of the data cluster, and determine a CopySet function by combining the association degree function and the scattering width function.
[0062] A second acquisition module S203, configured to acquire the bandwidth and data recovery rate of the disk node, determine a corresponding data recovery function according to the bandwidth and the data recovery rate, and solve based on the CopySet function and the data recovery function to obtain the number of CopySets of the CopySet function.
[0063] A replication module S204, configured to perform data replication on the data to be stored based on the number of CopySets to obtain corresponding replicated data, store the replicated data on corresponding N replica nodes, generate corresponding N-tuples and store them in the central table of the distributed database, and store the N-tuples in the unit tables of the N replica nodes.
[0064] A polling module S205, configured to poll the disk nodes through the polling module. When it detects that a disk node fails, query the unit table of the faulty disk node, determine the associated repair replica node through the query result, and perform data repair on the faulty disk node through the repair replica node.
[0065] In one embodiment, the apparatus further includes:
[0066] An early warning module, configured to, when the query operation of querying the unit table of the faulty disk node fails, obtain the N-tuple of the faulty disk node through the central table, generate an early warning information through the N-tuple, and send the early warning information to the bound terminal.
[0067] In one embodiment, the apparatus further includes:
[0068] A fitting module, configured to determine a minimum value function of the corresponding CopySet function based on the CopySet function, determine a maximum value function of the corresponding data recovery function based on the data recovery function, and perform fitting through the system of equations of the minimum value function of the CopySet function and the maximum value function of the data recovery function to obtain the number of CopySets of the CopySet function.
[0069] For the specific limitations of the device for cluster adaptive data backup and recovery based on CopySet, reference may be made to the limitations of the method for cluster adaptive data backup and recovery based on CopySet in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned device for cluster adaptive data backup and recovery based on CopySet can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0070] Figure 5 An entity structure diagram of an electronic device is exemplified, as Figure 5 shown, the electronic device may include: a processor 301, a memory 302, a communication interface 303, and a communication bus 304. Among them, the processor 301, the memory 302, and the communication interface 303 complete mutual communication through the communication bus 304. The processor 301 can call the logical instructions in the memory 302 to execute the following method: obtain the configuration information of the distributed database, the distributed database is stored on different data clusters respectively, the data cluster includes corresponding disk nodes, and the configuration information includes: the degree of association between disk nodes, the size of the data cluster; determine the association function between disk nodes based on the degree of association, calculate the scattering width function of the data cluster based on the degree of association and the size of the data cluster, and determine the CopySet function by combining the association function and the scattering width function; obtain the bandwidth and data recovery rate of the disk node, determine the corresponding data recovery function according to the bandwidth and the data recovery rate, and solve the CopySet number of the CopySet function based on the CopySet function and the data recovery function; perform data replication on the data to be stored based on the CopySet number to obtain the corresponding replica data, store the replica data on the corresponding N replica nodes, generate the corresponding N-tuple and store it in the central table of the distributed database, and store the N-tuple in the unit table of the N replica nodes; poll the disk nodes through the polling module, when it is detected that a disk node fails, query the unit table of the failed disk node, determine the associated repair replica node through the query result, and perform data repair on the failed disk node through the repair replica node.
[0071] In addition, when the logical instructions in the above-mentioned memory 302 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0072] On the other hand, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the transmission methods provided in the above embodiments. For example, it includes: obtaining configuration information of a distributed database, where the distributed database is stored on different data clusters respectively, and the data clusters include corresponding disk nodes. The configuration information includes: the degree of association between disk nodes and the size of the data cluster; determining an association degree function between disk nodes based on the degree of association, calculating a scattering width function of the data cluster based on the degree of association and the size of the data cluster, and determining a CopySet function by combining the association degree function and the scattering width function; obtaining the bandwidth and data recovery rate of disk nodes, determining a corresponding data recovery function according to the bandwidth and data recovery rate, and solving the CopySet number of the CopySet function based on the CopySet function and the data recovery function; performing data replication on the data to be stored based on the CopySet number to obtain corresponding replica data, storing the replica data on corresponding N replica nodes, generating corresponding N-tuples and storing them in the central table of the distributed database, and storing N-tuples in the unit tables of the N replica nodes; polling the disk nodes through a polling module. When a disk node failure is detected, querying the unit table of the faulty disk node, determining the associated repair replica node through the query result, and performing data repair on the faulty disk node by the repair replica node.
[0073] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0074] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for cluster adaptive data backup and recovery based on CopySet, characterized in that, Including: Obtain the configuration information of the distributed database, where the distributed database is stored on different data clusters respectively, and the data clusters include corresponding disk nodes, and the configuration information includes: the degree of association between disk nodes and the size of the data cluster; Determine the association degree function between disk nodes based on the degree of association, calculate the scattering width function of the data cluster based on the degree of association and the size of the data cluster, and determine the CopySet function by combining the association degree function and the scattering width function. The scattering width function limits the distribution range of data based on the maximum interval between the original data and the replicas in the data cluster; Obtain the bandwidth and data recovery rate of the disk node, determine the corresponding data recovery function according to the bandwidth and data recovery rate, and solve the CopySet number of the CopySet function based on the CopySet function and the data recovery function; Perform data replication on the data to be stored based on the CopySet number to obtain the corresponding replica data, store the replica data on the corresponding N replica nodes, generate the corresponding N-tuple and store it in the central table of the distributed database, and store the N-tuple in the unit tables of the N replica nodes; Poll the disk nodes through a polling module. When a disk node failure is detected, query the unit table of the faulty disk node, determine the associated repair replica node through the query result, and perform data repair on the faulty disk node through the repair replica node.
2. The method for cluster adaptive data backup and recovery based on CopySet according to claim 1, wherein The method further includes: When the query operation of querying the unit table of the faulty disk node fails, obtain the N-tuple of the faulty disk node through the central table, generate a warning message through the N-tuple, and send the warning message to the bound terminal.
3. The method for cluster adaptive data backup and recovery based on CopySet according to claim 1, wherein The fitting and solving of the CopySet number of the CopySet function based on the CopySet function and the data recovery function includes: Determine the minimum value function of the corresponding CopySet function based on the CopySet function, determine the maximum value function of the corresponding data recovery function based on the data recovery function, and perform fitting through the system of equations of the CopySet function minimum value function and the data recovery function maximum value function to obtain the CopySet number of the CopySet function.
4. The method for cluster adaptive data backup and recovery based on CopySet according to claim 3, characterized in that, The system of equations includes: max(v)=max(fun(b,m)) min(m)=min(fun(S,r)) where v is the data recovery speed, b is the bandwidth, m is the number of CopySets, S is the scattering width, and r is the association degree function.
5. The method for cluster adaptive data backup and recovery based on CopySet according to claim 1, characterized in that After performing data repair on the faulty disk node through the repair replica node, it further includes: Generate a log file of the faulty disk node and store the log file in the central table.
6. A device for cluster adaptive data backup and recovery based on CopySet, characterized in that, The device includes: A first acquisition module, configured to acquire configuration information of a distributed database, where the distributed database is respectively stored on different data clusters, the data clusters include corresponding disk nodes, and the configuration information includes: the degree of association between disk nodes, and the size of the data cluster; A calculation module, configured to determine an association degree function between disk nodes based on the degree of association, calculate a scattering width function of the data cluster based on the degree of association and the size of the data cluster, and determine a CopySet function by combining the association degree function and the scattering width function, where the scattering width function limits the distribution range of data based on the maximum interval between the original data and replicas in the data cluster; A second acquisition module, configured to acquire the bandwidth and data recovery rate of the disk node, determine a corresponding data recovery function according to the bandwidth and the data recovery rate, and solve for the number of CopySets of the CopySet function based on the CopySet function and the data recovery function; A replication module, configured to perform data replication on the data to be stored based on the number of CopySets to obtain corresponding replica data, store the replica data on corresponding N replica nodes, generate corresponding N-tuples and store them in a central table of the distributed database, and store the N-tuples in a unit table of the N replica nodes; A polling module, configured to poll the disk nodes through the polling module. When a disk node failure is detected, query the unit table of the faulty disk node, determine associated repair replica nodes through the query result, and perform data repair on the faulty disk node through the repair replica nodes.
7. The apparatus for cluster adaptive data backup and recovery based on CopySet according to claim 6, wherein The apparatus further includes: An early warning module, configured to, when the query operation of querying the unit table of the faulty disk node fails, obtain the N-tuples of the faulty disk node through the central table, generate early warning information through the N-tuples, and send the early warning information to a bound terminal.
8. The apparatus for cluster adaptive data backup and recovery based on CopySet according to claim 6, characterized in that The apparatus further includes: A fitting module, configured to determine a minimum value function of the corresponding CopySet function based on the CopySet function, determine a maximum value function of the corresponding data recovery function based on the data recovery function, and perform fitting through a system of equations of the minimum value function of the CopySet function and the maximum value function of the data recovery function to obtain the number of CopySets of the CopySet function.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the method for cluster adaptive data backup and recovery based on CopySet according to any one of claims 1 to 5 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method for cluster adaptive data backup and recovery based on CopySet according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Cluster node fault recovery method and device, electronic equipment and storage medium
CN111124755A
Method of data replication in a distributed data storage system and corresponding device
US20120072689A1