Distributed backup system management method and device, equipment and storage medium

By constructing personalized fault prediction models and similarity matching, combined with resource optimization strategies, the shortcomings of distributed backup systems in fault monitoring and early warning are addressed, achieving efficient fault early warning and rapid recovery, and improving the reliability and security of the power system.

CN121764731APending Publication Date: 2026-03-31CHINA THREE GORGES CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing distributed backup systems lack in-depth consideration of specific business scenarios, node-specific characteristics, and dynamic environmental changes in fault monitoring and early warning. They are unable to identify new fault causes in a timely manner, and early warnings are delayed, leading to data loss and service interruption.

Method used

By acquiring the operational, hardware, and environmental information of each backup node, a personalized fault prediction model is constructed. Combined with similarity matching and resource optimization strategies, target backup nodes are identified for data backup and fault warning, with priority given to maintaining high-risk nodes.

Benefits of technology

It improves the accuracy of fault prediction and the reliability of the power system, reduces fault recovery time, and enhances the system's safety and fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121764731A_ABST
    Figure CN121764731A_ABST
Patent Text Reader

Abstract

The invention provides a distributed backup system management method and device, equipment and a storage medium, and relates to the technical field of distributed storage, and the method comprises the steps: obtaining the operation information and hardware information of each backup node and the environment information of a region where the backup node is located; determining a fault prediction model corresponding to each backup node according to the hardware information of each backup node, and inputting the operation information of each backup node into the corresponding fault prediction model to obtain a predicted fault rate of each backup node; when a fault backup node is detected, comparing the operation information and the hardware information of each backup node with the environment information of the area where the backup node is located, and determining the similarity between each detected backup node and the fault backup node according to a comparison result; and determining a target backup node according to the similarity between each detection backup node and the fault backup node and the predicted fault rate of each detection backup node. Through the management method of the distributed system, fault early warning can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed storage technology, and more specifically to a management method, apparatus, device, and storage medium for a distributed backup system. Background Technology

[0002] With the digitalization and intelligentization of power systems, data plays a crucial role in the production, management, and decision-making of the power industry. The stable operation of power systems depends on the high availability and integrity of data; therefore, distributed systems have become a key technology for ensuring the security of power data.

[0003] However, existing distributed backup systems often rely solely on general models for fault prediction. These models are typically built upon historical data and common failure modes, lacking in-depth consideration of specific business scenarios, node-specific characteristics, and dynamic environmental changes. For example, with the continuous upgrading of hardware and the evolution of system architecture, new fault triggers (such as abnormal heat dissipation in AI computing modules and compatibility issues with heterogeneous storage devices) constantly emerge, and general models cannot adapt to and accurately identify these potential threats in a timely manner. Furthermore, existing systems largely depend on preset thresholds for fault judgment, lacking sensitivity to subtle signs before a fault occurs (such as intermittent abnormal data from hardware sensors and subtle fluctuations in service response time), making early warning difficult. They also lack multi-dimensional data correlation analysis capabilities, failing to uncover potential correlations between faults from massive monitoring metrics (such as CPU utilization, network latency, and disk I / O queue length), resulting in delayed fault warnings, missed optimal intervention opportunities, and potentially serious consequences such as data loss and service interruption. This fails to meet the urgent needs of enterprises for high system reliability and stability. Existing distributed backup systems have significant shortcomings in fault monitoring and early warning. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a management method, apparatus, device and storage medium for a distributed backup system, thereby overcoming the shortcomings of the aforementioned distributed backup system in terms of fault monitoring and early warning.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A management method for a distributed backup system, wherein the distributed system includes multiple backup nodes; the management method includes: Obtain the operating information, hardware information, and environmental information of each backup node's location; Based on the hardware information of each backup node, a fault prediction model corresponding to each backup node is determined, and the operating information of each backup node is input into the corresponding fault prediction model to obtain the predicted failure rate of each backup node. When a faulty backup node is detected, the operating information, hardware information and environmental information of the area where each backup node is located are compared, and the similarity between each detected backup node and the faulty backup node is determined based on the comparison results. The detected backup node is any backup node other than the faulty backup node. The target backup node is determined based on the similarity between each detected backup node and the faulty backup node, as well as the predicted failure rate of each detected backup node.

[0006] The aforementioned management methods also include: The dominant backup nodes are determined based on the similarity between each detected backup node and the faulty backup node, and the predicted failure rate of each detected backup node, wherein the number of dominant backup nodes is not less than the number of target backup nodes. Obtain the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node; Based on the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node, node pairing is performed to determine the target advantageous backup nodes that are paired with each target backup node. Backup data from each target backup node is backed up to the paired target dominant backup node.

[0007] The aforementioned distributed system includes multiple backup nodes of different levels, with higher-level backup nodes having a higher priority than lower-level nodes; the management method also includes: Obtain the priority of each data shard; The priority of each data shard is used to match the level of each backup node; Based on the matching information and continuity order of each data segment, the corresponding backup node for each data segment is determined, and each data segment is stored in the corresponding backup node.

[0008] The aforementioned management methods also include: Multiple advantageous backup nodes are determined based on the similarity between each detected backup node and the faulty backup node, as well as the predicted failure rate of each detected backup node. Obtain the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node; Based on the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node, determine the matching score of each advantageous backup node relative to each target backup node. The target advantageous backup nodes are determined based on the matching scores and rank distances of each advantageous backup node relative to each target backup node. Backup data from each target backup node is backed up to the paired target dominant backup node.

[0009] The aforementioned management methods also include: Upon receiving the maintenance completion instruction for the corresponding target backup node, data recovery processing is performed on the target backup node based on the backup data on the corresponding target superior backup node; If, within a preset time period after the target backup node performs data recovery processing, the target backup node is detected to be in normal operating condition, then the backup data belonging to the target backup node in the corresponding target dominant backup node will be deleted.

[0010] The aforementioned management methods also include: Obtain load monitoring data from the distributed system; Determine the resource requirements of each backup node based on load monitoring data; Resource allocation is based on the resource requirements and level of each backup node, with priority given to supplying resources to higher-level backup nodes.

[0011] The above-mentioned determination of the target backup node based on the similarity between each detected backup node and the failed backup node, and the predicted failure rate of each detected backup node, includes: A first evaluation score for each detection backup node is determined based on the similarity between each detection backup node and the faulty backup node, and a second evaluation score for each detection backup node is determined based on the predicted failure rate of each detection backup node. The first evaluation score is positively correlated with the similarity, and the second evaluation score is positively correlated with the predicted failure rate. The target evaluation score is obtained by weighted summing of the first and second evaluation scores. The detection backup node whose target evaluation score is greater than the first preset threshold is determined as the target backup node; Dominant backup nodes are determined based on the similarity between each detected backup node and the failed backup node, and the predicted failure rate of each detected backup node, including: The detection backup nodes whose target evaluation scores are less than the second preset threshold are identified as candidate backup nodes, wherein the second preset threshold is less than the first preset threshold; When the number of candidate backup nodes is greater than or equal to the number of target backup nodes, each candidate backup node is determined as the dominant backup node; When the number of candidate backup nodes is less than the number of target backup nodes, the second preset threshold is increased by a step value, and the process returns to the step of determining the detected backup nodes whose target evaluation scores are less than the second preset threshold as candidate backup nodes.

[0012] The aforementioned fault prediction model determines the corresponding fault prediction model for each backup node based on the hardware information of each backup node, and inputs the operating information of each backup node into the corresponding fault prediction model to obtain the predicted fault rate of each backup node, specifically including: Based on the hardware information of backup nodes, the aging patterns and failure modes of different hardware components are analyzed. The parameters contained in the aging patterns and failure modes of different hardware components are used as key parameters for predicting disk failures. By collecting historical hardware failure data and combining statistical methods and machine learning algorithms, a personalized failure prediction model for different hardware configurations is constructed.

[0013] The statistical methods described above employ survival analysis, and the machine learning algorithm used is the random forest algorithm. The specific process of constructing a personalized fault prediction model for different hardware configurations includes: Survival analysis method indicator definition: Survival function Hardware in time The probability that no fault occurs, i.e. ,in For the time of failure; Risk function The hardware has been running normally for the specified time. Under the condition of "", the probability of a failure occurring instantaneously in the next instant is, i.e. ; Kaplan-Meier survival function estimation: For samples sorted by failure time, ,in For the first The time of the fault is recorded. for The number of hardware devices still operating normally before the specified time, i.e., the "risk set size". for The number of hardware failures at any given time is then estimated using the following formula for the survival function: ; The input features of the random forest algorithm include static hardware configuration, dynamic running status, environmental parameters, and "risk indicators" output by survival analysis. The definition of random forest includes The first decision tree, the Tree samples In time The internal fault prediction results are (1 represents predicted failure, 0 represents normal operation), then the final failure probability is: ; Combining the "time trend" of survival analysis with the "feature association" of random forest, the output is tailored to the hardware. The personalized failure probability is: ; in, For hardware Configuration grouping, Let this be the conditional survival function for this group; The fusion weights are dynamically adjusted based on the sample size of the configured groups: the larger the sample size, the higher the fusion weight. The larger the sample size, the more it depends on the statistical regularities of survival analysis; when the sample size is small, Reduced, relying on feature learning from random forests; By configuring the conditional survival function for grouping and adjusting the fusion weights, the model can make differentiated predictions for different hardware outputs, avoiding the prediction bias of a "general model" for niche configurations.

[0014] The hardware information of the aforementioned backup nodes includes CPU model and number of cores, memory capacity and type, disk read / write performance parameters, and network interface bandwidth. The aging patterns and failure modes of different hardware components include parameters such as the mean time between failures (MTBF), power-on time, and number of read / write cycles for mechanical hard drives (HDDs).

[0015] In the above-mentioned process of comparing the operational information, hardware information, and environmental information of each backup node when a faulty backup node is detected, and determining the similarity between each detected backup node and the faulty backup node based on the comparison results, the formula for calculating the similarity is as follows: ; in, For the first The similarity score between each detection backup node and the failed backup node. , , The weighting coefficients for operational information, hardware information, and environmental information are respectively, satisfying... , , , The first The similarity values ​​of each backup node in terms of runtime information, hardware information, and environmental information; The runtime information includes multiple metrics, such as CPU utilization, memory usage, and disk I / O rate. The runtime information similarity Sim_r(i) can be calculated using cosine similarity. ; Where: R_i=[R_i1,R_i2,…,R_in] is the operation indicator vector of the backup node i, R_f=[R_f1,R_f2,…,R_fn] is the operation indicator vector of the faulty backup node, ⋅ represents the vector dot product, and ||⋅|| represents the vector norm; Hardware information includes a mixture of discrete and continuous metrics; hardware information similarity. The weighted combination method can be used to calculate: ; in, for The weight of each hardware metric, For the first Similarity of hardware metrics, continuous metrics: ; in, and These are the metric values ​​for detecting backup nodes and faulty backup nodes, respectively. This is the scaling factor; Discrete index: ; Environmental information includes geographical location, data center parameters, etc., and environmental information similarity. Distance functions can be used for calculation: ; in, The comprehensive distance for environmental parameters can be defined as: ; in, and These are environmental parameter values. The standard deviation of the parameter is used for normalization; This is the distance scaling factor, which controls the rate at which similarity decays.

[0016] The above-mentioned node pairing based on the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node, to determine the target advantageous backup node paired with each target backup node specifically includes: A matching score is used to evaluate each advantageous backup node. For the target backup node being matched, the evaluation considers three aspects: available storage space, network latency, and geographical distance. Available storage space is positively correlated with the matching score, while network latency is negatively correlated. Multiple distance ranges are set for geographical distance, with the highest matching score corresponding to the middle range. Increasing or decreasing the distance will lead to a decrease in the matching score. Based on the above matching rules, a corresponding matching formula is pre-constructed. The total matching score formula is: ; in, For the first The matching score of each advantageous backup node. , , The weighting coefficients for storage space, network latency, and geographical distance are respectively, satisfying... , , , These are the scoring functions for the corresponding dimensions; The formula for scoring available storage space is: ; in, For nodes Available storage space The maximum available storage space among all candidate nodes; The formula for calculating network latency score is: ; in, For nodes Network latency with the target node The scaling parameter controls the rate at which the delay decays with respect to the fraction; The formula for calculating geographical distance score is: ; in: , This is the dividing point between geographical distance intervals. ; The parameters satisfy: (, , .

[0017] The management device using the above-described distributed backup system management method includes: The acquisition module is used to acquire the operating information, hardware information, and environmental information of the region where each backup node is located; The prediction module is used to determine the corresponding fault prediction model for each backup node based on the hardware information of each backup node, and input the operating information of each backup node into the corresponding fault prediction model to obtain the predicted failure rate of each backup node. The comparison module is used to compare the operating information, hardware information and environmental information of each backup node when a faulty backup node is detected, and to determine the similarity between each detected backup node and the faulty backup node based on the comparison results. The detected backup node is a backup node other than the faulty backup node. The determination module is used to determine the target backup node based on the similarity between each detected backup node and the faulty backup node, as well as the predicted failure rate of each detected backup node.

[0018] An electronic device using the above-described management method for a distributed backup system includes a processor and a memory; the memory stores a computer program, wherein the computer program implements the management method for the distributed backup system when executed by the processor.

[0019] The storage medium of the aforementioned distributed backup system stores a computer program, which, when executed by a processor, implements the management method of the distributed backup system.

[0020] The management method, apparatus, device, and storage medium of the distributed backup system mentioned in this invention have the following beneficial effects: 1. The distributed system management method of this application obtains the operating information, hardware information, and environmental information of each backup node. Then, based on the hardware information of each backup node, a fault prediction model is determined for each backup node. The operating information of each backup node is input into the corresponding fault prediction model to obtain the predicted failure rate of each backup node. The fault prediction model is matched based on the hardware information to improve the prediction accuracy of the model. Furthermore, when a faulty backup node is detected, the operating information, hardware information, and environmental information of each backup node are compared, and the similarity between each detected backup node and the faulty backup node is determined based on the comparison results. Finally, the target backup node is determined based on the similarity between each detected backup node and the faulty backup node, as well as the predicted failure rate of each detected backup node. This method rapidly identifies candidate faulty backup nodes through a combination of similarity matching and model prediction, achieving fault early warning.

[0021] 2. This application uses the operational information, hardware information, and environmental information of each backup node for fault monitoring and early warning, taking into account a comprehensive range of factors and achieving high prediction accuracy. Furthermore, the target backup node can be identified as a key target for fault monitoring and priority maintenance, enabling rapid and effective acquisition of fault information and swift maintenance after a fault occurs. This results in shorter power system recovery time and effectively improves the reliability and security of the power system. Attached Figure Description

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart illustrating a management method for a distributed system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the data backup process for a target backup node in one embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the pairing of a target backup node and a dominant backup node in one embodiment of the present invention; Figure 4 This is a schematic diagram of the data backup process for the target backup node in another embodiment of the present invention; Figure 5 This is a schematic diagram illustrating the pairing of a target backup node and a dominant backup node in another embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a management device for a distributed system in one embodiment of the present invention. Detailed Implementation

[0023] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0024] Example 1: As described in the background section, with the digitalization and intelligentization of power systems, data plays a crucial role in the production, management, and decision-making of the power industry. The stable operation of power systems depends on the high availability and integrity of data; therefore, distributed systems have become a key technology for ensuring the security of power data.

[0025] However, existing distributed backup systems have significant shortcomings in fault monitoring and early warning.

[0026] Based on this, such as Figure 1 As shown, this application provides a management method for a distributed system, which includes multiple backup nodes. The management method for the distributed system includes the following steps S101 to S104.

[0027] S101: Obtain the operating information, hardware information, and environmental information of each backup node's location.

[0028] The operational information may include runtime, memory usage, CPU utilization, disk I / O rate, and network bandwidth consumption. In this embodiment, the operational information obtained can be at least one of these parameters to monitor node load and task processing efficiency in real time. Hardware information may include CPU model, memory capacity, and disk type. In this embodiment, the hardware information obtained can be at least one of these parameters. Environmental information about the backup node's location may include climate information, data center temperature, and humidity. In this embodiment, the environmental information obtained can be at least one of these parameters.

[0029] It's important to note that each backup node is responsible for storing a portion of the backup data. A consistent hashing algorithm can be used to distribute the data evenly across the nodes, ensuring high availability and fault tolerance. The consistent hashing algorithm maps data to a circular space and selects storage nodes based on the hash value of the data, thus achieving even data distribution and load balancing. In applications, the data distribution can be checked periodically, and the distribution can be dynamically adjusted based on node load.

[0030] As can be understood, the multi-backup node distributed architecture design described above achieves multi-node distributed storage, ensuring data availability even if some nodes fail. Furthermore, data backup across multiple nodes reduces the risk of single points of failure. Additionally, the use of a consistent hashing algorithm ensures even data distribution, preventing excessive load on any single node.

[0031] S102: Determine the corresponding fault prediction model for each backup node based on the hardware information of each backup node, and input the operating information of each backup node into the corresponding fault prediction model to obtain the predicted failure rate of each backup node.

[0032] In applications, based on the hardware information of backup nodes, such as CPU model and core count, memory capacity and type, disk read / write performance parameters, and network interface bandwidth, the aging patterns and failure modes of different hardware components can be analyzed. For example, the mean time between failures (MTBF), power-on time, and read / write cycles of hard disk drives (HDDs) can serve as key parameters for predicting disk failures; while the CPU's manufacturing process, thermal design, and long-term load rate are closely related to the CPU failure probability. By collecting a large amount of historical hardware failure data and combining statistical methods (such as survival analysis and reliability function modeling) and machine learning algorithms (such as random forests and LSTM neural networks), personalized failure prediction models for different hardware configurations can be constructed. For example, when backup nodes are categorized into different levels based on hardware performance, failure prediction models for different levels can be constructed.

[0033] For example, statistical methods can be survival analysis, which processes time-dependent fault data, and machine learning algorithms can be random forest algorithms, which capture the nonlinear correlation between features and faults.

[0034] Survival analysis methods focus on "the time from hardware operation to its first failure," or "failure time," and are particularly suitable for scenarios involving "censored data," such as when some hardware remains operational until the end of the observation period or is taken offline for non-failure reasons.

[0035] Definition of core indicators: Survival function Hardware in time The probability that no fault occurs, i.e. ,in This refers to the time of failure.

[0036] Risk function The hardware has been running normally for the specified time. Under the condition of "", the probability of a failure occurring instantaneously in the next instant is, i.e. .

[0037] Kaplan-Meier survival function estimation (nonparametric method, handling censored data): for samples sorted by failure time ( ,in For the first (time of the fault) recorded for The number of hardware devices still functioning normally before the specified time (i.e., the "risk set size"). for The number of hardware failures at any given time is then estimated using the following formula for the survival function: ; Random forests use ensemble learning of multiple decision trees to handle the nonlinear relationship between high-dimensional features (such as hardware configuration, operating status, and environmental parameters) and failure probability, and output the failure probability of specific hardware within a given time period.

[0038] The input features of the Random Forest algorithm can include static hardware configurations (such as the number of CPU cores). Hard drive capacity Dynamic operating status (such as CPU temperature) Memory usage ), environmental parameters (such as computer room humidity) Voltage fluctuations ), and the "risk indicators" output by survival analysis (such as current time) corresponding ).

[0039] The definition of random forest includes The first decision tree, the Tree samples In time The internal fault prediction results are (1 represents predicted failure, 0 represents normal operation), then the final failure probability is: ; Combining the "time trend" of survival analysis with the "feature association" of random forest, the output is tailored to the hardware. The personalized failure probability is: ; in, For hardware Configuration groups (e.g., "Model A Hard Drive", "Model B Server") This is the conditional survival function for this group (reflecting the common failure time patterns of hardware with the same configuration); The fusion weights are dynamically adjusted based on the sample size of the configured groups: the larger the sample size, the higher the fusion weight. The larger the sample size, the more it depends on the statistical regularities of survival analysis; when the sample size is small, (Reduces reliance on feature learning from random forests).

[0040] By configuring the conditional survival function for grouping and adjusting the fusion weights, the model can output differentiated predictions for different hardware (such as old servers and new servers), avoiding the prediction bias of the "general model" for niche configurations.

[0041] Based on this, the real-time operational information of the backup nodes is input into the corresponding fault prediction model. This operational information reflects the hardware's working condition in actual operating scenarios; for example, longer runtime increases the risk of failure; continuous high CPU load accelerates aging, leading to a higher failure rate; frequent random disk read / write operations exacerbate disk wear. Through in-depth analysis and feature extraction of the operational data, combined with pre-trained parameters, the fault prediction model can calculate the predicted failure rate of each backup node within a specific future time period.

[0042] S103: When a faulty backup node is detected, the operating information, hardware information and environmental information of the area where each backup node is located are compared, and the similarity between each detected backup node and the faulty backup node is determined based on the comparison results. The detected backup node is a backup node other than the faulty backup node.

[0043] In the application, a heartbeat detection mechanism can be deployed on each backup node. Each node periodically sends a heartbeat signal to other nodes, for example, every minute, to check their online status. Various methods can be used, such as heartbeat detection, scheduled detection, and event-driven detection, to quickly identify faulty nodes. Scheduled detection mechanisms can check node performance metrics at regular intervals, such as every 5 minutes, to detect potential faults. Deploying an event-driven detection mechanism can monitor system events to detect fault information in a timely manner.

[0044] It is understandable that the operational information, hardware information, and environmental information of each backup node can be compared in advance. Based on the comparison results, the similarity between each detected backup node and the failed backup node in these three aspects—operational information, hardware information, and environmental information—can be determined. By setting reasonable weighting coefficients, the comparison results of operational information, hardware information, and environmental information are quantified and integrated, and the similarity between each detected backup node and the failed backup node is finally calculated. This provides accurate guidance for subsequent fault prediction and node repair, ensuring that the distributed system can quickly recover and stabilize.

[0045] For example, the similarity calculation formula can be as follows: ; in, For the first The similarity score between each detection backup node and the failed backup node. , , The weighting coefficients for operational information, hardware information, and environmental information are respectively, satisfying... , , , The first The similarity values ​​of each backup node in terms of runtime information, hardware information, and environmental information.

[0046] Operational information includes multiple metrics such as CPU utilization, memory usage, and disk I / O speed. The similarity of operational information is also important. Cosine similarity can be used for calculation: ; in: To detect backup nodes The vector of operating indicators, This is a vector of operational metrics for the faulty backup node. Represents the vector dot product. Represents the vector norm.

[0047] Hardware information includes a mixture of discrete and continuous metrics; hardware information similarity. The weighted combination method can be used to calculate: ; in, for The weights of each hardware metric are as follows: CPU weight 0.3, memory weight 0.2, disk weight 0.5. For the first Similarity of hardware metrics, including continuous metrics (such as CPU clock speed and memory capacity): ; in, and These are the metric values ​​for detecting backup nodes and faulty backup nodes, respectively. This is the scaling factor.

[0048] Discrete metrics such as disk type (SSD / HDD) and network interface type: ; Environmental information includes geographical location, data center parameters, etc., and environmental information similarity. Distance functions can be used for calculation: ; in, The comprehensive distance for environmental parameters can be defined as: ; in, and These are environmental parameter values. The standard deviation of the parameter is used for normalization. This is the distance scaling factor, which controls the rate at which similarity decays.

[0049] For example, if , , The similarity calculation result of a certain backup node is as follows: , , .

[0050] The total similarity score for this node is: ; This formula allows the system to quantitatively assess the similarity between each backup node and the faulty node, providing a basis for subsequent fault warnings and maintenance decisions.

[0051] S104: Determine the target backup node based on the similarity between each detected backup node and the faulty backup node, as well as the predicted failure rate of each detected backup node.

[0052] It is understandable that a high similarity between detected backup nodes and faulty backup nodes indicates commonalities and a high potential risk of failure. The predicted failure rate directly reflects the likelihood of a node failing within a future period. Combining these two factors allows for the rapid and accurate identification of candidate faulty backup nodes, i.e., target backup nodes. These target backup nodes can then be designated as key targets for fault monitoring and maintenance. When a target backup node fails, its fault information can be quickly and effectively obtained, enabling rapid maintenance based on this information. This results in shorter power system recovery time and effectively improves the reliability and security of the power system.

[0053] The aforementioned distributed system management method acquires the operational information, hardware information, and environmental information of each backup node. Then, based on the hardware information of each backup node, a fault prediction model is determined for each node. The operational information of each backup node is input into the corresponding fault prediction model to obtain the predicted failure rate of each backup node. Matching the fault prediction model with hardware information improves the model's prediction accuracy. Furthermore, when a faulty backup node is detected, the operational information, hardware information, and environmental information of each backup node are compared, and the similarity between each detected backup node and the faulty backup node is determined based on the comparison results. Finally, based on the similarity between each detected backup node and the faulty backup node, and the predicted failure rate of each detected backup node, the target backup node is determined. This method, combining similarity matching and model prediction, rapidly identifies candidate faulty backup nodes, achieving fault early warning. Fault monitoring and early warning based on the operational information, hardware information, and environmental information of each backup node considers a comprehensive range of factors and achieves high prediction accuracy.

[0054] In some embodiments, such as Figure 2 As shown, the management method of this distributed system also includes the following steps S201 to S204.

[0055] S201: Determine the dominant backup nodes based on the similarity between each detected backup node and the faulty backup node, and the predicted failure rate of each detected backup node, wherein the number of dominant backup nodes is not less than the number of target backup nodes.

[0056] It's understandable that a low similarity between detected backup nodes and failed backup nodes means that the detected backup nodes are less likely to be affected by the same failure causes. Meanwhile, the predicted failure rate uses historical data and machine learning algorithms to quantitatively assess the probability of future node failures; nodes with lower failure rates have higher operational stability. Combining these two factors allows for the rapid and accurate identification of backup nodes with a low probability of failure at a specific time—these are the advantageous backup nodes.

[0057] S202: Obtain the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node.

[0058] S203: Based on the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node, perform node pairing to determine the target advantageous backup node paired with each target backup node.

[0059] It's understandable that available storage space determines the ability of a primary backup node to handle data migration from a failed node. Insufficient available space may lead to data write failures or partial data loss, impacting business continuity. Network latency directly affects data synchronization efficiency; low latency ensures rapid data replication and recovery during failover, preventing prolonged service interruptions. Geographical distance is related to disaster recovery capabilities. A more distant primary backup node can effectively withstand regional disasters such as earthquakes or data center fires, preventing simultaneous damage to both the target and primary backup nodes. Geographical distance also indirectly affects network latency and transmission costs. By considering these three key factors, the suitability of a primary backup node can be more scientifically assessed. Prioritizing primary backup nodes with sufficient available storage space, low network latency, and suitable geographical distance as the target primary backup node improves the reliability, performance, and disaster recovery capabilities of the distributed system.

[0060] For example, a matching score can be used to evaluate each advantageous backup node. For the target backup node being matched, the evaluation considers three aspects: available storage space, network latency, and geographical distance. Available storage space is positively correlated with the matching score, while network latency is negatively correlated. Multiple distance ranges can be set for geographical distance; the range with the highest matching score corresponds to the middle distance, and increasing or decreasing the distance will lead to a decrease in the matching score. In applications, a corresponding matching formula can be pre-constructed based on the above matching rules. For example, the formula for the total matching score is: ; in, For the first The matching score of each advantageous backup node. , , The weighting coefficients for storage space, network latency, and geographical distance are respectively, satisfying... , , , These are the scoring functions for the corresponding dimensions.

[0061] The formula for scoring available storage space can be: ; in, For nodes Available storage space The largest available storage space among all candidate nodes The formula for calculating network latency score can be: ; in, For nodes Network latency with the target node The scaling parameter controls the rate at which the delay decays with respect to the fraction. The formula for calculating geographic distance score (piecewise function) can be: ; in:, , The geographical distance interval boundary point ( The parameters satisfy: (Scores decrease in the nearest interval) (The scores in the middle range increase). (Further distance intervals result in decreasing scores). For example, if the parameters of a dominant backup node are: , , , , (Located in the middle section) Then its matching score is: ; The matching score of each advantageous backup node relative to the currently matched target backup node is determined based on the matching formula. Then, the advantageous backup node with the highest score is determined as the target advantageous backup node. For example... Figure 3 As shown, the first target backup node on the left and the third dominant backup node on the right have the highest matching score. Therefore, the first target backup node on the left and the third dominant backup node on the right are paired, and the third dominant backup node on the right is the target dominant backup node of the first target backup node on the left.

[0062] Once a dominant backup node is identified as the target dominant backup node, in one example, during the node pairing process for subsequent target backup nodes, the dominant backup node that has been identified as the target dominant backup node can be deleted from the pairing nodes; in another example, the available storage space of the target dominant backup node can be updated by subtracting the storage space occupied by the backup data in the corresponding target backup node from the original available storage space to obtain the new available storage space, and then the next node pairing process can be carried out based on the new available storage space.

[0063] S204: Back up the backup data of each target backup node to the paired target dominant backup node.

[0064] It can be understood that the target backup node is a backup node with a high failure rate at a specific time, while the target dominant backup node is a dominant backup node with a low failure rate at a specific time and is precisely matched with the corresponding target backup node. Therefore, even if the target backup node experiences an unexpected failure or data corruption, the paired target dominant backup node can respond quickly and achieve rapid recovery with complete and real-time updated backup data. This not only effectively reduces the risk of data loss but also enhances the system's fault tolerance in complex failure scenarios, effectively improving the data security and reliability of the distributed system.

[0065] In some embodiments, the distributed system includes multiple backup nodes of different levels, with higher-level backup nodes having a higher priority than lower-level nodes. The management method for this distributed system further includes: obtaining the priority of each data shard; matching the level of each backup node according to the priority of each data shard; determining the corresponding backup node for storing each data shard based on the matching information and continuity order of the data shards; and storing each data shard in the corresponding backup node.

[0066] It's understandable that dividing backup nodes into multiple priority levels allows for differentiated storage management based on data importance. For example, backup nodes can be categorized into high, medium, and low levels based on factors such as performance and reliability. High-level backup nodes utilize high-performance hardware like SSD storage and high-frequency CPUs, equipped with redundant power supplies and networks, providing stronger fault tolerance. Low-level nodes, on the other hand, may use ordinary hardware suitable for non-core data storage. Simultaneously, the system assigns corresponding priorities to each data shard by analyzing factors such as the data's business value, access frequency, and modification frequency. For instance, highly sensitive data like core user information is marked as high priority, while system logs and temporary cache data are classified as low priority.

[0067] Based on data and node priorities, the system employs a matching algorithm to precisely map sharded data to backup nodes: high-priority data is preferentially stored on high-level backup nodes to ensure data security and fast access; medium- and low-priority data are allocated to medium- and low-level nodes to achieve rational utilization of storage resources. Furthermore, to ensure data integrity and retrieval efficiency, the system also considers the sequential order of data, such as time-series data and relational table data, storing logically continuous sharded data on the same or adjacent nodes to reduce cross-node data retrieval overhead. This hierarchical matching and sequential storage strategy not only meets the high-reliability storage requirements of core data but also reduces overall storage costs and improves the comprehensive performance of the distributed storage system.

[0068] Based on the previous embodiment, in some embodiments, such as Figure 4As shown, the management method of this distributed system also includes the following steps S401 to S405.

[0069] S401: Based on the similarity between each detected backup node and the faulty backup node, and the predicted failure rate of each detected backup node, determine multiple advantageous backup nodes.

[0070] By combining the similarity between the detected backup node and the failed backup node, as well as the predicted failure rate of the detected backup node, we can quickly and accurately identify the backup node with a low probability of failure at a specific time, i.e., the superior backup node.

[0071] S402: Obtain the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node.

[0072] S403: Based on the available storage space of each advantageous backup node, as well as the network latency and geographical distance to each target backup node, determine the matching score of each advantageous backup node relative to each target backup node.

[0073] It's understandable that available storage space determines the ability of a primary backup node to handle data migration from a failed node. Insufficient available space may lead to data write failures or partial data loss, impacting business continuity. Network latency directly affects data synchronization efficiency; low latency ensures rapid data replication and recovery during failover, preventing prolonged service interruptions. Geographical distance is related to disaster recovery capabilities. A more distant primary backup node can effectively withstand regional disasters such as earthquakes or data center fires, preventing simultaneous damage to both the target and primary backup nodes. Geographical distance also indirectly affects network latency and transmission costs. By considering these three key factors, the suitability of a primary backup node can be more scientifically assessed. Prioritizing primary backup nodes with sufficient available storage space, low network latency, and suitable geographical distance as the target primary backup node improves the reliability, performance, and disaster recovery capabilities of the distributed system.

[0074] S404: Determine the target dominant backup node to be paired with each target backup node based on the matching score and level distance of each dominant backup node relative to each target backup node.

[0075] The matching score is a quantitative assessment of the compatibility between the superior backup node and the target backup node across multiple dimensions. Its calculation incorporates indicators such as available storage space, network latency, and geographical distance. The specific calculation principle can be found in the description of the aforementioned embodiments, and will not be repeated here. The rank distance reflects the priority differences between nodes; a large rank distance exists between a high-rank superior backup node and a low-rank target backup node, and vice versa. During the decision-making process, the system prioritizes superior backup nodes with high matching scores and close rank distances as pairing partners. This strategy, based on quantitative assessment and priority trade-offs, accurately determines the optimal pairing scheme, ensuring that high-priority tasks are executed first, and enhancing the system's fault tolerance and data management efficiency.

[0076] For example, such as Figure 5 As shown, backup nodes can be divided into high-level and low-level nodes, with a level distance of 1 between them. A level distance of 1 corresponds to a score of 50; a distance of 0 between nodes of the same level corresponds to a score of 100. If the weights corresponding to the matching score and level distance are both 0.5, then the final score determined based on the matching score and level distance can be calculated as follows: Figure 5 As shown, Figure 5 In the results, the first target backup node on the left and the second dominant backup node on the right have the highest final scores, which is 90 (0.5×80+0.5×100). The first target backup node on the left and the third dominant backup node on the right have the highest final scores, which is 80 (0.5×60+0.5×100). Therefore, the second dominant backup node on the right is the target dominant backup node of the first target backup node on the left, and the third dominant backup node on the right is the target dominant backup node of the second target backup node on the left.

[0077] S405: Back up the backup data of each target backup node to the paired target dominant backup node.

[0078] It can be understood that the target backup node is a backup node with a high failure rate at a specific time, while the target dominant backup node is a superior backup node with a low failure rate at a specific time and is precisely matched with the corresponding target backup node. Therefore, even if the target backup node experiences an unexpected failure or data corruption, the paired target dominant backup node can respond quickly and achieve rapid recovery with complete and real-time updated backup data. In addition, since the priority of the target dominant backup node is close to that of the target backup node, the target dominant backup node can also directly serve as a response node, ensuring that high-priority tasks are executed first.

[0079] In some embodiments, the management method of the distributed system further includes: after receiving the maintenance completion instruction of the corresponding target backup node, performing data recovery processing on the target backup node based on the backup data on the corresponding target dominant backup node; and after a preset time period following the data recovery processing of the target backup node, if it is detected that the target backup node is in normal operation, deleting the backup data belonging to the target backup node in the corresponding target dominant backup node.

[0080] It is understandable that upon receiving the maintenance completion instruction from the target backup node, it indicates that the target backup node should be in normal operating condition, and the system can initiate the data recovery process for the target backup node. At this time, the target dominant backup node paired with the target backup node serves as the core source of data recovery, and its stored backup data will be quickly and accurately transmitted back to the target backup node. This process can be achieved through efficient data transmission protocols and incremental synchronization technology, which can significantly reduce the amount of data transmission and recovery time, ensuring that business operations are quickly restored to normal.

[0081] After data recovery is complete, the system will enter a pre-set monitoring period. By collecting real-time operational metrics such as CPU utilization, memory usage, and disk I / O response, combined with log analysis and link tracing, the system will rigorously check whether the target backup node has fully recovered to normal operating status. If, during this period, all metrics of the target backup node meet health standards, and there is no data loss, service interruption, or abnormal errors, the data recovery is considered successful. The system will then automatically trigger a data cleanup mechanism, deleting redundant backup data belonging to the target backup node. This operation frees up storage resources, reduces system storage costs, avoids management complexity caused by data redundancy, and ensures the timeliness and validity of backup data, providing reliable support for subsequent fault recovery and maintaining the efficient and stable operation of the distributed storage system.

[0082] In some embodiments, the management method of the distributed system further includes: acquiring load monitoring data of the distributed system; determining the resource requirements of each backup node based on the load monitoring data; and allocating resources according to the resource requirements of each backup node and the level of each backup node, so as to prioritize the supply of resources to high-level backup nodes.

[0083] The load status of the power system includes CPU utilization, memory utilization, and network bandwidth utilization. For example, the load status of the power system can be monitored at regular intervals, such as every 5 minutes; and CPU, memory, and network bandwidth utilization data of each backup node can be collected.

[0084] In practice, specifically, based on load monitoring data and combined with historical load trends and business demand models, the system uses a resource demand prediction algorithm to accurately assess the resource requirements of each backup node in different scenarios. For example, during peak business periods, backup nodes with high concurrency access have significantly increased CPU and memory resource requirements, while nodes that handle a large amount of data writing rely more on disk I / O performance.

[0085] After determining resource requirements, the system allocates resources differentiatedly based on the backup node's priority level. High-priority backup nodes, responsible for core data storage and critical business support, have their reliability and performance directly impacting system stability and are therefore given the highest priority. During resource allocation, the system prioritizes resource supply to high-priority backup nodes to ensure rapid read / write of critical data and service continuity. Medium and low-priority backup nodes are dynamically allocated resources based on their needs and remaining available resources, meeting basic operational requirements while avoiding resource waste. Through this load-aware and priority-based resource allocation strategy, the distributed system improves the reliability of core services, achieves efficient resource utilization, and enhances the overall system performance and stability.

[0086] In related technologies, the load of power systems is dynamically changing, and traditional static resource allocation strategies cannot meet real-time demands. In this embodiment, however, the dynamic resource scheduling mechanism can dynamically adjust resource allocation based on real-time load, improving resource utilization, avoiding resource waste, and monitoring load changes in real time for rapid adjustments. Furthermore, it ensures that backup and recovery tasks for critical data are prioritized, optimizing system performance.

[0087] In some embodiments, determining a target backup node based on the similarity between each detected backup node and a faulty backup node, and the predicted failure rate of each detected backup node, includes: determining a first evaluation score for each detected backup node based on the similarity between each detected backup node and a faulty backup node, and determining a second evaluation score for each detected backup node based on the predicted failure rate of each detected backup node; performing a weighted summation of the first evaluation score and the second evaluation score to obtain a target evaluation score; and determining the detected backup node with a target evaluation score greater than a first preset threshold as the target backup node. Wherein, the first evaluation score is positively correlated with the similarity, and the second evaluation score is positively correlated with the predicted failure rate.

[0088] The process of determining dominant backup nodes based on the similarity between each detected backup node and the faulty backup node, and the predicted failure rate of each detected backup node, includes: identifying detected backup nodes with a target evaluation score less than a second preset threshold as candidate backup nodes; when the number of candidate backup nodes is greater than or equal to the number of target backup nodes, identifying each candidate backup node as a dominant backup node; when the number of candidate backup nodes is less than the number of target backup nodes, increasing the second preset threshold by a step value, and returning to the step of identifying detected backup nodes with a target evaluation score less than the second preset threshold as candidate backup nodes. The second preset threshold is less than the first preset threshold.

[0089] In application, when identifying target backup nodes, the similarity between the detected backup node and the failed backup node is first calculated based on their hardware configuration, operating status, and environmental parameters. Cosine similarity or Euclidean distance algorithms are used to map this similarity to a first evaluation score, giving higher scores to nodes with higher similarity, thus identifying nodes potentially affected by the same fault causes. Simultaneously, machine learning models such as random forests or LSTM models are used to analyze the input features of the detected backup nodes, predicting the failure probability of each node and converting this into a second evaluation score, giving higher scores to nodes with higher predicted failure rates. Subsequently, based on the system's focus on risk factors, different weights are assigned to the first and second evaluation scores (e.g., 40% for the first evaluation score and 60% for the second). A weighted sum is then used to obtain a target evaluation score that comprehensively reflects the node's risk level. The system identifies detected backup nodes with target evaluation scores greater than a first preset threshold as target backup nodes; these nodes, due to their higher risk values, are prioritized for inclusion in fault monitoring.

[0090] When selecting superior backup nodes, the system identifies backup nodes with a target evaluation score lower than a second preset threshold (this threshold is lower than the first preset threshold and is used to select relatively low-risk nodes) as candidate backup nodes. If the number of candidate backup nodes is greater than or equal to the required number of target backup nodes, it indicates that low-risk node resources are sufficient, and these candidate nodes are directly identified as superior backup nodes, providing reliable backup resources for failover. If the number of candidate backup nodes is insufficient, the system increases the second preset threshold by a step value, such as 5 points, and re-executes the selection operation to expand the candidate range until the number of candidate backup nodes meets the requirements. This dynamic selection mechanism ensures that superior backup nodes have low overall risk, and through flexible adjustment of the threshold, it ensures that the system can reserve sufficient backup nodes under various circumstances, improving the fault tolerance and fault recovery efficiency of the distributed system.

[0091] After identifying the dominant backup node, steps S202 to S204 can be executed to back up the backup data of each target backup node to the paired target dominant backup node. The target backup node is a backup node with a high failure rate at a specific time, while the target dominant backup node is a dominant backup node with a low failure rate at a specific time and is precisely matched to the corresponding target backup node. Therefore, even if the target backup node experiences unexpected failure or data corruption, the paired target dominant backup node can respond quickly and achieve rapid recovery with complete and real-time updated backup data. This not only effectively reduces the risk of data loss but also enhances the system's fault tolerance in complex failure scenarios, effectively improving the data security and reliability of the distributed system.

[0092] In some embodiments, the distributed system can also employ distributed transaction technology, using a two-phase commit protocol or distributed lock mechanism to ensure the atomicity and consistency of data operations among multiple nodes. The specific implementation steps are as follows: before data operations, initiate a distributed transaction; ensure that data operations on all nodes are synchronized using a two-phase commit protocol or distributed lock mechanism; if the transaction fails, roll back the data operations on all nodes. Data consistency can also be checked periodically, and if inconsistency is detected, immediately initiate data synchronization. The specific implementation steps are as follows: periodically check the data consistency of each node; if inconsistency is detected, synchronize data from other nodes; ensure that the data on all nodes remains consistent.

[0093] In some embodiments, incremental backup or hybrid backup strategies can be used to implement backups, thereby reducing the amount of backup data and improving backup efficiency. Simultaneously, by optimizing backup task scheduling, the time and bandwidth consumption during the backup process can be reduced.

[0094] The incremental backup strategy only backs up data that has changed since the last backup, significantly reducing the amount of data to be backed up. The specific implementation steps are as follows: Before each backup, check for data changes; only back up the data blocks that have changed; record the information of the backed-up data blocks for easy recovery later.

[0095] Hybrid backup combines the advantages of full and incremental backups, ensuring both data integrity and efficiency. The specific implementation steps are as follows: perform full backups regularly to ensure data integrity; perform incremental backups between full backups; optimize backup task scheduling to rationally arrange the execution order of full and incremental backups. The hybrid backup strategy combines the advantages of full and incremental backups, ensuring data integrity. It also improves overall system performance by optimizing backup task scheduling and allocating resources accordingly.

[0096] In some embodiments, please refer to Figure 6This application also provides a management device for a distributed system, the distributed system including multiple backup nodes. The management device for the distributed system includes: an acquisition module, a prediction module, a comparison module, and a determination module; wherein, The acquisition module is used to acquire the operating information, hardware information, and environmental information of the region where each backup node is located; The prediction module is used to determine the corresponding fault prediction model for each backup node based on the hardware information of each backup node, and input the operating information of each backup node into the corresponding fault prediction model to obtain the predicted failure rate of each backup node. The comparison module is used to compare the operating information, hardware information and environmental information of each backup node when a faulty backup node is detected, and to determine the similarity between each detected backup node and the faulty backup node based on the comparison results. The detected backup node is a backup node other than the faulty backup node. The determination module is used to determine the target backup node based on the similarity between each detected backup node and the faulty backup node, as well as the predicted failure rate of each detected backup node.

[0097] It should be noted that the distributed system management device provided in this application embodiment and the distributed system management method provided in this application embodiment are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned distributed system management method, and the repeated parts will not be described again.

[0098] In some embodiments, an electronic device provided in this application includes a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the above-described distributed system management method.

[0099] Specifically, the processor may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor, such as an application-specific integrated circuit (ASIC), etc. The processor may also include onboard memory for caching purposes. The processor may be a single processing unit or multiple processing units for performing different actions of the method flow according to embodiments of this application.

[0100] Memory can be any medium capable of containing, storing, transmitting, propagating, or transmitting instructions. For example, memory can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, instruments, or propagation media. Specific examples of memory include: magnetic storage devices such as magnetic tape or hard disk drives (HDDs); optical storage devices such as optical discs (CD-ROMs); random access memory (RAM) or flash memory; and / or wired / wireless communication links.

[0101] This application also provides a computer-readable medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned distributed system management method. This computer-readable medium may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into that device / apparatus / system. The aforementioned computer-readable medium carries one or more programs, which, when executed, implement the method as described in the embodiments of this application.

[0102] According to embodiments of this application, a computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wired, optical fiber, radio frequency signals, etc., or any suitable combination thereof.

[0103] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application. Therefore, the scope of this application should not be limited to the above embodiments, but should be defined not only by the appended claims, but also by their equivalents. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A management method of a distributed backup system, the distributed system comprising a plurality of backup nodes; characterized in that, The management method comprises: obtaining running information, hardware information and environment information of each backup node; determining a failure prediction model corresponding to each backup node according to the hardware information of each backup node, and inputting the running information of each backup node into the corresponding failure prediction model to obtain a predicted failure rate of each backup node; when a failure backup node is detected, comparing the running information, hardware information and environment information of the region of each backup node, and determining the similarity between each detected backup node and the failure backup node according to the comparison result, wherein the detected backup node is a backup node other than the failure backup node; determining a target backup node according to the similarity between each detected backup node and the failure backup node, and the predicted failure rate of each detected backup node.

2. The management method of a distributed backup system according to claim 1, characterized in that, The management method further comprises: determining a dominant backup node according to the similarity between each detected backup node and the failure backup node, and the predicted failure rate of each detected backup node, wherein the number of the dominant backup node is not less than the number of the target backup node; obtaining the available storage space of each dominant backup node, and the network delay and geographical distance of each target backup node; pairing the nodes based on the available storage space of each dominant backup node, and the network delay and geographical distance of each target backup node, to determine a target dominant backup node paired with each target backup node; backing up the backup data of each target backup node to the paired target dominant backup node.

3. The management method of a distributed backup system according to claim 1, characterized in that, The distributed system comprises a plurality of backup nodes of different levels, and the priority of a high-level backup node is higher than that of a low-level backup node. The management method further comprises: obtaining the priority of each shard data; matching the level of each backup node according to the priority of each shard data; determining the backup node corresponding to the storage of each shard data based on the matching information and the continuity order of each shard data, and storing each shard data in the corresponding backup node.

4. The management method of a distributed backup system according to claim 3, wherein, The management method further comprises: determining a plurality of dominant backup nodes according to the similarity between each detected backup node and the failure backup node, and the predicted failure rate of each detected backup node; obtaining the available storage space of each dominant backup node, and the network delay and geographical distance of each target backup node; determining the matching score of each dominant backup node relative to each target backup node based on the available storage space of each dominant backup node, and the network delay and geographical distance of each target backup node; determining a target dominant backup node paired with each target backup node according to the matching score of each dominant backup node relative to each target backup node and the level distance; backing up the backup data of each target backup node to the paired target dominant backup node.

5. The management method of a distributed backup system according to claim 2 or 4, characterized in that, The management method further comprises: after receiving a maintenance completion instruction corresponding to the target backup node, performing data recovery processing on the target backup node based on the backup data on the corresponding target dominant backup node; After a preset time period after the data recovery processing of the target backup node, if it is detected that the target backup node is in a normal operation state, the backup data belonging to the target backup node in the corresponding target backup node is deleted.

6. The management method of a distributed backup system according to claim 3, wherein, The management method further comprises: obtaining load monitoring data of the distributed system; determining resource requirements of each backup node according to the load monitoring data; allocating resources according to the resource requirements of each backup node and the grades of each backup node to preferentially supply resources to high-grade backup nodes.

7. The management method of a distributed backup system according to claim 2, wherein, The method for determining the target backup node according to the similarity of each detection backup node to the failed backup node and the predicted failure rate of each detection backup node comprises: determining a first evaluation score of each detection backup node based on the similarity of each detection backup node to the failed backup node, and determining a second evaluation score of each detection backup node based on the predicted failure rate of each detection backup node, wherein the first evaluation score is positively correlated with the similarity, and the second evaluation score is positively correlated with the predicted failure rate; performing weighted summation on the first evaluation score and the second evaluation score to obtain a target evaluation score; determining a detection backup node with a target evaluation score greater than a first preset threshold as the target backup node; determining the advantage backup node according to the similarity of each detection backup node to the failed backup node and the predicted failure rate of each detection backup node comprises: determining a detection backup node with a target evaluation score less than a second preset threshold as a candidate backup node, wherein the second preset threshold is less than the first preset threshold; when the number of candidate backup nodes is greater than or equal to the number of target backup nodes, determining each candidate backup node as the advantage backup node; when the number of candidate backup nodes is less than the number of target backup nodes, increasing the second preset threshold by a step value and returning to the step of determining a detection backup node with a target evaluation score less than a second preset threshold as a candidate backup node.

8. The management method of a distributed backup system according to claim 1, wherein, The failure prediction model determines a failure prediction model for each backup node according to the hardware information of each backup node, and inputs the running information of each backup node into the corresponding failure prediction model to obtain the predicted failure rate of each backup node, and specifically comprises: According to the hardware information of the backup node, the aging law and failure mode of different hardware components are analyzed, the parameters contained in the aging law and failure mode of different hardware components are used as key parameters for predicting disk failure, and through the collection of historical hardware failure data, combined with statistical methods and machine learning algorithms, personalized failure prediction models for different hardware configurations are constructed.

9. The management method of a distributed backup system according to claim 8, wherein, The statistical method adopts a survival analysis method, the machine learning algorithm adopts a random forest algorithm, and the specific process of constructing the personalized failure prediction model for different hardware configurations comprises: Survival analysis method index definition: Survival function : probability that hardware does not fail at time , i.e. , where is the time to failure; Risk function The hardware has been running normally for the specified time. Under the condition of "", the probability of a failure occurring instantaneously in the next instant is, i.e. ; Kaplan-Meier survival function estimate: for a sample ordered by failure times, where is the time of the th failure, and is the number of hardware units still operational at time , i.e. the "risk set size", is the number of hardware units that failed at time , then the estimate of the survival function is given by: ; The input features of the random forest algorithm include hardware static configuration, dynamic running state, environmental parameters, and the "risk index" output by the survival analysis; Definition of Random Forest includes a plurality of decision trees, wherein each of the trees is trained on a different subset of the training data a plurality of decision trees, wherein each of the trees is trained on a different subset of the training data a plurality of decision trees, wherein each of the trees is trained on a different subset of the training data a plurality of decision trees, wherein each of the trees is trained on a different subset of the training data a plurality of decision trees, wherein each of the trees is trained on a different subset of the training data ; Combining the "time trend" of survival analysis with the "feature association" of random forest, the output is the individualized failure probability for hardware : ; wherein, is a hardware configuration group, is a conditional survival function for the group; is a fusion weight, which dynamically adjusts according to the sample size of the configuration group: the larger the sample size, the larger the fusion weight, which depends on the statistical law of survival analysis; when the sample size is small, the fusion weight is reduced, which depends on the feature learning of random forest; By configuring the conditional survival function and fusion weight adjustment of the group, the model can make differentiated predictions for different hardware output differences, avoiding the prediction deviation of the "universal model" for small configuration.

10. The management method of a distributed backup system according to claim 8, wherein, The hardware information of the backup nodes includes CPU model and core number, memory capacity and type, disk read-write performance parameters, and network interface bandwidth, and the parameters contained in the aging law and failure mode of different hardware components include the mean time between failures (MTBF) of a mechanical hard disk (HDD), power-on time, and read-write frequency indicators.

11. The management method of a distributed backup system according to claim 1, wherein, In the step of comparing the running information, hardware information, and environmental information of the region of each backup node when a fault backup node is detected, and determining the similarity of each detected backup node and the fault backup node according to the comparison result, the calculation formula of the similarity is: ; in, For the first The similarity score between each detection backup node and the failed backup node. , , The weighting coefficients for operational information, hardware information, and environmental information are respectively, satisfying... , , , The first The similarity values ​​of each backup node in terms of runtime information, hardware information, and environmental information; The running information contains multiple indicators, including CPU usage, memory occupancy, disk I / O rate, etc., and the running information similarity Sim_r (i) is calculated using the cosine similarity: ; Wherein: R_i=[R_i1,R_i2,…,R_in] is the running indicator vector of the detected backup node i, R_f=[R_f1,R_f2,…,R_fn] is the running indicator vector of the fault backup node, · represents vector dot product, and ||·|| represents vector norm. The hardware information contains discrete and continuous mixed indicators, and the hardware information similarity The weighted combination method can be used for calculation: ; wherein, is the weight of the hardware indicator, is the similarity of the hardware indicator, is the similarity of the hardware indicator, ; wherein, and are index values for detecting a backup node and a failed backup node, respectively, is a scaling factor; Discrete indicators: ; The environmental information includes geographical position, machine room parameters, and environmental information similarity The distance function is used for calculation: ; wherein, The integrated distance, which is an environmental parameter, can be defined as: ; wherein, and is an environmental parameter value, is a parameter standard deviation for normalization; is a distance scaling factor, controlling the similarity decay speed.

12. The management method of a distributed backup system according to claim 2, wherein, The node pairing is performed based on the available storage space of each dominant backup node and the network delay and geographic distance of each target backup node, and the target dominant backup node paired with each target backup node is determined, which specifically includes: The matching score is used to evaluate each dominant backup node. For the currently matched target backup node, the available storage space, network delay, and geographic distance are evaluated. The available storage space is positively correlated with the matching score, the network delay is inversely correlated with the matching score, and the geographic distance is set to multiple distance intervals. The matching score corresponding to the interval in the middle of the geographic distance is the highest, and the matching score decreases as the distance increases or decreases. Based on the above matching rules, a corresponding matching formula is constructed, and the total matching score formula is: ; wherein, is the matching score of the i-th backup node, is the matching score of the j-th backup node, , , are weight coefficients of storage space, network delay, and geographical distance, respectively, satisfying , , , are scoring functions of the corresponding dimensions. The available storage space score formula is: ; wherein, the available storage space of the node the available storage space of the node the maximum available storage space among all candidate nodes; The network delay score calculation formula is: ; wherein, is a node a network latency to a target node, is a scaling parameter, controlling the decay rate of the latency to the score; The geographic distance score calculation formula is: ; wherein: , is a geographical distance interval boundary point, ; the parameters satisfy: (, , .

13. A management apparatus of a management method using the distributed backup system according to any one of claims 1 to 7, characterized by, The management device includes: An acquisition module is configured to acquire the running information, hardware information, and environmental information of the region of each backup node. A prediction module is configured to determine a fault prediction model corresponding to each backup node based on the hardware information of each backup node, and input the running information of each backup node into the corresponding fault prediction model to obtain the predicted failure rate of each backup node. A comparison module is configured to compare the running information, hardware information, and environmental information of the region of each backup node when a fault backup node is detected, and determine the similarity of each detected backup node and the fault backup node based on the comparison result, wherein the detected backup node is a backup node other than the fault backup node. A determination module is configured to determine a target backup node based on the similarity of each detected backup node and the fault backup node, and the predicted failure rate of each detected backup node.

14. An electronic device using a management method of the distributed backup system according to any one of claims 1 to 7, characterized by, The application relates to a distributed backup system, comprising a processor and a memory; the memory stores a computer program, wherein the computer program realizes a management method of the distributed backup system when executed by the processor.

15. A storage medium for use with the distributed backup system of any of claims 1-7, wherein, A storage medium stores a computer program, wherein the computer program realizes a management method of the distributed backup system when executed by a processor.