A network self-healing device based on node state data
By introducing a self-healing device based on node status data into the data center network, and combining multiple algorithm models for real-time monitoring and prediction, the problems of high communication latency, low accuracy of status data, and uneven resource allocation are solved, enabling rapid fault recovery and efficient resource management of network communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies suffer from high communication latency, low accuracy of monitoring status data, limited path selection, and uneven resource allocation in data center networks, resulting in insufficient network communication efficiency and stability.
A network self-healing device based on node state data is adopted, which combines extended Kalman filter algorithm, ant colony algorithm and simulated annealing algorithm to monitor network status in real time, perform intelligent path selection and resource allocation, predict potential faults through ARIMA model and take preventive measures before faults occur.
It significantly improves the stability and self-healing capabilities of network communication, enables rapid fault recovery and intelligent resource scheduling, and enhances network communication efficiency and service continuity.
Smart Images

Figure CN120434108B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network communication technology, specifically a network self-healing device based on node status data. It aims to improve the stability and self-healing capability of network communication by real-time monitoring of network status, prediction and detection of potential faults by a status observer, and taking preventive measures for intelligent decision-making path selection and resource allocation. Background Technology
[0002] With the continuous advancement of network technology, the stability and reliability of network communication have become core requirements for various application scenarios. In practical applications, networks face various sudden failures, such as equipment failure and link interruption. Firstly, relying on simple fault detection and recovery mechanisms, such as manual experience-based link-by-link checks and command-line operations, often results in slow response times and difficulty in quickly restoring network communication, severely impacting communication quality and user experience. Secondly, there is a lack of in-depth observation and intelligent analysis capabilities of network status. Network status includes multiple dimensions such as traffic load, link quality, and equipment health, which are crucial for predicting and detecting potential faults. Existing technologies do not comprehensively observe and utilize this status information, making it impossible to take effective preventative measures before network failures occur. Thirdly, path selection and resource allocation are key factors affecting network communication efficiency and stability. Traditional path selection algorithms often struggle to find the optimal path globally, easily getting trapped in local optima, leading to uneven distribution of network resources and reduced communication efficiency. Summary of the Invention
[0003] This invention aims to address key issues in current data center network self-healing technologies, such as high communication latency, low accuracy of monitoring status data, limited path selection, and uneven resource allocation. By designing a network self-healing device and introducing advanced algorithm models and status observation mechanisms, it achieves real-time and efficient communication assurance while improving the network's ability to predict and respond to potential faults.
[0004] The technical solution adopted in this invention is as follows:
[0005] A network self-healing device based on node state data includes a state monitoring module, a state estimation and filtering module, a fault prediction and prevention module, a fault recovery module, and a resource reallocation module. It also includes a device operation status detection module, which is connected to other modules via a communication network to collect multi-dimensional operation status data in real time, including the runtime of the fault recovery module. Data volume The number of iterations of the optimization algorithm called during path planning Network status data and the residuals of the state estimation and filtering modules By combining health assessment models, a closed-loop management system can be achieved from data collection to optimization.
[0006] The device operation status monitoring module collects multi-dimensional operation status data in real time, including: the EKF filtering error of the status estimation and filtering module, the path planning iteration count of the fault recovery module, and the network fluctuation status of the fault prediction and prevention module. The entropy value is calculated by normalizing the above indicators. ,
[0007]
[0008] in It is a prediction correction factor introduced by the fault prediction and prevention module, which dynamically adjusts the entropy value calculation based on the output of the ARIMA model. The frequency of network topology changes in the fault recovery module is controlled by calculating the number of network topology changes in the past hour to determine the number of iterations of the ant colony algorithm.
[0009] Calculated indicator weights To achieve high failure probability predicted by ARIMA, the weights are adjusted accordingly. Reduce and suppress high-risk modules in advance; and when the filter residual In case of anomalies, a two-stage filtering enhancement mode is activated to dynamically adjust the covariance of the observation noise. This is used to handle filtering faults, where higher weights result in more sensitive filtering.
[0010] Furthermore, the device operation status detection module collects the core indicators of the status estimation and filtering module in real time, including EKF filtering error and residual. and observation noise covariance ;
[0011] After normalization, and in calculating the entropy value When introducing a filter residual correction term Dynamically adjust weights When residuals are abnormal, a two-stage filtering enhancement mode is triggered to dynamically correct the observation noise covariance. This enables adaptive adjustment of filter sensitivity; furthermore, it is based on the average filter error over a historical 24-hour period. and standard deviation Set dynamic thresholds and use health scores ,in, The normalized indicator triggers a tiered response.
[0012] If the filtering error drops to a moderately abnormal level due to exceeding the standard, a forced switch to the backup filtering model is initiated. At the same time, an additional computing resource is requested from the resource reallocation module, and the filtered and optimized data is fed back to the fault prediction and prevention module to improve the accuracy of the ARIMA model input data and assist the ARIMA model in correcting the prediction factor λ, forming a positive cycle of "improved filtering accuracy → enhanced prediction reliability".
[0013] Furthermore, regarding the fault prediction and prevention module, the device operation status detection module dynamically corrects the entropy calculation based on the predicted fault probability λ output by the ARIMA model. As λ increases, the entropy value increases, the weight decreases, and the resource consumption of the fault prediction and prevention module is suppressed.
[0014] Network fluctuation data for fault prediction and prevention module A sliding window approach is used, leveraging the TSDS (Tracking Data Sheets) generated by the fault prediction and prevention module itself (to analyze the differences in short-term and long-term trends against a historical benchmark); if short-term fluctuations... If the threshold is exceeded, the ARIMA model parameters are updated online, and the difference order is adjusted; if the health score When a severe anomaly occurs, the fault prediction and prevention module isolates the abnormal node and activates the backup path planning instance; simultaneously, the isolated event is fed back to the status monitoring module to update the network topology data. .
[0015] Furthermore, in the fault recovery module, the device operation status detection module monitors the convergence of the ant colony algorithm in real time and continuously tracks the number of iterations of the ant colony algorithm. The ratio of the frequency of network topology changes to the network topology change frequency; when this ratio is greater than a threshold... When the network becomes too dynamic, the algorithm immediately triggers a two-level response mechanism; the first-level response generates an algorithm failure alarm signal. The accompanying diagnostic conclusion was "an abnormal surge in the number of ant colony algorithm iterations," confirming the root cause as "frequent changes in network topology leading to difficulty in algorithm convergence." The second-level response sent optimization instructions to the resource reallocation module. For example, the algorithm switching command is: "Enable simulated annealing algorithm" and increase the corresponding module's computing resources.
[0016] Furthermore, the weight of the device operation status detection module Guiding the decision-making of the resource reallocation module;
[0017] When the weights of the state estimation and filtering modules increase, more CPU resources are allocated to improve data processing speed. When the weights of the prediction module decrease due to the increase of λ, its memory usage is reduced to avoid resource waste, and resource utilization information is fed back to the state monitoring module in real time. This achieves a self-healing closed-loop process: network topology mutation → ant colony algorithm iteration surge → health score decline → triggering simulated annealing algorithm switch → resource reallocation → new path data feedback to the state monitoring module → updating ARIMA model input → improving prediction accuracy.
[0018] Compared with the prior art, the effective gains achieved by this invention are as follows:
[0019] (1) The original data is denoised by a two-stage EKF algorithm, which greatly improves the accuracy and reliability of the state data.
[0020] (2) Construct relevant time series datasets based on network status data continuously collected by the monitoring module, and use them as inputs to the ARIMA model for modeling and predicting network status, so as to predict potential faults.
[0021] (3) By combining the ant colony algorithm and the simulated annealing algorithm, the optimal path selection and resource allocation of the global and local systems were realized, the network fluctuation faults were quickly recovered, and the network communication efficiency was improved.
[0022] (4) By combining the state observation mechanism with the intelligent algorithm model, the network state changes can be predicted in real time, potential faults can be detected and prevented in a timely manner, and the network's self-healing ability and service continuity can be significantly improved.
[0023] (5) Through dynamic weight allocation and closed-loop feedback mechanisms, intelligent scheduling of system resources and real-time optimization of module health status are achieved. The device operation status detection module can accurately assess the health status of each module by collecting multi-dimensional operation status data in real time, combined with entropy calculation and dynamic weight allocation. When an anomaly is detected, the module triggers a hierarchical response mechanism and dynamically optimizes the allocation of resources such as CPU and memory through collaboration with the resource reallocation module. Finally, a closed-loop self-healing process of "monitoring → analysis → optimization → feedback" is formed, which significantly improves the overall stability and resource utilization efficiency of the device. Attached Figure Description
[0024] Figure 1 This is a diagram of a network self-healing system device according to an embodiment of the present invention.
[0025] Figure 2 This is a structural diagram of the status monitoring module in an embodiment of the present invention.
[0026] Figure 3 This is a flowchart illustrating the process of the fault prediction and prevention module in an embodiment of the present invention.
[0027] Figure 4 This is a flowchart illustrating the fault recovery module processing procedure in an embodiment of the present invention.
[0028] Figure 5 This is a structural diagram of the device operation status detection module. Detailed Implementation
[0029] The following is in conjunction with the appendix Figure 1-5 The present invention will be further explained and described below.
[0030] This invention employs a state-detection-based network self-healing mechanism, primarily aiming to provide a novel automation solution for network control systems. It integrates Extended Kalman Filter (EKF), ant colony optimization, and simulated annealing algorithms to achieve efficient denoising of raw network state data, potential fault prediction, intelligent path selection, and globally optimized resource allocation. This solution can monitor network state in real time, design a network self-healing device capable of predicting and detecting potential faults, and execute intelligent decision-making for path selection and resource allocation, thereby improving the stability and self-healing capabilities of network communication. It not only significantly enhances the stability and self-healing capabilities of network communication but is also particularly suitable for complex, large-scale network environments, such as data center networks.
[0031] Reference Figure 1 A network self-healing device based on node state data, comprising:
[0032] The status monitoring module is used to capture the status information of the network and each device node in real time, including the link quality, traffic dynamics and device load information in the network, and transmit the collected device status data to the status estimation and filtering module. At the same time, the resource utilization information and network status information transmitted by the resource reallocation module are also transmitted to the fault prediction and prevention module.
[0033] The state estimation and filtering module is used to receive equipment state data, use the extended Kalman filter algorithm to finely process the equipment state data, obtain filtered equipment state data, and pass it to the fault prediction and prevention module and the fault recovery module.
[0034] The fault prediction and prevention module receives filtered equipment status data from the status estimation and filtering module and network status data from the status monitoring module. For equipment status data, if a fault is detected in the current equipment, a relevant equipment control strategy is generated. For network status data, if the fault data is detected, the fault data is analyzed and diagnosed in conjunction with resource utilization information transmitted by the status monitoring module and historical fault data of network nodes to construct a network time series dataset. A control strategy is generated based on the local control strategy library. If the status data is detected as normal, the existing ARIMA prediction model is used to predict nodes that may fail within a set time period in the future, and a corresponding control strategy is generated and transmitted to the fault recovery module.
[0035] The fault recovery module receives control strategies generated by the fault prediction and prevention module. For network fluctuation faults, it selects different algorithms for path planning based on the complexity of the current network. For equipment faults, it temporarily isolates the relevant equipment and generates equipment alarm signals. The isolation status is lifted after the relevant equipment returns to normal. After taking corresponding measures, it determines whether the measures are effective. If effective, the fault recovery result is passed to the resource reallocation module. If it fails, relevant alarm information is generated.
[0036] The device operation status detection module is used to monitor the operation status of each module in the network self-healing device in real time. By analyzing key performance indicators, it determines whether the module is in normal operation and ensures the normal operation of the entire system.
[0037] The resource reallocation module is used to reallocate network resources based on the fault recovery status and the current network node resource usage, and to feed back the resource utilization information to the status monitoring module in real time, forming a closed-loop optimization mechanism.
[0038] Furthermore, the state estimation and filtering module employs a two-stage Kalman filter model to perform fine processing on the received device state data, including:
[0039] In the first stage, the extended Kalman filter (EKF) is used to conduct a preliminary assessment of the equipment status and obtain the original equipment status data.
[0040] In the second stage, based on the original equipment status data from the first stage, standard Kalman filtering is used for further fine-tuning to achieve data preprocessing, wave-controlled denoising, and normalization, thereby obtaining effective features.
[0041] Furthermore, the specific process of the fault prediction and prevention module is as follows:
[0042] The fault prediction and prevention module first determines whether the data is device status data or network status data. If it is determined to be device status data and fault data, relevant device control strategies are generated for device isolation and recovery. If it is determined to be network status data, it is processed according to whether a fault has occurred. If it is fault data, key fault information is extracted and combined with resource utilization information transmitted by the status monitoring module, which is then fed into the fault judgment vector machine to diagnose the cause of the fault and query the local preset control strategy library to generate a control strategy. The resource utilization information is used to help diagnose environmental anomalies. If it is determined to be normal status data, a time series dataset is constructed based on the data packet transmission status and network latency in the status data. This dataset is used as input to the ARIMA prediction model for prediction. Based on the prediction results, the trend of network fluctuations and potential fault points within a set future time period are determined, and corresponding control strategies are generated.
[0043] Furthermore, when performing path planning, the fault recovery module adopts different path planning strategies based on the complexity of different networks: for networks with fewer than a set number of nodes, fixed network topology, and simple and predictable traffic patterns, the ant colony algorithm is used to dynamically plan the optimal or suboptimal communication path; for networks with more than a set number of nodes, dynamically changing network topology, and complex and unpredictable traffic patterns, the simulated annealing algorithm is enabled to accept non-optimal paths based on probability.
[0044] Furthermore, the device operation status detection module collects real-time operation status data from each module, such as response time, data processing volume, and algorithm execution efficiency. It analyzes the collected status data to determine whether each module is operating normally. If an anomaly is detected, it generates corresponding warning information and performance optimization suggestions for the resource reallocation module to enable device self-adjustment.
[0045] In the network control system described in this invention, several key components are deployed for status monitoring, such as... Figure 2 As shown, the condition monitoring module includes sensors (or data acquisition units), controllers, actuators, and the communication network between them. The specific process is as follows:
[0046] Sensors capture real-time status information of the network and each device node (controlled device), such as link quality, traffic dynamics, device load, and physical parameters such as temperature and pressure.
[0047] State monitoring modules deployed in sensor networks periodically collect data from sensors using a controller to construct state vector data. The collected raw device status data is transmitted via a communication network. The resource utilization information and network state data transmitted by the resource reallocation module are passed to the state estimation and filtering module. It is passed to the fault prediction and prevention module.
[0048] As the core processing unit in the device, the state estimation and filtering module employs a two-stage Kalman filter model to finely process the received equipment state data. Combining the state transition function and observation function defined by the system dynamic model, and the Kalman gain calculated based on observation noise and process noise, it performs a preliminary estimate of the current state, obtaining filtered state information. This information is then passed to the fault prediction and prevention module and the fault recovery module.
[0049] The embodiments of the present invention employ a two-stage Kalman filter model to perform fine processing on the received data, specifically as follows:
[0050] Construct a two-stage model: In the first stage, use EKF to conduct a preliminary assessment of the system state, expressed by the formula:
[0051]
[0052] In the formula, Estimate the current state. This is the state transition function. For Kalman gain, These are observations (from the sensor). For observation functions (mapping the system state to the observation space);
[0053] In the second stage, based on the estimation in the first stage, standard Kalman filtering is used for further fine-tuning to perform preprocessing steps such as denoising and normalization of the data in order to obtain effective features.
[0054] like Figure 3 As shown, the fault prediction and prevention module receives data from the state estimation and filtering module. and network status data from the status monitoring module Next, it will determine whether the data is faulty. If it is faulty, it will analyze the status data, extract key fault information, and input it into the fault judgment vector machine to diagnose the cause of the fault. It will also query the local preset control strategy library (dictionary) StD to generate a control strategy. If the data is from a normal state, a time series dataset (TSDS) is constructed based on the data packet transmission status and network latency in the state data. This dataset is used as input to the ARIMA model for prediction, and the prediction results are used to determine the future time period. Identify trends and potential failure points in the internal network fluctuations and generate control strategies. Examples of control policies include dynamically adjusting data packet sending rates, optimizing routing, caching critical data, and backing up communication paths. Send to the fault recovery module.
[0055] The ARIMA model is used for time series forecasting. This model mainly consists of three parts: autoregressive (AR), differencing (I), and moving average (MA). By integrating different components (such as trend and residuals) from the time series dataset TSDS, the non-stationary time series is transformed into a stationary time series, and then forecasting is performed using the autoregressive and moving average models. Specifically, the autoregressive and moving average models are used to obtain the linear relationship between the current value and historical values and historical error terms.
[0056] like Figure 4 As shown, after the fault recovery module obtains the control strategy, it will first determine the control strategy. If the fault is a device malfunction, the device will be temporarily isolated and an alarm signal will be generated. Once the device is back to normal, the isolation will be lifted. After appropriate solutions are implemented, the fault recovery module will determine whether the solutions are effective. If effective, the fault recovery result will be sent to the resource reallocation module; otherwise, relevant alarm information will be generated. If the failure is due to network fluctuations, different algorithms are flexibly selected for path planning based on the complexity of the current network. In simple networks, the ant colony algorithm is used to dynamically plan the optimal or suboptimal communication path; while in complex and large networks, the simulated annealing algorithm is used to enhance the global search capability by probabilistically accepting non-optimal paths.
[0057] When performing path planning, the fault recovery module can adopt different path planning strategies based on the complexity of different networks. For simple networks, i.e., networks with a small number of nodes (e.g., tens to hundreds of nodes), relatively fixed network topology, and relatively simple and predictable traffic patterns, the ant colony algorithm is used to dynamically plan the optimal or suboptimal communication path. For complex and large networks, i.e., networks with a large number of nodes (e.g., tens of thousands of nodes), dynamically changing network topology, and complex and unpredictable traffic patterns, the simulated annealing algorithm is enabled. By accepting non-optimal paths with probability, the global search capability is enhanced, avoiding getting trapped in local optima. In addition, to improve response speed, a parallel construction method is used during the network path query process.
[0058] The ant colony algorithm is used to dynamically plan the optimal or suboptimal communication path. The specific process is as follows:
[0059] First, initialize the pheromone on each edge of the network. ,in It is a constant. This represents the initial amount of pheromone; assuming the ant... On the side Move upwards, and the pheromones are updated after the move. ,in It is the pheromone volatility coefficient, representing the proportion of pheromones that evaporate over time. It's the number of ants. For ants Within a time period on the edge The amount of pheromones left on the surface; and the ant chose to be on the edge. The probability of moving upward is ,in For the edge pheromone concentration Power of 1 For the edge Heuristic information on (reciprocal of side length) Power of 1 For all the Only the edges that the ant is allowed to choose are summed, here This represents the set of nodes that the ant can move to in its current state. For the edge pheromone concentration Power of 1 For the edge Heuristic information on The power of, where These are algorithm parameters used to control the relative importance of pheromones and heuristic information; then a path record table is used. Recording ants The nodes currently visited, after At each point, all ants have completed one cycle; finally, the path length of each ant is calculated, the shortest path length is saved, and the process jumps back to the pheromone initialization stage at the beginning of the model to prepare for the next cycle.
[0060] However, the ant colony algorithm can get stuck in local optima when dealing with complex and large networks. Therefore, to improve performance in complex networks, a scheme combining the ant colony algorithm and simulated annealing was designed. For the current solution... And new interpretation ,if Then accept This is the current solution. Otherwise, with probability... ,accept .in It represents the weight difference between the new solution and the current solution. This means that a probabilistic approach of accepting non-optimal paths is used to enhance global search capabilities, and a parallel construction method is employed to reduce the time consumed during path lookup.
[0061] When performing path planning, simulated annealing (SA) is used to enhance global search capabilities by probabilistically accepting non-optimal paths, thus addressing the local overfitting problem in ant colony dynamic programming. Its formula is expressed as:
[0062]
[0063] in Due to differences in path quality, The current temperature is [value]. In addition, to reduce path construction time, this device employs a parallel construction method to find paths.
[0064] like Figure 5 As shown, the device operation status detection module is connected to each module through a communication network to collect multi-dimensional operation status data in real time (such as the running time of the fault recovery module). Data volume The number of iterations of the optimization algorithm called during path planning Network status data and the residuals of the state estimation and filtering modules By combining health assessment models, a closed-loop management system can be achieved from data collection to optimization.
[0065] This module uses dynamic weight allocation and a sliding window threshold to determine the module's health status in real time. The specific implementation is as follows: First, the module collects core indicators from other modules: the EKF filtering error of the state estimation and filtering module, the path planning iteration count of the fault recovery module, and the network fluctuation status of the fault estimation and prevention module, etc. The entropy value is then calculated by normalizing the above indicator information. ,in It is a prediction correction factor introduced by the fault prediction and prevention module, which dynamically adjusts the entropy value calculation based on the output of the ARIMA model. To control the frequency of network topology changes in the fault recovery module, the number of network topology changes within the past hour is calculated to determine the iteration count of the ant colony algorithm. Then, the indicator weights are calculated. Ultimately, this will be achieved when ARIMA predicts a high probability of failure ( Increase) -> Entropy Increase -> Weight Reduce and suppress high-risk modules in advance; and when the filter residual In case of anomalies, a two-stage filtering enhancement mode is activated to dynamically adjust the covariance of the observed noise. This is used to handle filtering faults, where higher weights result in more sensitive filtering.
[0066] Meanwhile, this module employs a sliding window-based dynamic threshold adjustment technique to analyze historical operational data in real time (such as the average of indicators over the last 24 hours). and standard deviation Adaptive thresholds were established for each monitoring indicator, and health scores were used. (in (For normalized indicators) triggers a tiered response. When When a mild anomaly is detected, a suggestive alert is generated. Fine-tune parameters (such as Kalman gain); when When in a moderately abnormal state, force a switch to the path optimization algorithm and generate a resource reallocation instruction; when In the event of a severe anomaly, the isolation module is activated and a backup node is engaged. During this process, the device operation status detection module can directly generate executable instructions and form a "monitoring-analysis-decision" closed loop through real-time interaction with the resource reallocation module and the status monitoring module. Taking path planning algorithm monitoring as an example: the module continuously tracks the number of iterations of the ant colony algorithm. The ratio of the frequency of network topology changes to the network topology change frequency; when this ratio is greater than a threshold... When the network becomes too dynamic, the algorithm immediately triggers a two-level response mechanism. The first-level response generates an algorithm failure alarm signal. The accompanying diagnostic conclusion was "an abnormal surge in the number of iterations of the ant colony algorithm," and related indicators were analyzed using fuzzy logic (such as the simultaneous detection of...). (Too large), confirming the root cause as "frequent changes in network topology cause difficulty in algorithm convergence." The second-level response sends optimization instructions to the resource reallocation module. For example, the algorithm switching command is: "Enable simulated annealing algorithm" and increase the corresponding module's computing resources.
[0067] The resource reallocation module reallocates corresponding computing resources to the relevant modules based on alarm information from the operation monitoring module, and generates corresponding allocation results recorded in the device status log. In addition, this module also reallocates network resources (such as bandwidth, processor time, etc.) based on fault recovery status and current network node resource usage, and feeds back resource utilization information to the status monitoring module in real time, forming a closed-loop optimization mechanism.
[0068] This module uses dynamic weight allocation and a sliding window threshold to determine the module's health status in real time. The specific implementation method is as follows:
[0069] The operational status monitoring module collects key metrics from the status estimation and filtering modules in real time, including EKF filtering error and residual. Observation noise covariance And so on. The entropy value is calculated through normalization. When introducing a filter residual correction term Dynamically adjust weights When residuals are abnormal, a two-stage filtering enhancement mode is triggered to dynamically correct the observation noise covariance. This enables adaptive adjustment of the filter sensitivity. Furthermore, it also considers the average filter error over a historical 24-hour period. and standard deviation Set dynamic thresholds. If the health score... (in If the normalized index drops to a moderately abnormal level due to excessive filtering error, it will be forcibly switched to the backup filtering model. At the same time, it will request additional computing resources from the resource reallocation module and feed back the filtered and optimized data to the fault prediction module to improve the accuracy of the ARIMA model input data and assist the ARIMA model in correcting the prediction factor λ, forming a positive cycle of "improved filtering accuracy → enhanced prediction reliability".
[0070] For the fault prediction and prevention module, the operation status detection module dynamically corrects the entropy value calculation based on the predicted fault probability (reflected as λ value) output by the ARIMA model. As λ increases (higher predicted fault risk), the entropy increases and the weight decreases, thereby suppressing the module's resource consumption and preventing fault prediction from failing due to data overload. Furthermore, network fluctuation data for the fault prediction module is also considered. A sliding window approach is used, leveraging the TSDS (Time Series Dataset) generated by the fault prediction module itself as a historical benchmark to analyze the differences in short-term (e.g., 5-minute) and long-term (24-hour) trends. If short-term fluctuations... If the threshold is exceeded, an online update of the ARIMA model parameters is triggered, adjusting the difference order. (If the health score...) When a severe anomaly occurs, the fault prediction and prevention module isolates the abnormal node and activates the backup path planning instance. Simultaneously, the isolated event is reported back to the status monitoring module to update the network topology data. .
[0071] In the fault recovery module, the operation status detection module monitors the convergence of the ant colony algorithm in real time and continuously tracks the number of iterations of the ant colony algorithm. The ratio of the frequency of network topology changes to the network topology change frequency; when this ratio is greater than a threshold... When the network becomes too dynamic, the algorithm immediately triggers a two-level response mechanism. The first-level response generates an algorithm failure alarm signal. The accompanying diagnostic conclusion was "an abnormal surge in the number of ant colony algorithm iterations," confirming the root cause as "frequent changes in network topology leading to difficulty in algorithm convergence." The second-level response sent optimization instructions to the resource reallocation module. For example, the algorithm switching command is: "Enable simulated annealing algorithm" and increase the corresponding module's computing resources.
[0072] In addition, the weight of the running status detection module This directly guides the resource reallocation module's decisions. For example, when the weight of the state estimation and filtering module increases, more CPU resources are allocated to improve data processing speed; when the weight of the prediction module decreases due to an increase in λ, its memory usage is reduced to avoid resource waste, and resource utilization information is fed back to the state monitoring module in real time. Ultimately, this achieves network topology mutation (…). ↑) → Ant colony algorithm iterations surge ( (↑) → Health score declines → Triggers simulated annealing algorithm switch → Resource reallocation → New path data is fed back to the status monitoring module → Updates ARIMA model input → Improves prediction accuracy—this is a closed-loop process from local optimization to global self-healing. In this process, all modules share unified status data, and the consistency of decision-making is ensured by integrating the entropy calculation from the running status monitoring module with the health score. Using a sliding window and dynamic weights, global adaptive adjustment is achieved from filtering and denoising to fault prediction and path selection.
Claims
1. A network self-healing device based on node state data, comprising a state monitoring module, a state estimation and filtering module, a fault prediction and prevention module, a fault recovery module, and a resource reallocation module, characterized in that, It also includes a device operation status detection module, which is connected to other modules via a communication network to collect multi-dimensional operation status data in real time, including the running time of the fault recovery module. Data volume The number of iterations of the optimization algorithm called during path planning Network status data and the state estimation and filtering residuals of the filtering module By combining health assessment models, a closed-loop management system can be achieved from data collection to optimization. The device operation status monitoring module collects multi-dimensional operation status data in real time as indicators, including: the EKF filtering error of the state estimation and filtering module, the number of path planning iterations of the fault recovery module, and the network fluctuation status of the fault prediction and prevention module. The entropy value is calculated by normalizing the above indicators. , ; in It is a prediction correction factor introduced by the fault prediction and prevention module, which dynamically adjusts the entropy value calculation based on the output of the ARIMA model. e represents the frequency of network topology changes in the fault recovery module. i The original information entropy value is used to control the number of iterations of the ant colony algorithm by calculating the number of network topology changes in the past hour. Calculated indicator weights To achieve high failure probability predicted by ARIMA, the weights are adjusted accordingly. Reduce and suppress high-risk modules in advance; and when the filter residual In case of anomalies, a two-stage filtering enhancement mode is activated to dynamically adjust the covariance of the observation noise. This is used to handle filtering faults, where R is the covariance of the original observed noise, and the higher the weight, the more sensitive the filtering.
2. The network self-healing device based on node state data according to claim 1, characterized in that, The device operation status detection module collects core indicators from the status estimation and filtering module in real time, including EKF filtering error and filtering residual. Covariance of original observation noise ; After normalization, and in calculating the entropy value When introducing a filter residual correction term Dynamically adjust weights ; When the filter residual is abnormal, a two-stage filter enhancement mode is triggered to dynamically correct the observation noise covariance. This enables adaptive adjustment of filter sensitivity; Based on the historical 24-hour average filtering error and standard deviation Set dynamic thresholds and use health scores ,in, The normalized indicator triggers a tiered response. If the filtering error drops to a moderately abnormal level due to exceeding the standard, a forced switch to the backup filtering model is initiated. At the same time, an additional computing resource is requested from the resource reallocation module, and the filtered and optimized data is fed back to the fault prediction and prevention module to improve the accuracy of the ARIMA model input data. This assists the ARIMA model in correcting the prediction correction factor, forming a positive cycle of "improved filtering accuracy → enhanced prediction reliability".
3. A network self-healing device based on node state data according to claim 2, characterized in that, For the fault prediction and prevention module, the device operation status detection module dynamically corrects the entropy calculation using the prediction correction factor λ output by the ARIMA model. As λ increases, the entropy value increases, the weight decreases, and the resource consumption of the fault prediction and prevention module is suppressed. Network fluctuation data for fault prediction and prevention module Using a sliding window approach, the Time Series Dataset (TSDS) generated by the fault prediction and prevention module itself is used as a historical benchmark to analyze the differences between its short-term and long-term trends; if short-term fluctuations... If the threshold is exceeded, the ARIMA model parameters are updated online, and the difference order is adjusted; if the health score When a severe anomaly occurs, the fault prediction and prevention module isolates the abnormal node and activates the backup path planning instance; simultaneously, the isolated event is fed back to the status monitoring module to update the network topology data. .
4. A network self-healing device based on node state data according to claim 3, characterized in that, In the fault recovery module, the device operation status detection module monitors the convergence of the ant colony algorithm in real time and continuously tracks the number of iterations of the ant colony algorithm. The ratio of the frequency of network topology changes to the network topology change frequency; when this ratio is greater than a threshold... When the network becomes too dynamic, the algorithm immediately triggers a two-level response mechanism. First-level response generation algorithm failure alarm signal The accompanying diagnostic conclusion was "an abnormal surge in the number of ant colony algorithm iterations," confirming the root cause as "frequent changes in network topology leading to difficulty in algorithm convergence." The second-level response sent optimization instructions to the resource reallocation module. .
5. A network self-healing device based on node state data according to claim 4, characterized in that, Weight of the device operation status detection module Guiding the decision-making of the resource reallocation module; When the weight of the state estimation and filtering module increases, more CPU resources are allocated to improve data processing speed. When the weight of the fault prediction and prevention module decreases due to the increase of λ, its memory usage is reduced to avoid resource waste, and resource utilization information is fed back to the state monitoring module in real time. This achieves a self-healing closed-loop process: network topology mutation → ant colony algorithm iteration surge → health score decline → triggering simulated annealing algorithm switch → resource reallocation → new path data feedback to the state monitoring module → updating ARIMA model input → improving prediction accuracy.
Citation Information
Patent Citations
Network self-healing device based on node state data
CN119854097A
Fault diagnosis and adaptive reconstruction method for communication network of power distribution network
CN120050159A