Automatic Fault Recovery Method for Internet of Things Communication System Based on Fault Tolerant Design
By building a multi-source data monitoring network and reinforcement learning algorithm, the rapid fault detection and adaptive repair of the Internet of Things communication system are achieved, solving the problems of slow recovery speed and inability to monitor internal faults in the existing technology, and improving the fault handling efficiency and system reliability.
Patent Information
- Application Number
- CN202411027023.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-07-29
AI Technical Summary
When the existing IoT communication systems have a large number of nodes or are widely distributed, the recovery speed is slow, the recovery efficiency is low, and the internal failures of the node cannot be monitored, affecting the safety and operational efficiency of the building.
By deploying multiple types of building risk monitors to build a multi-source data monitoring network, collect building status data in real time, use IoT technology and weighted average method to fusion data, combine preset thresholds and anomaly detection algorithms, use reinforcement learning algorithms to build a fault prediction model, adaptively automatically repair and generate repair suggestions, and record fault events and their processing processes.
It realizes fast and accurate fault detection and positioning, improves fault handling efficiency, reduces the delay and constraints of manual operations, and enhances the fault tolerance and reliability of the system.
Smart Images

Figure CN118748642B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the Internet of Things, and specifically relates to an automatic fault recovery method for an Internet of Things communication system based on fault tolerance design. Background Art
[0002] With the rapid development of the Internet of Things technology, more and more devices and systems are connected to the Internet, making the application of the Internet of Things communication system more and more extensive in modern society, such as intelligent transportation, smart home, and remote healthcare. However, due to factors such as device aging, network fluctuations, and software errors, the system fails, affecting the normal operation of the system.
[0003] In terms of building monitoring, with the increase in the number of building monitors and the diversification of functions, the complexity of the Internet of Things communication system is also continuously increasing, including various sensors, actuators, gateways, and cloud server components, which work together to achieve comprehensive monitoring of the building environment, device status, and energy consumption. In a complex Internet of Things communication system, the risk of failure is inevitable. These failures may be caused by device failures, communication link problems, software errors, or external interference. In a building monitor, any failure may lead to data loss, false alarms, or system paralysis, thus affecting the safety and operation efficiency of the building.
[0004] For example, the patent with the authorization announcement number CN109714733B discloses a method for detecting and recovering Internet of Things communication faults and an Internet of Things system, including: the first UE transmits a wireless fault notification signal when it is in a faulty state, and the second UE, when it is operating normally, periodically scans for wireless signals based on a preset scanning frequency band, establishes a short-distance communication wireless connection with the first UE, and performs a fault recovery operation on the first UE based on the short-distance communication wireless connection. The method and the Internet of Things system of this technical solution utilize short-distance communication technology to be able to construct a point-to-point network between adjacent nodes, and achieve the recovery of abnormal terminals in a small range and by gradually spreading, which can solve the island problem caused by wireless communication coverage, terminal faults, etc., and bring convenience to subsequent fault troubleshooting and operation and maintenance, and solve the positioning and system update problems of abnormal terminals in a cellular Internet of Things.
[0005] The above existing technologies all have the following problems: 1) When the number of nodes is large or the distribution is wide, the recovery speed is slow and the recovery efficiency is low; 2) It is impossible to monitor internal node faults. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention proposes an automatic fault recovery method for an Internet of Things communication system based on fault-tolerant design. By deploying multiple types of building risk monitors, a multi-source data monitoring network is constructed to collect building status data in real time, and data fusion is performed using Internet of Things technology and the weighted average method. Combining preset thresholds and anomaly detection algorithms, the system can automatically detect anomalies and risk situations in the building. Based on the fault features extracted from the data, a fault prediction model is constructed using a reinforcement learning algorithm. After fault isolation, the system performs adaptive automatic repair and generates repair suggestions. The fault events and their handling processes are recorded, and fault reports and improvement suggestions are generated through big data analysis, improving the accuracy and timeliness of fault detection.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] An automatic fault recovery method for an Internet of Things communication system based on fault-tolerant design, including:
[0009] Step S1: Construct a multi-source data monitoring network by deploying multiple types of building risk monitors, collect building status data, and perform data fusion to obtain fused building status data;
[0010] Step S2: Extract fault features of vibration, inclination, and settlement from the fused building status data. Based on historical data, real-time building status data, and fault features, use a reinforcement learning algorithm to construct and optimize a fault prediction model. And by real-time analyzing the fused building status data, combining preset thresholds and anomaly detection algorithms, automatically detect anomalies and risk situations in the building. When the CPU detects an anomaly or fault, combine with the fault prediction model to locate the fault point and isolate the fault point;
[0011] Step S3: After fault isolation, the system performs adaptive automatic repair, and according to the fault type and repair situation, the system automatically generates repair suggestions. If the automatic repair fails, manual intervention is performed;
[0012] Step S4: Record the fault events and their handling processes, and regularly generate fault reports and improvement suggestions;
[0013] The specific steps of the said Step S2 include:
[0014] S2.1: Obtain the fused building status data, and use the TPCA algorithm to extract features from the fused building status data to obtain building status feature data Among them, represents the vibration fault feature data, represents the inclination fault feature data, represents the settlement fault feature data, and n, m, r respectively represent the quantities of vibration, inclination, and settlement fault feature data;
[0015] S2.2: Obtain historical building status data, and extract the failure mode, frequency, and trend from the historical building status data;
[0016] S2.3: Based on and the historical building status data, use the reinforcement learning algorithm to construct a failure prediction model;
[0017] The specific steps of step S2 also include:
[0018] S2.4: According to the real-time building status data and the prediction results of the failure prediction model, use the improved online gradient descent algorithm to adjust the parameters of the failure prediction model in real time;
[0019] S2.5: Use the improved g-σ principle outlier detection algorithm to analyze the fused building status data. The formula of the g-σ principle outlier detection algorithm is:
[0020]
[0021] where, represents the fused building status data, w j represents the data weight, d max represents the upper outlier threshold, d min represents the lower outlier threshold, b represents the number of fused building status data, j represents the position index of the fused building status data, and g represents a positive number;
[0022] The specific steps of step S2 also include:
[0023] S2.6: If or then it is determined that the building status data is abnormal, and an exception alarm is triggered. When the CPU detects an exception, an exception handling process is triggered;
[0024] S2.7: Combine the failure prediction model to analyze the building status parameters with CPU exceptions. Through data analysis and model reasoning, locate the failure point or risk area. Once the failure point is located, the system immediately takes isolation measures to isolate the failure point;
[0025] The specific steps of S2.3 include:
[0026] S2.31: According to define the state space of the building system and the action space of the building system, set the reward function according to the state change of the system and the actions taken, and perform the division of the training set and the test set;
[0027] S2.32: Take successful fault prevention as a reward and fault occurrence as a punishment, initialize the Q-value network and the target Q-value network of DQN, and use a deep neural network as the model structure;
[0028] S2.33: For each state in the training set, randomly select an action with probability ε according to the ε-greedy policy, select the action with the highest Q-value with probability (1 - ε), execute the selected action, and record the obtained reward and the next state;
[0029] The specific steps of S2.3 also include:
[0030] S2.34: Store the experience data (state; action; reward; next state) in the experience replay buffer, and randomly sample a batch of experience data from the experience replay buffer for training the DQN model;
[0031] S2.35: Calculate the reward after executing the action plus the highest Q-value of the next state, and update the weights of the Q-value network using the gradient descent algorithm. The formula is:
[0032]
[0033] where y k represents the target Q-value, p k represents the immediate reward obtained in state s k when the action is executed, η represents the discount factor, μ represents the weight parameter of the environment, U cer (s k+1 ) represents the uncertainty function of state s k+1 , Q(s k+1 , l dz ; θ) represents the Q-value obtained according to the target Q-value network θ in the next state s k+1 , ψ represents the weight parameter of the entropy that controls action selection, E tro (l dz ) represents the entropy function of action l dz , s k+1 represents the next state reached after executing the action;
[0034] The specific steps of S2.3 also include:
[0035] S2.36: Every fixed number of steps, update the weights of the target Q-value network using the weights of the Q-value network, and evaluate the performance of the DQN model;
[0036] S2.37: Use the trained DQN model for fault prediction, take corresponding preventive measures according to the prediction results, and adjust the fault prediction strategy using the information feedback by the reward function according to the actual operation effect after implementing the preventive measures;
[0037] S2.38: Evaluate and optimize the adjusted fault prediction model.
[0038] Specifically, the formula for the improved online gradient descent algorithm in S2.4 is as follows:
[0039]
[0040] where α t+1 and α t represent the parameters of the fault prediction model at the (t + 1)-th and t-th iterations respectively, β represents the learning rate, δ represents a constant, m t-1 and v t-1 represent the first-order moment estimate and second-order moment estimate of the gradient respectively, λ1 and λ2 represent the exponential decay rates of the moment estimates, and ▽L(α t ; row t , col t ) represents the gradient of the loss function L with respect to the parameter α t at the position of the data point (row t , col t ).
[0041] Specifically, an electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the automatic fault recovery method for an Internet of Things communication system based on fault tolerance design.
[0042] Specifically, a computer-readable storage medium stores computer instructions, and when the computer instructions run, they execute the steps of the automatic fault recovery method for an Internet of Things communication system based on fault tolerance design.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] 1. The present invention proposes an automatic fault recovery method for an Internet of Things communication system based on fault tolerance design. By deploying multi-type building risk monitors to construct a multi-source data monitoring network and combining a preset threshold and an anomaly detection algorithm, it can perform fault detection quickly and accurately; by constructing and optimizing a fault prediction model, the system can predict potential fault risks in advance, accurately locate the fault points, and improve the efficiency of fault handling.
[0045] 2. The present invention proposes an automatic fault recovery method for an Internet of Things communication system based on fault tolerance design. By isolating the fault points, it realizes adaptive automatic repair and fault isolation, avoids the delay problem caused by manual operation, and at the same time reduces the constraints of manual operation. Description of the Drawings
[0046] Figure 1Schematic diagram of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design of the present invention;
[0047] Figure 2 Flowchart of the working process of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design of the present invention;
[0048] Figure 3 Flowchart of the implementation of fault prediction and isolation of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design of the present invention;
[0049] Figure 4 Flowchart of the construction of the fault prediction model of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design of the present invention. Detailed implementation manners
[0050] Embodiment 1
[0051] Please refer to Figure 1 - Figure 2 A kind of embodiment provided by the present invention: an automatic fault recovery method for the Internet of Things communication system based on fault tolerance design, including the following steps:
[0052] Step S1: Construct a multi-source data monitoring network by deploying multi-type building risk monitors, collect building status data through the multi-source data monitoring network, transmit the building status data to the CPU by means of the Internet of Things, and perform data fusion by using the weighted average method to obtain the fused building status data;
[0053] Among them, the multi-type building risk monitors include vibration, inclination, and settlement monitors, ensuring the redundancy and accuracy of the data, and at the same time ensuring that when one monitor fails, the remaining monitors can still work normally and provide monitoring data; the Internet of Things method refers to the WIFI method.
[0054] Furthermore, the specific formula of the weighted average method is:
[0055]
[0056] Among them, W av represents the average weight after weighted processing, x i represents the building status data value obtained from the i-th monitor, w i represents the weight of the building status data value of the i-th monitor, which is usually determined according to the accuracy and reliability factors of the monitor, and q represents the total number of monitors.
[0057] Step S2: Extract the fault features of vibration, inclination, and settlement from the building status data. Based on the historical data, real-time building status data, and fault features, use the reinforcement learning algorithm to construct and optimize the fault prediction model. And through real-time analysis and fusion of the building status data, combined with the preset threshold and anomaly detection algorithm, automatically detect the anomalies and risk situations in the building. When the CPU detects an anomaly or fault, combine it with the fault prediction model to locate the fault point and isolate the fault point;
[0058] Step S3: After the fault isolation, the system performs adaptive automatic repair, and according to the fault type and repair situation, the system automatically generates repair suggestions. If the automatic repair fails, manual intervention is carried out;
[0059] Furthermore, the specific steps of step S3 include:
[0060] (1) Fault isolation: Through system detection or manual diagnosis, determine the specific location or scope where the fault occurs, and according to the fault type, such as distribution network fault, equipment fault, perform corresponding isolation operations, such as disconnecting the circuit breaker, closing the switch;
[0061] (2) Adaptive automatic repair: The system automatically attempts to repair according to the preset fault handling strategy. The repair strategies include restarting the device, switching to the standby power supply, and adjusting the parameter settings;
[0062] (3) Generating repair suggestions: The system generates corresponding repair suggestions according to the fault type and repair situation, including replacing components, upgrading software, and adjusting the operation mode;
[0063] (4) Automatic repair evaluation: The system evaluates the effect of the automatic repair, judges whether the fault has been resolved. If the fault is resolved, record the repair process and results;
[0064] (5) Manual intervention: If the automatic repair fails, the system triggers the manual intervention process, notifies the relevant personnel, provides the fault information and repair suggestions, and is further processed manually.
[0065] Step S4: Record the fault events and their handling processes, and use big data analysis and data mining algorithms to intelligently analyze the fault records, and regularly generate fault reports and improvement suggestions.
[0066] Embodiment 2
[0067] Please refer to Figure 3 - Figure 4 , the specific steps of this embodiment in step S2 include:
[0068] S2.1: Obtain the fused building status data, and use the TPCA algorithm to extract features from the fused building status data to obtain the building status feature data Among them, Represents vibration fault feature data, Represents tilt fault feature data, Represents settlement fault feature data, where n, m, and r respectively represent the quantities of vibration, tilt, and settlement fault feature data;
[0069] S2.2: Obtain historical building status data, and extract fault patterns, frequencies, and trends from the historical building status data;
[0070] S2.3: Based on and the historical building status data, use a reinforcement learning algorithm to construct a fault prediction model;
[0071] S2.4: According to the real-time building status data and the prediction results of the fault prediction model, use an improved online gradient descent algorithm to adjust the parameters of the fault prediction model in real time. The formula for the improved online gradient descent algorithm is:
[0072]
[0073] where α t+1 , α t respectively represent the parameters of the fault prediction model in the (t + 1)-th and t-th iterations, β represents the learning rate, δ represents a constant used to prevent division by zero, m t-1 , v t-1 respectively represent the first-order moment estimate and second-order moment estimate of the gradient, λ1 and λ2 represent the exponential decay rates of the moment estimates, ▽L(α t ; row t , col t ) represents the gradient of the loss function L with respect to the parameter α t at the data point (row t , col t ). In the present invention, by collecting and processing building status data in real time, combining an online learning algorithm to update and optimize the fault prediction model in real time, and monitoring the performance of the model through performance metrics, it can be ensured that the model can accurately and timely predict building faults, improving the safety and reliability of the building;
[0074] It should be noted that in the existing online gradient descent algorithm, a fixed learning rate is used, resulting in too large a step size in the initial stage of learning, missing the optimal solution, or too slow a convergence rate in the later stage. Moreover, in the calculation process, the problems of gradient disappearance or gradient explosion are common. Especially when the number of network layers is large, the sensitivity to noise data will be relatively high. To address these two types of problems, the present invention introduces the concepts of momentum method and adaptive learning rate on the basis of the existing online gradient descent algorithm. Among them, introducing a momentum variable helps to accelerate the convergence of gradient descent in the relevant direction and suppress oscillations, while introducing the concept of an adaptive learning rate variable can perform deviation correction in real time and accelerate the convergence process of the optimal solution. Therefore, the inventors of the present application creatively added the multi-order moment estimation variable of the gradient itself and its exponential decay rate, automatically adjusted the learning rate of each parameter according to the historical changes of the gradient, realized the adaptivity of the learning rate, and at the same time, reduced the sensitivity to noise, making the parameter update more stable and better able to handle the problems of gradient disappearance / explosion.
[0075] S2.5: Analyze the fused building state data using an improved g-σ principle outlier detection algorithm. The formula of the g-σ principle outlier detection algorithm is:
[0076]
[0077] where, represents the fused building state data, w j represents the data weight, d max represents the upper outlier threshold, d min represents the lower outlier threshold, b represents the number of fused building state data, j represents the position index of the fused building state data, g represents a positive number, represents the weighted standard deviation.
[0078] In the present invention, the range of outliers is redefined on the basis of the existing outlier detection algorithm. The inventors of the present application creatively found that each data point has different importance, and introduced the concept of weight based on this different importance; at the same time, the processing process of missing value variables was introduced to avoid error problems caused by data missing; at the same time, the inventors of the present application creatively abandoned the original three-fold principle and used different distance metrics; therefore, in the present invention, the missing values are filled with the standard deviation, and a weight is assigned to each data point, and the final range of outliers is obtained by adjusting the distance metric. The range of outliers obtained in this way is more accurate than the range of outliers obtained by other methods.
[0079] For example, for an Internet of Things system that is responsible for monitoring data from multiple sensor nodes, such as temperature, humidity, and air pressure, this sensor data is crucial for understanding the environmental status in real time, but occasionally there may be outliers or missing values. Among them,
[0080] (1) Different sensor data has different importance for environmental monitoring. For example, under extreme weather conditions, temperature data may be more important than humidity data. During the calculation process, a weight is assigned to each sensor data. For example, the weight of temperature data is 2, the weight of humidity data is 1, and the weight of air pressure data is 1.5. When calculating outliers, the value of each data point is multiplied by its corresponding weight;
[0081] (2) To avoid the impact of missing values on outlier detection and reduce errors, a processing process for missing value variables is introduced. Assume that the humidity data is missing at a certain time point, and the mean value of the humidity data is 60% and the standard deviation is 10%. Then the filled missing value is 60% ± 10%, and 60% + 10% = 70% can be selected as the filled value;
[0082] (3) The traditional three - standard - deviation method may not be applicable to all datasets, especially when the data distribution is not normal. Therefore, it is very necessary to use different distance metrics; assume that the weighted data and the filled missing values are already available. Then, the Manhattan distance is used to calculate the distance from each data point to the mean of all data points. Among them, the k - value is set to 1.5, that is, if the distance of a certain data point is greater than 1.5 times the mean distance of all points, then this data point is considered an outlier.
[0083] In summary, through this method, the Internet of Things system can more accurately identify outliers, effectively handle missing values at the same time, and improve the reliability of the data and the fault tolerance of the system.
[0084] S2.6: If or then it is determined that the building status data is abnormal and an abnormal alarm is triggered. When the CPU detects an abnormality, an exception handling process is triggered;
[0085] S2.7: Combine the fault prediction model to analyze the building status parameters with CPU abnormalities. Through data analysis and model reasoning, locate the fault point or risk area. Once the fault point is located, the system immediately takes isolation measures to isolate the fault point.
[0086] Furthermore, the specific implementation steps for locating the fault point include:
[0087] (1) Data collection: Real - time collect the status parameter data corresponding to the CPU in the building system, such as CPU usage, temperature, and power consumption;
[0088] (2) Anomaly detection: Real-time analysis of the collected status parameter data through the trained fault prediction model to determine whether there are any anomalies;
[0089] (3) Feature analysis: In-depth analysis of the abnormal data to extract key features, such as a sudden spike in CPU usage or an abnormal increase in temperature;
[0090] (4) Fault location: Using the 7-layer network structure analysis method, combined with the abnormal features and parameter values, to locate the fault point or risk area. The 7-layer network structure analysis method is the prior art content in this field and is not the creative solution of this application, so it will not be elaborated here;
[0091] (5) Verification and confirmation: Verify the accuracy of the fault location through actual tests or user feedback.
[0092] The isolation measures include: 1) Hardware isolation. If the fault point is located in a specific hardware device, such as the CPU or memory, physical isolation measures can be taken, such as shutting down the device or disconnecting the connection; 2) Software isolation: Through software means, such as restarting the service, closing the process, or disabling the function, isolate the fault point within a smaller range to prevent it from affecting other parts; 3) Network isolation: If the fault point is in the network, isolate the fault point from the network through network configuration or security devices, such as a firewall; 4) System isolation: For the entire system or application, take measures such as backup and recovery or switching to a standby system to isolate the fault point from the system and ensure the continuity of the business; 5) Data isolation: If the fault point involves data, take measures such as data backup, data migration, and data isolation to ensure the security and integrity of the data.
[0093] The specific steps of S2.3 include:
[0094] S2.31: According to Define the state space of the building system and the action space of the building system, set the reward function according to the state changes of the system and the actions taken, and Perform the division of the training set and the test set. Among them, in the state space of the building system, the state is represented by the extracted fault features and historical building state data, the action space contains all maintenance or operation strategies, such as replacing components and adjusting parameters, and the reward function reflects the impact of different actions on the system performance. For example, an action that successfully prevents a fault should receive a high reward, while an action that causes a fault should receive a low reward.
[0095] S2.32: Take the successful prevention of faults as a reward and the occurrence of faults as a punishment, and initialize the Q-value network and the target Q-value network of the DQN, using a deep neural network as the model structure;
[0096] S2.33: For each state in the training set, randomly select an action with probability ε according to the ε-greedy policy, select the action with the highest Q value with probability (1 - ε), execute the selected action, and record the obtained reward and the next state;
[0097] S2.34: Store the experience data (state; action; reward; next state) in the experience replay buffer, and randomly sample a batch of experience data from the experience replay buffer for training the DQN model;
[0098] S2.35: Calculate the reward obtained after executing the action plus the highest Q value of the next state, and update the weights of the Q value network using the gradient descent algorithm. The formula is:
[0099]
[0100] where y k represents the target Q value, p k represents the immediate reward obtained in state s k when the action is executed, η represents the discount factor, μ represents the weight parameter of the environment, U cer (s k+1 ) represents the uncertainty function of state s k+1 , Q(s k+1 , l dz ; θ) represents the Q value obtained according to the target Q value network θ in the next state s k+1 , ψ represents the weight parameter of the entropy that controls action selection, E tro (l dz ) represents the entropy function of action l dz , s k+1 represents the next state reached after executing the action;
[0101] It should be noted that in the present invention, when calculating the target Q value, there are factors such as environmental uncertainty, time decay, and uncertainty in action selection, which will affect the Q value calculation. Therefore, it is necessary to consider these uncertainties in the algorithm design. Based on the existing technology, the uncertainty of environmental state transition is considered, making it more robust when facing unpredictable or dynamically changing environments, and helping to reduce the performance degradation caused by environmental mutations. Considering the uncertainty in action selection during the decision-making process can enable the algorithm to find a better balance between exploration and exploitation, avoiding overconfidence in a certain action that may not be optimal. By introducing uncertainty into the target Q value calculation, the algorithm can more intelligently select the exploration direction, preferentially explore areas or actions with higher uncertainty, and thus is more likely to discover better strategies. Therefore, it is very necessary to increase the consideration of uncertainty in the present invention.
[0102] S2.36: At fixed intervals of steps, update the weights of the target Q-value network using the weights of the Q-value network, and evaluate the performance of the DQN model;
[0103] The weight update formula for the target Q-value network is:
[0104]
[0105] where ξ represents the learning rate of the target Q-value network, represents the partial derivative, represents the loss function with respect to the network parameters gradient, represents in the next state s k under, according to the target Q-value network parameters obtained Q-value;
[0106] S2.37: Use the trained DQN model for fault prediction, take corresponding preventive measures according to the prediction results, and adjust the fault prediction strategy using the information fed back by the reward function according to the actual operation effect after implementing the preventive measures;
[0107] S2.38: Evaluate and optimize the adjusted fault prediction model.
[0108] Embodiment 3
[0109] An electronic device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design.
[0110] A computer-readable storage medium, on which computer instructions are stored, and when the computer instructions run, they execute the steps of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design.
[0111] The embodiments of the present invention have been described above in conjunction with the accompanying drawings, but the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make changes, modifications, substitutions, and variations to the above embodiments without departing from the purpose and scope of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. An automatic fault recovery method for an Internet of Things communication system based on fault-tolerant design, characterized in that, Including: Step S1: Construct a multi-source data monitoring network by deploying multi-type building risk monitors, collect building status data, and perform data fusion to obtain fused building status data; Step S2: Extract fault features of vibration, inclination, and settlement from the fused building status data. Based on historical data, real-time building status data, and fault features, use a reinforcement learning algorithm to construct and optimize a fault prediction model. And through real-time analysis of the fused building status data, combined with preset thresholds and anomaly detection algorithms, automatically detect anomalies and risk situations in the building status data. When the CPU detects an anomaly or fault, combine with the fault prediction model to locate the fault point and isolate the fault point; Step S3: After fault isolation, the system performs adaptive automatic repair, and according to the fault type and repair situation, the system automatically generates repair suggestions. If the automatic repair fails, manual intervention is carried out; Step S4: Record the fault events and their handling processes, and regularly generate fault reports and improvement suggestions; The specific steps of the said Step S2 include: S2.1: Obtain the fused building status data, and use the TPCA algorithm to extract features from the fused building status data to obtain the building status feature data Among them, represents the vibration fault feature data, represents the inclination fault feature data, represents the settlement fault feature data, and n, m, and r respectively represent the quantities of the vibration, inclination, and settlement fault feature data; S2.2: Obtain historical building status data, and extract fault patterns, frequencies, and trends from the historical building status data; S2.3: Based on and historical building status data, a fault prediction model is constructed using a reinforcement learning algorithm; The specific steps of the said Step S2 further include: S2.4: According to the real-time building status data and the prediction results of the fault prediction model, use an improved online gradient descent algorithm to adjust the parameters of the fault prediction model in real time; S2.5: Use an improved g-σ principle outlier detection algorithm to analyze the fused building status data. The formula of the g-σ principle outlier detection algorithm is: Among them, represents the fused building status data, w j represents data is the weight of, d max represents the upper threshold of outliers, d min represents the lower threshold of outliers, b represents the number of fused building status data, j represents the position index of the fused building status data, and g represents a positive number; The specific steps of the said Step S2 further include: S2.6: If or it is determined that the building status data is abnormal and an exception alarm is triggered. When the CPU detects an exception, an exception handling process is triggered; S2.7: Combine with the fault prediction model, analyze the building status parameters with CPU anomalies, and through data analysis and model reasoning, locate the fault point or risk area. Once the fault point is located, the system immediately takes isolation measures to isolate the fault point; The specific steps of the said S2.3 include: S2.31: According to Define the state space of the building system and the action space of the building system, set the reward function according to the state change of the system and the actions taken, and Divide the training set and the test set; S2.32: Take successful fault prevention as a reward and fault occurrence as a punishment, and initialize the Q-value network and target Q-value network of DQN, using a deep neural network as the model structure; S2.33: For each state in the training set, randomly select an action with a probability of ε according to the ε-greedy strategy, select the action with the highest Q value with a probability of (1 - ε), execute the selected action, and record the obtained reward and the next state; The specific steps of the said S2.3 further include: S2.34: Store the experience data in the experience replay buffer, and randomly extract a batch of experience data from the experience replay buffer for training the DQN model; the experience data is a quadruple data of state, action, reward, and next state; S2.35: Calculate the reward after executing the action plus the highest Q value of the next state, and use the gradient descent algorithm to update the weights of the Q-value network. The formula is: Among them, y k represents the target Q value, p k represents the immediate reward obtained after executing the action in state s k , η represents the discount factor, μ represents the weight parameter of the environment, U cer (s k+1 ) represents the uncertainty function of state s k+1 , Q(s k+1 , l dz ; θ) represents the Q value obtained according to the target Q value network θ in the next state s k+1 , ψ represents the weight parameter of the entropy for controlling action selection, E tro (l dz ) represents the entropy function of action l dz , s k+1 represents the next state reached after executing the action; The specific steps of the said S2.3 further include: S2.36: Every fixed number of steps, use the weights of the Q-value network to update the weights of the target Q-value network, and evaluate the performance of the DQN model; S2.37: Use the trained DQN model for fault prediction, take corresponding preventive measures according to the prediction results, and adjust the fault prediction strategy using the information feedback by the reward function according to the actual operation effect after implementing the preventive measures; S2.38: Evaluate and optimize the adjusted fault prediction model.
2. The automatic fault recovery method for the Internet of Things communication system based on fault-tolerant design according to claim 1, characterized in that The formula of the improved online gradient descent algorithm in the above S2.4 is: Among them, α t+1 and α t respectively represent the parameters of the fault prediction model at the (t + 1)-th and t-th iterations, β represents the learning rate, δ represents a constant, m t-1 and v t-1 respectively represent the first-order moment estimate and the second-order moment estimate of the gradient, λ1 and λ2 represent the exponential decay rates of the moment estimates, represents the gradient of the loss function L with respect to the parameter α t at the position of the data point (row t , col t ).
3. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design described in any one of claims 1-2.
4. A computer-readable storage medium, characterized in that, Stored thereon are computer instructions, and when the computer instructions run, they execute the steps of the automatic fault recovery method for the Internet of Things communication system based on fault tolerance design described in any one of claims 1-2.
Citation Information
Patent Citations
Methods for detecting and recovering from IoT communication failures and IoT systems
CN109714733B
Intelligent equipment management system and method based on Internet of Things
CN116823227A
Method for detecting steady-state information transmission anomaly of power grid regulation and control system
CN117639265A