Power system server fault tolerance control method and device based on noise enhancement
By introducing noise enhancement data sets and multi-objective optimization strategy models in power system server fault tolerance control, the problem of insufficient adaptability of the existing technology in complex noise environments and data is solved, and accurate fault tolerance control and improved robustness are achieved.
Patent Information
- Application Number
- CN202411817094.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
AI Technical Summary
In the face of complex noise environments and data loss, the power system server fault fault tolerance control is insufficient, resulting in low fault positioning accuracy, long diagnosis process, and difficult to cope with complex situations of multiple fault coupling.
A power system server fault tolerance control method based on noise enhancement is proposed. By obtaining the power system server running data, a noise enhancement data set containing different noises is constructed, the objective function is designed and a multi-objective optimization strategy model is constructed, and the model is trained to perform server fault tolerance control.
This method can achieve accurate fault tolerance control in a noisy environment, improve the robustness and adaptability of the model, reduce the dependence on real-time data acquisition and manual intervention, and is suitable for complex power system scenarios.
Smart Images

Figure CN119938398A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of power system control technology, and in particular to a method and device for fault tolerance control of a power system server based on noise enhancement. Background Art
[0002] With the development and expansion of smart grids, the reliability and stability of power system servers have become critical. However, traditional fault-tolerant control methods rely on preset rules and expert systems, which are simple and direct to implement, but their limitations are gradually revealed in the face of the complex operating environment of modern power systems, especially in terms of rule base maintenance, system scalability and new fault mode identification.
[0003] Most existing fault diagnosis systems rely on manual experience and judgment, resulting in low fault location accuracy, a long diagnostic process, and difficulty in coping with complex situations of multiple fault coupling or achieving fast and accurate diagnosis and predictive maintenance. In recent years, machine learning methods such as support vector machine (SVM) fault classification and neural network fault prediction have been widely used in the field of fault-tolerant control, but these methods rely on real-time data acquisition, have a large computational load, and have high requirements on data quality. Their performance will be affected by noise interference and data loss. In addition, traditional reinforcement learning methods have the problems of high computational complexity and easy to fall into local optimal solutions when dealing with high-dimensional state spaces, especially in continuous action spaces, where their performance is unsatisfactory.
[0004] At the same time, traditional methods have limited ability to resist interference, and there is room for improvement in the use of historical data. The model's generalization ability and ability to handle data distribution deviation are insufficient, especially in the case of sample imbalance, it is difficult to achieve ideal results.
[0005] Therefore, there is an urgent need for a fault-tolerant control scheme for power system servers based on noise enhancement, aiming to solve the shortcomings of power system servers in fault-tolerant control, especially the lack of adaptability when facing complex noise environments and data loss. Summary of the invention
[0006] In view of this, it is necessary to provide a power system server fault tolerance control method and system based on noise enhancement to solve the problem of insufficient adaptability in the prior art when facing complex noise environments and data loss.
[0007] A method for fault tolerance control of a power system server based on noise enhancement, comprising:
[0008] Obtain the power system server operation data and construct a noise enhancement data set containing different noises;
[0009] Design objective functions and build multi-objective optimization strategy models;
[0010] Train the multi-objective optimization strategy model according to the noise-enhanced data set;
[0011] The power system performs server fault tolerance control according to the trained multi-objective optimization strategy model.
[0012] Preferably, design the objective function, including:
[0013] Define the system state space S, where each state is represented by a multi-dimensional vector;
[0014] Construct the objective function:
[0015] J(θ) = E[R(s,a)] + λE[D(s,n)]
[0016] where (,) represents the state-action value function, D(s,n) represents the noise identification function, and λ is the weight factor.
[0017] Preferably, the construction of the multi-objective optimization strategy model includes:
[0018] Use the weighted TD error to construct and update the value network Q(s,a), satisfying:
[0019] Q′(s,α) = Q(s,a) + α[r + γmax(Q(s′,a′)) - Q(s,a)]·W(n)
[0020] where W(n) is a weighting factor dynamically adjusted according to the noise intensity, α is the learning rate, and γ is the discount factor;
[0021] Construct and update the noise recognition network D(s,n) through an adversarial mechanism, satisfying:
[0022] L(D) = E[logD(s,n)] + E[log(1 - D(s,G(s)))]
[0023] where G(s) represents the noise generator, which is responsible for introducing appropriate noise into the training data; D(s,n) represents the noise recognition network, and its loss function is L(D);
[0024] Use the policy gradient method to construct and optimize the policy network π θ (a丨s), satisfying:
[0025]
[0026] where Q w (s,a) represents the state-action value estimate provided by the value network; D(s,n) represents the noise level evaluation provided by the noise recognition network.
[0027] Preferably, training the multi-objective optimization strategy model according to the noise-enhanced data set includes:
[0028] Training the multi-objective optimization strategy model based on a priority experience replay mechanism with noise enhancement.
[0029] Preferably, the sample priority calculation formula of the priority experience replay mechanism satisfies:
[0030] p(i) = |δ i + η|D(s i , n i )| α
[0031] where δ i is the TD error, D(s i , n i ) is the noise identification value; α is the exponential priority, and η is the noise weight.
[0032] Preferably, the power system performs server fault tolerance control according to the trained multi-objective optimization strategy model, including:
[0033] Obtaining the current system state data of the power system;
[0034] According to the trained multi-objective optimization strategy model and the current system state data of the power system, obtaining the optimal action; performing server fault tolerance control based on the optimal action.
[0035] Preferably, after performing server fault tolerance control based on the optimal action, it further includes: recording the data after the power system performs server fault tolerance control and using it for subsequent model updates.
[0036] A power system server fault tolerance control device based on noise enhancement includes:
[0037] A data acquisition module for acquiring the operation data of the power system server and constructing a noise-enhanced data set containing different noises;
[0038] A model construction module for designing an objective function and constructing a multi-objective optimization strategy model;
[0039] A model training module for training the multi-objective optimization strategy model according to the noise-enhanced data set;
[0040] A fault tolerance control module for enabling the power system to perform server fault tolerance control according to the trained multi-objective optimization strategy model.
[0041] Compared with the prior art, the beneficial effect of the present application is that the present application proposes a noise-enhanced offline reinforcement learning power system server fault-tolerant control, which simulates different MDP (Markov Decision Process) environments by introducing appropriate noise into the offline data set, thereby training a more robust and adaptable model. This method does not require real-time data collection or manual intervention, and can achieve accurate fault-tolerant control, which is particularly suitable for complex power system scenarios with significant noise impact. The core design includes a main control network and a noise recognition network, which ensures that the system can make accurate decisions in a noisy environment through an adversarial mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flow chart of a method for fault tolerance control of a power system server based on noise enhancement provided in an embodiment of the present application.
[0043] Figure 2 It is a schematic diagram of constructing a target optimization strategy model provided in an embodiment of the present application.
[0044] Figure 3 It is a structural diagram of a power system server fault tolerance control device 300 based on noise enhancement provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] See also Figure 1-Figure 2 , Figure 1 It is a flow chart of a method for fault tolerance control of a power system server based on noise enhancement provided in an embodiment of the present application; Figure 2 The target optimization strategy model construction diagram provided in the embodiment of the present application is shown in FIG. The power system server fault tolerance control method based on noise enhancement includes the following steps:
[0047] S101: Acquire power system server operation data and construct a noise enhancement data set containing different noises.
[0048] The operation data of the power system server includes the historical data of the power system server under various operation states, including key operation parameters (such as CPU usage, memory, network throughput, etc.) and noise interference information. When obtaining the operation data, the sampling frequency can be 100ms, and the duration is not less than 6 months to ensure that the data covers different working scenarios.
[0049] Then, different intensities and types of noise are introduced into the above running data to change its original MDP environment. The noise-enhanced data set simulates various possible environmental changes and constructs more representative training samples. Each type of noise is divided into several levels according to actual application requirements, such as adding different degrees of Gaussian noise, random noise or systematic interference to the original data to change the state transition probability under the decision condition.
[0050] S102: Design objective function and build a multi-objective optimization strategy model.
[0051] Design objective functions, including:
[0052] Define the system state space S, where each state is represented by a multidimensional vector; it contains key system operating indicators and noise characteristics to reflect the impact of different noise levels on the state. For example: the system state space is designed to contain: S = {normal operation state s1, mild anomaly s2, severe anomaly s3, memory leak s4, timeout s5, crash s6}, and each state is represented by a multidimensional vector, which contains key system operating indicators and noise characteristics to reflect the impact of different noise levels on the state. This multidimensional vector representation enables the model to capture the interaction between various factors in a complex environment and make more accurate decisions based on this.
[0053] Construct the objective function:
[0054] J(θ)=E[R(s,a)]+λE[D(s,n)]
[0055] Where (,) represents the state-action value function, D(s,n) represents the noise identification function, and λ is the weight factor used to balance the reinforcement and noise identification of the policy network.
[0056] Furthermore, the construction of the multi-objective optimization strategy model includes:
[0057] Use the weighted TD error to construct and update the value network Q(s, a), satisfying:
[0058] Q′(s,a)=Q(s,a)+α[r+γmax(Q(s′,a′))-Q(s,a)]·W(n)
[0059] Where W(n) is a weighting factor dynamically adjusted according to the noise intensity, α is the learning rate, and γ is the discount factor;
[0060] The noise recognition network D(s,n) is constructed and updated through the adversarial mechanism to meet the following requirements:
[0061] L(D)=E[logD(s,n)]+E[log(1-D(s,G(s)))]
[0062] Among them, G(s) represents a noise generator, which is responsible for introducing appropriate noise into the training data; D(s,n) represents a noise recognition network, and its loss function is L(D);
[0063] The policy network π θ (a丨s) is constructed and optimized by using the policy gradient method, satisfying:
[0064]
[0065] Among them, Q w (s, a) represents the state-action value estimation provided by the value network; D(s,n) represents the noise level evaluation provided by the noise recognition network.
[0066] In this embodiment, the state space defines the data format input to the value network and the policy network, that is, how the agent perceives the environment. Each state in the state space is associated with a series of possible actions. The agent selects the optimal action according to the current state, so as to achieve effective fault tolerance control. Each state in the state space is associated with a series of possible actions, and the objective function is to find a policy that maximizes the long-term reward. By defining an elaborate and accurate state space, it can be ensured that the objective function accurately captures the best action plan in different states.
[0067] S103: Train the multi-objective optimization policy model according to the noise-enhanced data set.
[0068] The overall training process includes:
[0069] N1. Initialize network parameters
[0070] N2. Sample from the experience pool: (including the historical data generated during the interaction between the agent and the environment. Each piece of experience usually includes the environmental state perceived by the agent at a certain moment, the action taken by the agent in this state, the reward immediately obtained after executing the action, the new state transferred to after executing the action, and whether the completion state is reached). Store these experiences to obtain the experience pool, which can be repeatedly utilized in subsequent training without relying on real-time interaction with the environment every time. This not only improves the data utilization rate but also allows the algorithm to be trained in an offline mode.
[0071] N3. Calculate the TD error after enhancing the noise and the noise identification loss
[0072] N4. Update the value network, the policy network, and the noise recognition network
[0073] N5. More target optimization policy model
[0074] N6. Repeat steps N2 - N5 until convergence or the maximum number of iterations is reached
[0075] Preferably, during training, training the multi - objective optimization policy model according to the noise - enhanced data set includes:
[0076] Training the multi - objective optimization policy model based on a priority experience replay mechanism with noise enhancement. The sample priority calculation formula of the priority experience replay mechanism satisfies:
[0077] p(i) = 丨δ i + η丨D(s i ,n i )丨 α
[0078] where δ i is the TD error, D(s i ,n i ) is the noise identification value; α is the exponential priority, and η is the noise weight.
[0079] After introducing noise, some samples that originally seemed ordinary may become more challenging and informative. These noise - enhanced samples can provide additional learning signals to help the model better understand and cope with uncertainties in practical applications. Through priority sampling, the model can correct early learning errors faster without having to repeatedly process already - mastered knowledge points. This is particularly important for dealing with complex power system anomalies because some fault patterns may be rare but crucial.
[0080] As training progresses, the importance of samples changes. PER allows dynamic adjustment of the weight of each sample to ensure that the model always focuses on the currently most valuable data points, thus reducing unnecessary waste of computing resources.
[0081] By effectively reusing existing data, the convergence speed of the model can be significantly accelerated without adding new data. Noise enhancement not only increases the complexity of the data set but also provides more learning signals for the model. PER ensures that these enhanced samples can be fully utilized, thus accelerating the conversion process from experience to knowledge, reducing redundant learning, improving generalization ability, accelerating the convergence speed, and enhancing robustness, etc., significantly improving the training efficiency of the reinforcement learning model.
[0082] S104: The power system performs server fault tolerance control according to the trained multi - objective optimization policy model.
[0083] Performing server fault tolerance control includes:
[0084] Obtain the current system state data of the power system;
[0085] According to the trained multi-objective optimization strategy model and the current system status data of the power system, the optimal action is obtained; based on the optimal action, server fault tolerance control is performed: control actions are executed and system responses are monitored.
[0086] Finally, after performing server fault tolerance control based on the optimal action, it also includes: recording data after the power system performs server fault tolerance control, and using it for subsequent model updates.
[0087] To verify the effectiveness of the proposed method, an experimental evaluation was conducted on the standard D4RL benchmark test set and the power system server operation and maintenance data set. The experimental environment configuration is as follows:
[0088] 1. Experimental Setup
[0089] Hardware platform: RTX 3090GPU, number of training rounds: 1M steps, batch size: 256, learning rate: 3e-4.
[0090] 2. Benchmark environment: Use the Mujoco environment in the D4RL dataset for evaluation.
[0091] 3. Comparison of algorithms
[0092] -Conservative Q-Learning (CQL)
[0093] -Implicit Q-Learning (IQL)
[0094] -Adversarial Trained Actor Critic(ATAC)
[0095] 4. Experimental Results
[0096]
[0097] The power system server contains 6 types of states: healthy and faulty, and the data is collected from the operation data. In view of the huge cost of memory fault implantation, the health status is also regarded as a different state. For this purpose, an offline dataset A is constructed, where dataset A contains 5 types of server faults (S1-S5) and 1 type of healthy state S6. The details of dataset A are shown in Table 2.
[0098]
[0099]
[0100] Table 2: Dataset A
[0101] Based on the comparison results of various methods on data set A, Table 3 is obtained. It can be seen from Table 3 that the method of the present application is significantly better than various baseline algorithms in power system server fault control and has a good performance in fault tolerant control.
[0102]
[0103] Table 3: Comparison results of dataset A
[0104] like Figure 3 As shown, based on the same inventive concept, the present application also provides a power system server fault tolerance control device 300 based on noise enhancement, comprising:
[0105] The data acquisition module 301 is used to acquire the operation data of the power system server and construct a noise enhancement data set containing different noises;
[0106] A model building module 302 is used to design an objective function and build a multi-objective optimization strategy model;
[0107] A model training module 303, used to train the multi-objective optimization strategy model according to the noise enhancement data set;
[0108] The fault-tolerant control module 304 is used to enable the power system to perform server fault-tolerant control according to the trained multi-objective optimization strategy model.
[0109] In the above embodiment, the specific implementation method of each functional module of the power system server fault fault-tolerant control device 300 based on noise enhancement can refer to the introduction of steps S101-S104 in the embodiment of the power system server fault fault-tolerant control method based on noise enhancement and any optional embodiment thereof, and this embodiment will not repeat it again.
[0110] In summary, the power system server fault-tolerant control scheme based on noise enhancement proposed in this application does not need to rely on expert experience guidance. The controller can be trained only by constructing and enhancing offline data sets containing different fault types. Specifically, by introducing appropriate noise in the offline data set and constructing an environment different from the original data set to enhance the training effect, the adaptability of the controller to different fault types can be effectively improved, and a fault-tolerant strategy that performs well in practical applications can be automatically generated. This application solves the problems of insufficient training data and limited adaptability in different fault scenarios. Through enhanced data sets and noise-based fault disturbances, the fault tolerance range of the controller can be expanded, thereby achieving highly automated and more accurate fault response. In addition, compared with traditional methods, this application greatly reduces the system's early warning and response process, eliminates the link of manual intervention, avoids the inefficient steps of alarm-expert verification-manual repair, significantly reduces economic losses and manpower consumption, and effectively improves production efficiency and reliability. In future work, the method of this application is expected to be expanded to more high-risk scenarios that require fault-tolerant control, and a more flexible and adaptable control strategy can be achieved by introducing different types of noise.
[0111] The applicant of this application has made a detailed explanation and description of the implementation examples of this application in conjunction with the drawings in the specification. However, those skilled in the art should understand that the above implementation examples are only preferred implementation schemes of this application, and the detailed description is only to help readers better understand the spirit of this application, and it is not a limitation on the scope of protection of this application. On the contrary, any improvements or modifications based on the inventive spirit of this application should fall within the scope of protection of this application.
Claims
1. A method for fault tolerance control of power system servers based on noise enhancement, characterized in that: include: Obtain the power system server operation data and construct a noise enhancement data set containing different noises; Design objective functions and build multi-objective optimization strategy models; Training the multi-objective optimization strategy model according to the noise enhancement data set; The power system performs server fault tolerance control according to the trained multi-objective optimization strategy model.
2. The method according to claim 1, characterized in that Design objective functions, including: Define the system state space S, where each state is represented by a multi-dimensional vector; Construct the objective function: J(θ)=E[R(s,a)]+λE[D(s,n)] Where (,) represents the state-action value function, D(s,n) represents the noise identification function, and λ is the weight factor.
3. The method according to claim 2, characterized in that The multi-objective optimization strategy model is constructed as follows: Use the weighted TD error to construct and update the value network Q(s, a), satisfying: Q′(s,a)=Q(s,α)+α[r+γmax(Q(s′,a′))-Q(s,a)]·W(n) Where W(n) is a weighting factor dynamically adjusted according to the noise intensity, α is the learning rate, and γ is the discount factor. The noise recognition network D(s,n) is constructed and updated through the adversarial mechanism to satisfy: L(D)=E[logD(s,n)]+E[log(1-D(s,G(s)))] Among them, G(s) represents the noise generator, which is responsible for introducing appropriate noise into the training data; D(s,n) represents the noise recognition network, and its loss function is L(D); Construct and optimize the policy network π using the policy gradient method θ (a|s), satisfying: Among them, Q w (s, a) represents the state-action value estimate provided by the value network; D(s, n) represents the noise level estimate provided by the noise recognition network.
4. The method according to claim 3, characterized in that Training the multi-objective optimization strategy model according to the noise enhancement data set includes: The multi-objective optimization strategy model is trained based on a noise-enhanced priority experience replay mechanism.
5. The method according to claim 4, characterized in that The sample priority calculation formula of the priority experience replay mechanism satisfies: p(i) = |δ i + η|D(s i , n i )| α Among them, δ i is the TD error, D(s i ,n i ) is the noise identification value; α is the exponential priority, and η is the noise weight.
6. The method according to claim 5, characterized in that The power system performs server fault tolerance control according to the trained multi-objective optimization strategy model, including: Obtain current system status data of the power system; According to the trained multi-objective optimization strategy model and the current system status data of the power system, the optimal action is obtained; and the server fault tolerance control is performed based on the optimal action.
7. The method according to claim 6, characterized in that After performing server fault tolerance control based on the optimal action, it also includes: recording data after the power system performs server fault tolerance control, and using it for subsequent model updates.
8. A power system server fault tolerance control device based on noise enhancement, characterized in that: include: A data acquisition module is used to acquire the operation data of the power system server and construct a noise enhancement data set containing different noises; Model building module, used to design objective functions and build multi-objective optimization strategy models; A model training module, used for training the multi-objective optimization strategy model according to the noise enhancement data set; The fault-tolerant control module is used to enable the power system to perform server fault-tolerant control according to the trained multi-objective optimization strategy model.