Current sudden change risk prediction method based on reinforcement learning
Through the improved dual deep Q network structure, combined with multi-strategy evaluation and dynamic reward function, the shortcomings of existing current monitoring methods in adaptability and recognition accuracy are solved, efficient identification and real-time warning of current mutations are achieved, and the safety and reliability of the power system are improved.
Patent Information
- Application Number
- CN202510809629.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-19
AI Technical Summary
Existing current monitoring methods lack adaptability, have limited ability to model nonlinear changes, and have difficulty in achieving continuous learning when identifying current mutations and predicting risks. This leads to low recognition accuracy and insensitive response to complex changes in current signals in power systems.
An improved dual-depth Q-network structure is adopted, combined with a shared feature extraction subnetwork, a steady-state and mutation dual-action value evaluation subnetwork, and a Q-value fusion module. Through the ε-greedy strategy and dynamic reward function, multi-strategy parallel evaluation and weighted fusion of current states are realized, thereby improving the generalization ability and recognition accuracy of the model.
It improves the recognition accuracy and response efficiency of current mutation risks, can realize real-time current mutation warning and safety control in complex power scenarios, enhances the adaptability and stability of the model, and reduces equipment failure rate.
Smart Images

Figure CN120671049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electrical monitoring and artificial intelligence technology, and in particular to a current mutation risk prediction method based on reinforcement learning. Background Art
[0002] During power system operation, the current state of electrical equipment is often closely related to its operational safety. This is especially true in critical equipment such as transformers, distribution cabinets, and cable lines. Sudden current changes often indicate potential abnormal risks, such as equipment aging, short-circuit faults, or system disturbances. Therefore, timely and accurate identification of current sudden changes and prediction of their risk levels are crucial for ensuring the stable operation of power systems. Existing current monitoring methods are primarily based on threshold judgment, pattern matching, or statistical learning models, identifying abnormal behavior by setting fixed upper and lower limits or rate of change criteria. These methods are practical in specific scenarios, but they struggle to adapt to the complex variations in current signals under different operating conditions, and are particularly limited in identifying early signs before sudden changes occur.
[0003] In recent years, some research has begun to introduce support vector machines and decision tree models for current anomaly detection. These methods typically rely on manual feature extraction and offline training of fixed model parameters for classification and recognition. While these methods can improve prediction accuracy to a certain extent when sufficient training data is available, they lack adaptability in dynamic scenarios, have limited ability to model nonlinear variations in current signals, and rely on manual intervention to update the models, making long-term stable operation difficult. Furthermore, these methods face recognition bias caused by changes in sample distribution and lack the ability to continuously learn.
[0004] Deep reinforcement learning technology has been increasingly applied to intelligent control and prediction in recent years, demonstrating its superior time series modeling and autonomous learning capabilities. Existing methods based on deep Q networks have been explored for electrical equipment anomaly detection tasks, optimizing decision-making strategies by constructing a state-action-reward interaction mechanism. However, traditional DQN architectures typically utilize only a single action-value estimation network, making it difficult to effectively distinguish between steady-state behavior and sudden changes in current states, resulting in insufficient sensitivity when predicting sudden risks. Furthermore, current methods generally employ fixed reward strategies, failing to fully consider the impact of prediction timing deviations on response effectiveness, limiting the model's ability to assess sudden risk levels.
[0005] Therefore, how to provide a current mutation risk prediction method based on reinforcement learning is an urgent problem that those skilled in the art need to solve. Summary of the Invention
[0006] One purpose of the present invention is to propose a current mutation risk prediction method based on reinforcement learning. The present invention makes full use of the dual-depth Q network structure, multi-scale current feature extraction and time-series reinforcement learning strategy to construct an intelligent risk identification model for the operating status of electrical equipment. By introducing a shared feature extraction subnetwork, a steady-state and mutation dual-action value assessment subnetwork, and a Q-value fusion module, the joint modeling of steady-state fluctuations and mutation trends in the current state is effectively achieved, and the generalization ability of the model under different operating conditions is improved. The present invention has the advantages of clear structure, high risk identification accuracy and strong adaptability, and can be widely used in real-time current mutation warning and safety control tasks in complex power scenarios such as power grids, rail transit, and industrial power supply.
[0007] A method for predicting current mutation risk based on reinforcement learning according to an embodiment of the present invention includes the following steps:
[0008] S1. Collect the current signal sequence of the monitored electrical equipment and perform feature extraction operations to construct a current state vector;
[0009] S2. Input the current state vector into an improved dual deep Q network, perform forward reasoning operations, and generate an action value prediction result in the current state. The improved dual deep Q network includes a main Q network and a target Q network with the same structure. The main Q network includes a shared feature extraction subnetwork, a first action value evaluation subnetwork, a second action value evaluation subnetwork, and a Q value fusion module.
[0010] S3. Based on the action value prediction results, the ε-greedy strategy is used to perform action selection operations and generate the selected action;
[0011] S4. Construct a reward function to generate a corresponding reward value based on whether the current mutation is successfully predicted, the prediction deviation time, and the preset penalty factor;
[0012] S5. Record the current state vector, the selected action, the reward value, and the next current state vector generated by the environment feedback to form an experience quadruple.
[0013] S6. Based on the experience quadruple, use the target Q network to calculate the target action value, update the parameters of the main Q network, and periodically synchronize the parameters of the main Q network to the target Q network;
[0014] S7. After completing the training optimization, the improved dual deep Q network is deployed to the current monitoring system, which receives the real-time input current state vector sequence and outputs the current mutation risk prediction result.
[0015] Optionally, the current state vector includes an original current sampling sequence consisting of N consecutive sampling points before the current moment and a multidimensional current state feature set; the multidimensional current state feature set includes a current change rate feature, a current first-order difference mean, a window range feature, and a short-time energy feature.
[0016] Optionally, the S2 specifically includes:
[0017] S21, constructing a main Q network of the improved dual deep Q network, wherein the main Q network includes a shared feature extraction subnetwork, a first action value evaluation subnetwork, a second action value evaluation subnetwork and a Q value fusion module;
[0018] S22, inputting the current state vector into the shared feature extraction sub-network, performing a nonlinear feature mapping operation, and generating a unified feature representation vector, wherein the nonlinear mapping process is performed according to the weight parameters of the shared feature extraction sub-network;
[0019] S23. Input the unified feature representation vector into the first action value evaluation subnetwork, perform an action value evaluation operation based on steady-state conditions, and output a corresponding first action value vector, where the first action value vector is generated by the first action value evaluation subnetwork based on the feature representation vector and internal weight parameters;
[0020] S24, inputting the unified feature representation vector into the second action value evaluation subnetwork, performing an action value evaluation operation based on the mutation trend condition, and outputting a corresponding second action value vector, where the second action value vector is generated by the second action value evaluation subnetwork based on the feature representation vector and internal weight parameters;
[0021] S25, inputting the first action value vector and the second action value vector into the Q value fusion module, performing a weighted fusion operation on the action value vectors, and generating an action value prediction result in the current state;
[0022] S26. Construct a target Q network of the improved dual deep Q network, wherein the target Q network structure is consistent with the main Q network structure, adopts a fixed parameter set, and maintains update consistency with the main Q network through a periodic synchronization mechanism to generate a target action value.
[0023] Optionally, the S25 specifically includes:
[0024] S251. Obtain a first action value vector output by a first action value evaluation subnetwork, where the first action value vector represents an action value distribution under a steady-state response;
[0025] S252: Obtain a second action value vector output by the second action value evaluation subnetwork, where the second action value vector represents the action value distribution under the current mutation trend;
[0026] S253. In the Q-value fusion module, a fusion weight factor α is set to adjust the relative contribution of the first action value vector and the second action value vector in the fusion process, where α represents the steady-state condition weight and (1-α) represents the mutation condition weight.
[0027] S254. Perform linear weighted operations on the first action value vector and the second action value vector according to the weighted fusion strategy, fuse them to generate the final action value prediction result vector in the current state, multiply the first action value vector by the fusion weight factor α, multiply the second action value vector by (1-α), and sum the two to obtain the fused action value prediction result.
[0028] Optionally, the S26 specifically includes:
[0029] S261: Construct a target Q network that is consistent with the main Q network structure. The target Q network includes the same shared feature extraction subnetwork, first action value evaluation subnetwork, second action value evaluation subnetwork, and Q value fusion module as the main Q network. All network layers of the target Q network use independent parameter sets W. t To define;
[0030] S262: Input the current state vector in the training sample into the target Q network, pass it through the shared feature extraction subnetwork, two action value evaluation subnetworks and the Q value fusion module in sequence, and output the target action value prediction result;
[0031] S263, set parameter synchronization period T sync , controls the update frequency of the target Q network parameters;
[0032] S264: After each parameter synchronization cycle, perform a parameter copy operation to assign all parameter values in the current main Q network to the target Q network.
[0033] Optionally, the S3 specifically includes:
[0034] S31, define action set A={a1, a2,…, a n}, represents all candidate actions that can be selected under the current current state, and the action set is a pre-set finite set that includes all legal control operations;
[0035] S32, receiving the action value prediction result vector under the current current state and setting the action selection parameter ε;
[0036] S33. Generate a random number r in the interval [0,1] and compare it with ε. If r < ε, execute the exploration strategy and randomly select an action a from the action set. rand as the currently selected action;
[0037] S34. If r ≥ ε, execute the utilization strategy and select the action corresponding to the maximum value of the action value prediction result from the action set as the currently selected action;
[0038] S35. Output selected action a select , and record the state-action pair consisting of the current state vector and the selected action.
[0039] Optionally, the S4 specifically includes:
[0040] S41. Set the reward function based on the selected action a select Whether the current mutation event is successfully predicted to construct the initial reward signal, if the prediction is successful, a positive reward value r is given pos , if the prediction fails, a negative reward value r is given neg ;
[0041] S42, define the prediction deviation time ΔT=|T pred -T actual |, where T pred The current mutation time predicted for the selected action, T actual is the actual time when the current mutation occurs;
[0042] S43. Set a penalty factor λ, and use the success or failure of the prediction of the mutation event and the prediction time deviation as the basis for reward calculation;
[0043] S44, if the selected action is a select Successfully predicting a current mutation event will result in a positive reward value r pos Based on, deduct the product of the predicted time deviation and the penalty factor to get the immediate reward value. If the selected action is a select If the prediction fails, a negative reward value r neg Based on the prediction time, the product of the penalty factor is deducted to obtain the immediate reward value.
[0044] Optionally, the S5 specifically includes:
[0045] S51. After each action is executed, collect the current state vector s at the current moment t and the selected action a select , where t represents the current time step;
[0046] S52, receiving the immediate reward value R(s) generated by environmental feedback t ,a select ), collect the next moment current state vector s generated by the environment feedback after the current action is executed t+1 ;
[0047] S53, combining the current state vector at the current moment, the selected action, the immediate reward value, and the current state vector at the next moment to form an experience quadruple;
[0048] S54. The generated experience quadruple is stored in an experience replay pool. The experience replay pool adopts a first-in-first-out cache mechanism to maintain a fixed-capacity sample set.
[0049] Optionally, the S6 specifically includes:
[0050] S61. Randomly sample a batch of experience quadruple samples from the experience playback buffer, each sample is represented by (s t ,a select ,R(s t ,a select ),r t );
[0051] S62, the current state vector s at the next moment t+1 Input to the target Q network, perform forward reasoning operations, and obtain all candidate actions in s t+1 The target action value prediction result vector in the state and the target Q value corresponding to the maximum value is selected:
[0052]
[0053] Among them, y t is the target action value, R(s,a select ) is the immediate reward value, s t is the current state vector at the previous moment, a select is the selected action, γ is the discount factor, a i For each action in the action set, Q target (s t+1 ,a i ) represents the state s t+1 The target Q network responds to action a i The value prediction results;
[0054] S63, the current state vector s at the current moment t Input into the main Q network, calculate the current action value prediction result vector of all candidate actions, and take the currently selected action a t The corresponding action value prediction result;
[0055] S64, construct an error function between the current action value prediction result and the target action value, the error function adopts the mean square error calculation method to convert the target action value y t With the main Q network for the currently selected action a t The difference between the corresponding action value prediction results is squared;
[0056] S65, performing an error back propagation operation on the main Q network, and updating the parameter sets of each layer in the main Q network based on the loss function;
[0057] S66, when the set parameter synchronization period T is met sync Under these conditions, all parameter sets in the current master Q network are copied to the target Q network.
[0058] Optionally, the S7 specifically includes:
[0059] S71. After completing the training process of the improved dual deep Q network, solidifying the final parameters of the shared feature extraction subnetwork, the first action value evaluation subnetwork, the second action value evaluation subnetwork, and the Q value fusion module in the main Q network;
[0060] S72, deploying the trained improved dual deep Q network into the current monitoring system, establishing a data receiving interface with the current acquisition terminal, and receiving the current signal collected by the sensor in real time;
[0061] S73, pre-process the real-time input current signal and construct the real-time current state vector s real,t , the real-time current state vector includes a current amplitude sequence and a feature set of multiple sampling points in the current window;
[0062] S74, the real-time current state vector s real,t Input into the main Q network structure of the improved dual deep Q network after training, perform forward reasoning operations, and output the corresponding real-time action value prediction vector;
[0063] S75. Based on the real-time action value prediction vector, select the action a corresponding to the maximum action value predict,t , represents the optimal mutation prediction behavior at the current moment;
[0064] S76, the selected action a predict,t The corresponding action meaning is mapped into the current mutation risk prediction result, which is output to the early warning module of the current monitoring system. The prediction result includes the mutation risk level, expected occurrence time and response suggestions.
[0065] The beneficial effects of the present invention are:
[0066] The present invention constructs a current mutation risk prediction method based on reinforcement learning. In response to the problems of low mutation recognition accuracy, poor adaptability and insufficient model updating capability in the prior art, an improved dual-depth Q network architecture is proposed, which effectively improves the recognition capability and prediction response efficiency of abnormal current behavior. The network structure includes a main Q network and a target Q network, wherein the main Q network introduces a shared feature extraction subnetwork, a steady-state action value assessment subnetwork, a mutation action value assessment subnetwork and a Q-value fusion module, thereby realizing independent modeling and joint judgment of different trend behaviors in the current state. Compared with the traditional reinforcement learning model that relies on a single strategy evaluation mechanism, the present invention improves the expression capability of the action value prediction results through multi-strategy parallel design and weighted fusion strategy, thereby improving the model's recognition accuracy of current mutation behavior under complex state changes.
[0067] In terms of feature construction, the present invention defines a multi-dimensional current state feature set including the current change rate, the first-order difference mean, the window range and the short-time energy, which further enhances the model's ability to perceive the microscopic trend of current signal changes. By introducing a dynamic reward function based on mutation prediction accuracy and time deviation, the model can simultaneously consider the balance between behavioral results and timeliness during the learning process, which helps to optimize the strategy output more reasonably. In addition, in the training mechanism, the experience replay strategy and the periodic parameter synchronization mechanism are adopted, so that the main Q network can stably drive the update of the target Q network while continuously optimizing, thereby ensuring the convergence and stability of the model during long-term training.
[0068] By deploying the trained reinforcement learning model into an actual current monitoring system and integrating it with real-time current data collected by sensors, the system continuously outputs predictions that include the mutation risk level, expected occurrence time, and recommended responses. This enables online risk identification and response support tailored to the actual operating status of electrical equipment. Overall, this invention addresses shortcomings of existing technologies in terms of model structure, feature expression, training mechanism, and deployment response, improving the accuracy and real-time performance of current mutation identification, and possesses excellent engineering applicability and promotional value. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0070] Figure 1 This is a flow chart of a current mutation risk prediction method based on reinforcement learning proposed by the present invention;
[0071] Figure 2This is a schematic diagram of the structure of an improved dual deep Q network in a current mutation risk prediction method based on reinforcement learning proposed in the present invention;
[0072] Figure 3 This is a schematic diagram of the internal structure of the main Q network in the current mutation risk prediction method based on reinforcement learning proposed in the present invention. DETAILED DESCRIPTION
[0073] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0074] refer to Figure 1-3 , a current mutation risk prediction method based on reinforcement learning, comprising the following steps:
[0075] S1. Collect the current signal sequence of the monitored electrical equipment and perform feature extraction operations to construct a current state vector;
[0076] S2. Input the current state vector into an improved dual deep Q network, perform forward reasoning operations, and generate an action value prediction result in the current state. The improved dual deep Q network includes a main Q network and a target Q network with the same structure. The main Q network includes a shared feature extraction subnetwork, a first action value evaluation subnetwork, a second action value evaluation subnetwork, and a Q value fusion module.
[0077] S3. Based on the action value prediction results, the ε-greedy strategy is used to perform action selection operations and generate the selected action;
[0078] S4. Construct a reward function to generate a corresponding reward value based on whether the current mutation is successfully predicted, the prediction deviation time, and the preset penalty factor;
[0079] S5. Record the current state vector, the selected action, the reward value, and the next current state vector generated by the environment feedback to form an experience quadruple.
[0080] S6. Based on the experience quadruple, use the target Q network to calculate the target action value, update the parameters of the main Q network, and periodically synchronize the parameters of the main Q network to the target Q network;
[0081] S7. After completing the training optimization, the improved dual deep Q network is deployed to the current monitoring system, which receives the real-time input current state vector sequence and outputs the current mutation risk prediction result.
[0082] The present invention proposes a current mutation risk prediction method based on reinforcement learning. By introducing an improved dual-depth Q network, efficient feature extraction of electrical equipment current signals and accurate risk prediction under complex conditions are achieved. This method can adaptively learn the implicit laws of current mutations, fully integrate multi-dimensional feature information, and improve the stability and robustness of the prediction through the collaborative optimization of the main Q network and the target Q network. By using the ε-greedy strategy and refined reward function design, the exploration and utilization are effectively balanced, and the early warning capability of mutation events is improved. After training and optimization, the system can realize real-time monitoring and intelligent warning of sudden current risks in the actual operating environment, reduce equipment failure rate, improve the operating safety and reliability of the power grid or power system, and provide strong technical support for the intelligent operation and maintenance of electrical equipment.
[0083] In this embodiment, the current state vector includes an original current sampling sequence consisting of N consecutive sampling points before the current moment and a multidimensional current state feature set; the multidimensional current state feature set includes a current change rate feature, a current first-order difference mean, a window range feature, and a short-time energy feature;
[0084] The current change rate feature is defined as the amplitude change rate between the current value at the current sampling point and the current value at the starting sampling point; the current first-order difference mean represents the mean of the first-order differences between adjacent sampling points of the current signal within the current sampling window; the window range feature is defined as the difference between the maximum and minimum current values within the current sampling window; and the short-time energy feature represents the energy accumulation level of the current state segment.
[0085] By constructing a current state vector that fuses the original sampling sequence with multidimensional current state features, this method comprehensively characterizes the timing characteristics of current signals and the signs of sudden changes. This method can enhance the sensitivity of identifying current sudden change risks, improve the speed and accuracy of the prediction model's response to abnormal conditions, and provide more reliable data support for safety monitoring and fault warning of electrical equipment.
[0086] In this embodiment, S2 specifically includes:
[0087] S21, constructing a main Q network of the improved dual deep Q network, wherein the main Q network includes a shared feature extraction subnetwork, a first action value evaluation subnetwork, a second action value evaluation subnetwork and a Q value fusion module;
[0088] S22, inputting the current state vector into the shared feature extraction sub-network, performing a nonlinear feature mapping operation, and generating a unified feature representation vector, wherein the nonlinear mapping process is performed according to the weight parameters of the shared feature extraction sub-network;
[0089] S23. Input the unified feature representation vector into the first action value assessment subnetwork, perform an action value assessment operation based on steady-state conditions, and output a corresponding first action value vector. The first action value vector is generated by the first action value assessment subnetwork based on the feature representation vector and internal weight parameters, and is used to characterize the expected value level of each optional action in the steady-state response mode under the current state;
[0090] S24. Input the unified feature representation vector into the second action value assessment subnetwork, perform an action value assessment operation based on the mutation trend condition, and output a corresponding second action value vector. The second action value vector is generated by the second action value assessment subnetwork based on the feature representation vector and internal weight parameters, and is used to characterize the expected risk value level of each optional action under the current mutation trend in the current state;
[0091] S25. Input the first action value vector and the second action value vector into the Q-value fusion module, perform a weighted fusion operation on the action value vectors, and generate an action value prediction result under the current state, which is used to adjust the relative importance of the two prediction mechanisms in the final action judgment;
[0092] S26. Construct a target Q network of the improved dual deep Q network, wherein the target Q network structure is consistent with the main Q network structure, adopts a fixed parameter set, and maintains update consistency with the main Q network through a periodic synchronization mechanism, generates a target action value, and provides a stable optimization target for the parameter iteration of the main Q network.
[0093] The present invention achieves efficient extraction of multi-dimensional features of current states and complex risk discrimination by introducing a shared feature extraction subnetwork and two types of action value assessment subnetworks into the main Q network. The first action value assessment subnetwork targets the steady-state response mode and can accurately assess the expected value of each action under normal conditions; the second action value assessment subnetwork focuses on the current mutation trend, which improves the sensitivity to abnormal and sudden risks. The two types of action values are weighted and integrated through the Q-value fusion module, which effectively balances the influence of steady-state and mutation predictions in the final decision and enhances the comprehensive discrimination ability of the model. The introduction of the target Q network and the synchronization of periodic parameters further improve the stability of the training process and reduce the value estimation deviation. The overall method improves the accuracy and robustness of current mutation risk prediction, and provides a more reliable decision-making basis for the intelligent monitoring and early warning system of electrical equipment.
[0094] In this embodiment, the S25 specifically includes:
[0095] S251. Obtain a first action value vector output by a first action value evaluation subnetwork, where the first action value vector represents an action value distribution under a steady-state response;
[0096] S252: Obtain a second action value vector output by the second action value evaluation subnetwork, where the second action value vector represents the action value distribution under the current mutation trend;
[0097] S253. In the Q-value fusion module, a fusion weight factor α is set to adjust the relative contribution of the first action value vector and the second action value vector in the fusion process, where α represents the steady-state condition weight and (1-α) represents the mutation condition weight.
[0098] S254. Perform linear weighted operations on the first action value vector and the second action value vector according to the weighted fusion strategy, fuse them to generate the final action value prediction result vector in the current state, multiply the first action value vector by the fusion weight factor α, multiply the second action value vector by (1-α), and sum the two to obtain the fused action value prediction result.
[0099] This method introduces a weighted fusion strategy into the Q-value fusion module. Based on a set fusion weight factor, the action value under steady-state conditions is linearly weighted proportionally with the action value under sudden change conditions, achieving flexible integration of different risk characteristics. The weight factor is used to adjust the contribution of the two types of values to the final prediction, thereby improving the scientific and targeted nature of action selection. This method effectively enhances the model's risk discrimination ability under different current conditions and improves the accuracy and adaptability of current sudden change risk prediction results.
[0100] In this embodiment, the S26 specifically includes:
[0101] S261: Construct a target Q network that is consistent with the main Q network structure. The target Q network includes the same shared feature extraction subnetwork, first action value evaluation subnetwork, second action value evaluation subnetwork, and Q value fusion module as the main Q network. All network layers of the target Q network use independent parameter sets W. t To define;
[0102] S262: Input the current state vector in the training sample into the target Q network, pass it through the shared feature extraction subnetwork, the two action value evaluation subnetworks and the Q value fusion module in sequence, and output the target action value prediction result. The target action value prediction result is used to calculate the target Q value and participate in the error back propagation and parameter update process of the main Q network;
[0103] S263, set parameter synchronization period T sync , controls the update frequency of the target Q network parameters;
[0104] S264: After each parameter synchronization cycle, perform a parameter copy operation to assign all parameter values in the current main Q network to the target Q network.
[0105] This method provides a stable optimization target for the training process of the main Q network by setting a target Q network with the same structure as the main Q network and using an independent parameter set and periodic parameter synchronization mechanism. Parameter synchronization effectively reduces bias in value estimation and avoids oscillation and non-convergence during training. This improves the training efficiency and model convergence stability of current mutation risk prediction, thereby enhancing the practical application reliability of the system.
[0106] In this embodiment, S3 specifically includes:
[0107] S31, define action set A={a1, a2,…, a n}, represents all candidate actions that can be selected under the current current state, and the action set is a pre-set finite set that includes all legal control operations;
[0108] S32, receiving the action value prediction result vector under the current current state and setting the action selection parameter ε;
[0109] S33. Generate a random number r in the interval [0,1] and compare it with ε. If r < ε, execute the exploration strategy and randomly select an action a from the action set. rand as the currently selected action;
[0110] S34. If r ≥ ε, execute the utilization strategy and select the action corresponding to the maximum value of the action value prediction result from the action set as the currently selected action;
[0111] S35. Output selected action a select , and record the state-action pair consisting of the current state vector and the selected action as the basis for subsequent training sample construction.
[0112] This paper defines a set of actions, combines them with action value predictions based on the current state, and introduces an ε-greedy strategy to achieve an effective balance between exploration and exploitation during action selection. By setting the probability parameter ε, actions are randomly selected at a certain ratio to enrich the diversity of experience, while actions with the highest predicted value are prioritized to improve decision optimality. This mechanism helps to improve the convergence speed of model training and the ability to search for the global optimal solution, providing a scientific basis for the construction of subsequent experience samples.
[0113] In this embodiment, the S4 specifically includes:
[0114] S41. Set the reward function based on the selected action a select Whether the current mutation event is successfully predicted to construct the initial reward signal, if the prediction is successful, a positive reward value r is given pos If the prediction fails, a negative reward value r is givenneg ;
[0115] S42, define the prediction deviation time ΔT=|T pred -T actual |, where T pred The current mutation time predicted for the selected action, T actual is the actual time when the current mutation occurs;
[0116] S43. Set a penalty factor λ, and use the success or failure of the prediction of the mutation event and the prediction time deviation as the basis for reward calculation;
[0117] S44, if the selected action is a select Successfully predicting a current mutation event will result in a positive reward value r pos Based on, deduct the product of the predicted time deviation and the penalty factor to get the immediate reward value. If the selected action is a select If the prediction fails, a negative reward value r neg Based on the prediction time, the product of the penalty factor is deducted to obtain the immediate reward value.
[0118] This invention achieves real-time incentives and constraints on model behavior by setting a reward function and combining it with the matching of action predictions with actual current mutation events. The formula ΔT represents the deviation between the predicted time and the actual time of occurrence. Introducing this as a penalty factor in the reward calculation effectively guides the model to improve prediction accuracy while optimizing prediction timeliness. For successful predictions, the reward value is deducted based on the time deviation, while for failed predictions, a negative reward is applied, prompting the model to continuously improve its predictive capabilities. This mechanism helps improve the accuracy and response speed of current mutation event detection, enhancing the practical application value of the system.
[0119] In this embodiment, the S5 specifically includes:
[0120] S51. After each action is executed, collect the current state vector s at the current moment t and the selected action a select , where t represents the current time step;
[0121] S52, receiving the immediate reward value R(s) generated by environmental feedback t ,a select ), collect the next moment current state vector s generated by the environment feedback after the current action is executed t+1 ;
[0122] S53, combining the current state vector at the current moment, the selected action, the immediate reward value, and the current state vector at the next moment to form an experience quadruple;
[0123] S54. The generated experience quadruple is stored in an experience replay pool. The experience replay pool adopts a first-in-first-out cache mechanism to maintain a fixed-capacity sample set.
[0124] This invention effectively accumulates and reuses historical experience by constructing a quadruple of experiences, including the current state, action, reward, and next state, and storing and managing them in an experience replay pool. Experience replay uses a first-in, first-out mechanism to ensure the diversity and representativeness of the sample set, helping to improve the stability and generalization capabilities of model training and enhance the practical application of the current state intelligent decision-making system.
[0125] In this embodiment, S6 specifically includes:
[0126] S61. Randomly sample a batch of experience quadruple samples from the experience playback buffer, each sample is represented by (s t ,a select ,R(s t ,a select ),r t );
[0127] S62, the current state vector s at the next moment t+1 Input to the target Q network, perform forward reasoning operations, and obtain all candidate actions in s t+1 The target action value prediction result vector in the state and the target Q value corresponding to the maximum value is selected:
[0128]
[0129] Among them, y t is the target action value, R(s,a select ) is the immediate reward value, s t is the current state vector at the previous moment, a select is the selected action, γ is the discount factor, a i For each action in the action set, Q target (s t+1 ,a i ) represents the state s t+1 The target Q network responds to action a i The value prediction results;
[0130] S63, the current state vector s at the current moment t Input into the main Q network, calculate the current action value prediction result vector of all candidate actions, and take the currently selected action a t The corresponding action value prediction result;
[0131] S64, construct an error function between the current action value prediction result and the target action value, the error function adopts the mean square error calculation method to convert the target action value y t With the main Q network for the currently selected action a t The difference between the corresponding action value prediction results is squared;
[0132] S65, performing an error back propagation operation on the main Q network, and updating the parameter sets of each layer in the main Q network based on the loss function;
[0133] S66, when the set parameter synchronization period T is met sync Under these conditions, all parameter sets in the current master Q network are copied to the target Q network.
[0134] The present invention randomly samples experience samples from the experience replay area, and calculates the target action value and the current action value based on the target Q network and the main Q network respectively, and uses the mean square error function to measure the deviation between the two. In the formula, the target action value is calculated by the weighted sum of the immediate reward and the maximum action value at the next moment, and the discount factor is used to balance the impact of future rewards. Using the square of the difference between the current action value and the target action value as the loss function helps to guide the main Q network to continuously approach the optimal strategy. Utilizing the backpropagation mechanism, the main Q network parameters are dynamically updated based on the loss function, and the main Q network parameters are periodically synchronized to the target Q network to ensure the stability and convergence speed of the training process. This method improves the action value estimation accuracy and training efficiency of the model, and provides a solid foundation for intelligent decision-making under current conditions.
[0135] In this embodiment, the S7 specifically includes:
[0136] S71. After completing the training process of the improved dual deep Q network, solidifying the final parameters of the shared feature extraction subnetwork, the first action value evaluation subnetwork, the second action value evaluation subnetwork, and the Q value fusion module in the main Q network;
[0137] S72, deploying the trained improved dual deep Q network into the current monitoring system, establishing a data receiving interface with the current acquisition terminal, and receiving the current signal collected by the sensor in real time;
[0138] S73, pre-process the real-time input current signal and construct the real-time current state vector s real,t , the real-time current state vector includes a current amplitude sequence and a feature set of multiple sampling points in the current window;
[0139] S74, the real-time current state vector s real,tInput into the main Q network structure of the improved dual deep Q network after training, perform forward reasoning operations, and output the corresponding real-time action value prediction vector;
[0140] S75. Based on the real-time action value prediction vector, select the action a corresponding to the maximum action value predict,t , represents the optimal mutation prediction behavior at the current moment;
[0141] S76, the selected action a predict,t The corresponding action meaning is mapped into the current mutation risk prediction result, which is output to the early warning module of the current monitoring system. The prediction result includes the mutation risk level, expected occurrence time and response suggestions.
[0142] This invention deploys a trained, improved dual-deep Q network in a current monitoring system to achieve intelligent analysis and prediction of actual current signals. By leveraging the network's extracted characteristic parameters and real-time action value assessment, it outputs optimal mutation prediction behavior for the current state. This mechanism improves the timeliness and accuracy of current mutation risk prediction, facilitates early warning and risk stratification in monitoring systems, and provides strong technical support for the safe and stable operation of power systems.
[0143] Example 1:
[0144] To verify the feasibility of the present invention in practice, the present invention was applied to a current monitoring scenario in a typical industrial power system. The system covers multiple high-voltage distribution cabinets and low-voltage feeder circuits, and is often accompanied by varying degrees of current fluctuations and sudden load changes during operation. Traditional current anomaly detection methods mainly rely on fixed threshold strategies. When the load type, operating temperature, or power supply disturbance changes, they often suffer from high misjudgment rates and delayed mutation response. To address this problem, the reinforcement learning-based current mutation risk prediction system proposed in the present invention was deployed on-site to perform continuous real-time monitoring of key circuits and intelligent risk warnings.
[0145] In actual application, the system collects current signal sequences from each monitoring loop and constructs a current state vector based on rate of change, first-order difference, window range, and short-term energy. This state vector is then fed into a trained improved dual-deep Q-network for inference processing. The model outputs the optimal mutation risk prediction for the current state. The system then maps the meaning of the output action to generate the corresponding warning level and expected mutation time information. After deployment, the system operates stably, processing an average of 20 real-time input signals per second with a response latency of less than 50 milliseconds, meeting the application requirements of high-frequency scenarios.
[0146] In order to compare the recognition effect of the method of the present invention with that of the traditional method, three typical high-voltage circuits were selected as monitoring objects, and their abnormal detection results for one week were recorded respectively, including key indicators such as prediction accuracy, advance recognition time, false alarm rate and missed alarm rate. A total of 43 actual current mutation events were recorded during the test. The traditional threshold method successfully identified 28 times, with 15 missed alarms, and most of them triggered alarms after the mutation occurred; while the system of the present invention successfully predicted 41 mutation events, of which 36 output risk predictions in advance within 0.3 to 1.2 seconds before the mutation, significantly improving the pre-warning effectiveness. At the same time, the false alarm rate of the system of the present invention in the absence of abnormalities is controlled at 3.1%, which is much lower than the 11.4% of the traditional method, and is particularly evident during periods of frequent current fluctuations.
[0147] The following table compares the key performance of different methods under the same operating environment during the test:
[0148] Table 1 Performance comparison data of the method of the present invention and the traditional current detection method
[0149]
[0150] As can be seen from the above table, the present invention has significantly improved multiple key performance indicators of current mutation risk prediction compared to traditional current anomaly detection methods. First, in terms of prediction accuracy, the system of the present invention identified a total of 41 mutation events in three typical monitoring loops, with an overall recognition rate of 95.3%, while the traditional method successfully identified only 28 mutation events, with an accuracy rate of 65.1%, an improvement of more than 30 percentage points. Looking further, in the three loops A, B, and C, the system of the present invention successfully predicted 14, 13, and 14 mutation events, respectively, of which loop B achieved a 100% prediction success rate, demonstrating excellent model generalization capabilities.
[0151] Secondly, in terms of the timeliness of early warning responses, the system of the present invention can output prediction results within 0.3 to 1.2 seconds before the current mutation occurs, with an average advance warning time of 0.83 seconds for the three circuits. Traditional methods, on the other hand, mostly respond after the mutation occurs, with an average delay of only 0.16 seconds. This difference demonstrates that the present invention can achieve more proactive risk control and enhance the system's accident prevention capabilities, making it particularly suitable for scenarios with rapid load changes or high risk of fault propagation.
[0152] Furthermore, in terms of false alarm control capabilities, the system of the present invention achieved false alarm rates of 2.7%, 3.8%, and 2.8% in the three loops, respectively, with an overall average of 3.1%, significantly outperforming the 11.4% average false alarm rate of traditional methods. Traditional methods often falsely trigger alarms when no actual mutations occur due to unreasonable threshold settings or drastic signal fluctuations, leading to misjudgments and waste of resources by operations and maintenance personnel. However, the present invention effectively reduces misjudgments caused by atypical fluctuations by introducing multidimensional feature modeling and reinforcement learning strategy tuning.
[0153] From the above performance indicators, it can be seen that the present invention not only significantly improves the recognition accuracy and response speed of mutation risks, but also greatly reduces the probability of false alarms by improving the dual deep Q network structure, dynamic reward mechanism and current feature fusion modeling. It has stronger robustness and adaptability, and can better meet the dual requirements of actual power systems for safety and intelligence levels.
[0154] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A current mutation risk prediction method based on reinforcement learning, characterized in that: The steps include: S1. Collect the current signal sequence of the monitored electrical equipment and perform feature extraction operations to construct a current state vector; S2. Input the current state vector into an improved dual deep Q network, perform forward reasoning operations, and generate an action value prediction result in the current state. The improved dual deep Q network includes a main Q network and a target Q network with the same structure. The main Q network includes a shared feature extraction subnetwork, a first action value evaluation subnetwork, a second action value evaluation subnetwork, and a Q value fusion module. S3. Based on the action value prediction results, the ε-greedy strategy is used to perform action selection operations and generate the selected action; S4. Construct a reward function to generate a corresponding reward value based on whether the current mutation is successfully predicted, the prediction deviation time, and the preset penalty factor; S5. Record the current state vector, the selected action, the reward value, and the next current state vector generated by the environment feedback to form an experience quadruple. S6. Based on the experience quadruple, use the target Q network to calculate the target action value, update the parameters of the main Q network, and periodically synchronize the parameters of the main Q network to the target Q network; S7. After completing the training optimization, the improved dual deep Q network is deployed to the current monitoring system, which receives the real-time input current state vector sequence and outputs the current mutation risk prediction result.
2. The method for predicting current mutation risk based on reinforcement learning according to claim 1, characterized in that: The current state vector includes an original current sampling sequence consisting of N consecutive sampling points before the current moment and a multidimensional current state feature set; the multidimensional current state feature set includes a current change rate feature, a current first-order difference mean, a window range feature, and a short-time energy feature.
3. The method for predicting current mutation risk based on reinforcement learning according to claim 1, characterized in that: The S2 specifically includes: S21, constructing a main Q network of the improved dual deep Q network, wherein the main Q network includes a shared feature extraction subnetwork, a first action value evaluation subnetwork, a second action value evaluation subnetwork and a Q value fusion module; S22, inputting the current state vector into the shared feature extraction sub-network, performing a nonlinear feature mapping operation, and generating a unified feature representation vector, wherein the nonlinear mapping process is performed according to the weight parameters of the shared feature extraction sub-network; S23. Input the unified feature representation vector into the first action value evaluation subnetwork, perform an action value evaluation operation based on steady-state conditions, and output a corresponding first action value vector, where the first action value vector is generated by the first action value evaluation subnetwork based on the feature representation vector and internal weight parameters; S24, inputting the unified feature representation vector into the second action value evaluation subnetwork, performing an action value evaluation operation based on the mutation trend condition, and outputting a corresponding second action value vector, where the second action value vector is generated by the second action value evaluation subnetwork based on the feature representation vector and internal weight parameters; S25, inputting the first action value vector and the second action value vector into the Q value fusion module, performing a weighted fusion operation on the action value vectors, and generating an action value prediction result in the current state; S26. Construct a target Q network of the improved dual deep Q network, wherein the target Q network structure is consistent with the main Q network structure, adopts a fixed parameter set, and maintains update consistency with the main Q network through a periodic synchronization mechanism to generate a target action value.
4. The method for predicting current mutation risk based on reinforcement learning according to claim 3, characterized in that: The S25 specifically includes: S251. Obtain a first action value vector output by a first action value evaluation subnetwork, where the first action value vector represents an action value distribution under a steady-state response; S252: Obtain a second action value vector output by the second action value evaluation subnetwork, where the second action value vector represents the action value distribution under the current mutation trend; S253. In the Q-value fusion module, a fusion weight factor α is set to adjust the relative contribution of the first action value vector and the second action value vector in the fusion process, where α represents the steady-state condition weight and (1-α) represents the mutation condition weight. S254. Perform linear weighted operations on the first action value vector and the second action value vector according to the weighted fusion strategy, fuse them to generate the final action value prediction result vector in the current state, multiply the first action value vector by the fusion weight factor α, multiply the second action value vector by (1-α), and sum the two to obtain the fused action value prediction result.
5. The method for predicting current mutation risk based on reinforcement learning according to claim 3, characterized in that: The S26 specifically includes: S261: Construct a target Q network that is consistent with the main Q network structure. The target Q network includes the same shared feature extraction subnetwork, first action value evaluation subnetwork, second action value evaluation subnetwork, and Q value fusion module as the main Q network. All network layers of the target Q network use independent parameter sets W. t To define; S262: Input the current state vector in the training sample into the target Q network, pass it through the shared feature extraction subnetwork, two action value evaluation subnetworks and the Q value fusion module in sequence, and output the target action value prediction result; S263, set parameter synchronization period T sync , controls the update frequency of the target Q network parameters; S264: After each parameter synchronization cycle, perform a parameter copy operation to assign all parameter values in the current main Q network to the target Q network.
6. The method for predicting current mutation risk based on reinforcement learning according to claim 1, characterized in that: The S3 specifically includes: S31, define action set A={a1, a2,…, a n }, represents all candidate actions that can be selected under the current current state, and the action set is a pre-set finite set that includes all legal control operations; S32, receiving the action value prediction result vector under the current current state and setting the action selection parameter ε; S33. Generate a random number r in the interval [0,1] and compare it with ε. If r < ε, execute the exploration strategy and randomly select an action a from the action set. rand as the currently selected action; S34. If r ≥ ε, execute the utilization strategy and select the action corresponding to the maximum value of the action value prediction result from the action set as the currently selected action; S35. Output selected action a select , and record the state-action pair consisting of the current state vector and the selected action.
7. The method for predicting current mutation risk based on reinforcement learning according to claim 1, characterized in that: The S4 specifically includes: S41. Set the reward function based on the selected action a select Whether the current mutation event is successfully predicted to construct the initial reward signal, if the prediction is successful, a positive reward value r is given pos , if the prediction fails, a negative reward value r is given neg ; S42, define the prediction deviation time ΔT=|T pred -T actual |, where T pred The current mutation time predicted for the selected action, T actual is the actual time when the current mutation occurs; S43. Set a penalty factor λ, and use the success or failure of the prediction of the mutation event and the prediction time deviation as the basis for reward calculation; S44, if the selected action is a select Successfully predicting a current mutation event will result in a positive reward value r pos Based on, deduct the product of the predicted time deviation and the penalty factor to get the immediate reward value. If the selected action is a select If the prediction fails, a negative reward value r neg Based on the prediction time, the product of the penalty factor is deducted to obtain the immediate reward value.
8. The method for predicting current mutation risk based on reinforcement learning according to claim 1, characterized in that: The S5 specifically includes: S51. After each action is executed, collect the current state vector s at the current moment t and the selected action a select , where t represents the current time step; S52, receiving the immediate reward value R(s) generated by environmental feedback t ,a select ), collect the next moment current state vector s generated by the environment feedback after the current action is executed t+1 ; S53, combining the current state vector at the current moment, the selected action, the immediate reward value, and the current state vector at the next moment to form an experience quadruple; S54. The generated experience quadruple is stored in an experience replay pool. The experience replay pool adopts a first-in-first-out cache mechanism to maintain a fixed-capacity sample set.
9. The method for predicting current mutation risk based on reinforcement learning according to claim 1, characterized in that: The S6 specifically includes: S61, randomly sample a batch of experience quadruple samples from the experience playback buffer, each sample is represented by (s t ,a select ,R(s t ,a select ),r t ); S62, the current state vector s at the next moment t+1 Input to the target Q network, perform forward reasoning operations, and obtain all candidate actions in s t+1 The target action value prediction result vector in the state and select the target Q value corresponding to the maximum value: Among them, y t is the target action value, R(s,a select ) is the immediate reward value, s t is the current state vector at the previous moment, a select is the selected action, γ is the discount factor, a i For each action in the action set, Q target (s t+1 ,a i ) represents the state s t+1 The target Q network responds to action a i The value prediction results; S63, the current state vector s at the current moment t Input into the main Q network, calculate the current action value prediction result vector of all candidate actions, and take the currently selected action a t The corresponding action value prediction result; S64, construct an error function between the current action value prediction result and the target action value, the error function adopts the mean square error calculation method to convert the target action value y t With the main Q network for the currently selected action a t The difference between the corresponding action value prediction results is squared; S65, performing an error back propagation operation on the main Q network, and updating the parameter sets of each layer in the main Q network based on the loss function; S66, when the set parameter synchronization period T is met sync Under these conditions, all parameter sets in the current master Q network are copied to the target Q network.
10. The method for predicting current mutation risk based on reinforcement learning according to claim 1, characterized in that: The S7 specifically includes: S71. After completing the training process of the improved dual deep Q network, solidifying the final parameters of the shared feature extraction subnetwork, the first action value evaluation subnetwork, the second action value evaluation subnetwork, and the Q value fusion module in the main Q network; S72, deploying the trained improved dual deep Q network into the current monitoring system, establishing a data receiving interface with the current acquisition terminal, and receiving the current signal collected by the sensor in real time; S73, pre-process the real-time input current signal and construct the real-time current state vector s real,t , the real-time current state vector includes a current amplitude sequence and a feature set of multiple sampling points in the current window; S74, the real-time current state vector s real,t Input into the main Q network structure of the improved dual deep Q network after training, perform forward reasoning operations, and output the corresponding real-time action value prediction vector; S75. Based on the real-time action value prediction vector, select the action a corresponding to the maximum action value predict,t , represents the optimal mutation prediction behavior at the current moment; S76, the selected action a predict,t The corresponding action meaning is mapped into the current mutation risk prediction result, which is output to the early warning module of the current monitoring system. The prediction result includes the mutation risk level, expected occurrence time and response suggestions.