Power transmission and distribution facility vulnerability insurance prediction method based on deep reinforcement learning
By applying deep reinforcement learning methods in the damage rate prediction of power transmission and distribution facilities, combined with Dueling DQN and Double DQN architectures, the problem of traditional models being difficult to capture dynamic changes and lack of environmental feedback is solved, and higher prediction accuracy and reliability are achieved, supporting insurance companies' more accurate decision-making and reducing operational risks.
Patent Information
- Application Number
- CN202510339782.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional machine learning regression models are difficult to capture the dynamic trends in time series data when predicting the damage rate of power transmission and distribution facilities, and lack learning mechanisms for environmental feedback, resulting in low prediction accuracy and unreliability.
Using a deep reinforcement learning method, model training and real-time prediction are carried out by establishing a reinforcement learning network model, including environmental modules, estimation networks, target networks, experience pools and loss functions. This method combines Dueling DQN and Double DQN architectures to realize the learning and evaluation of action value of environmental feedback.
It realizes a learning mechanism for environmental feedback, improves the accuracy and reliability of damage prediction of power transmission and distribution facilities, helps insurance companies to make more accurate decisions in risk assessment, dynamic rate formulation and claims management, and reduces claims costs and operational risks.
Smart Images

Figure CN120198079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of risk prediction. Specifically, it relates to a vulnerability insurance prediction method for power transmission and distribution facilities based on deep reinforcement learning. Background Art
[0002] Accurately predicting the damage rate of power transmission and distribution facilities in typhoon-prone areas is not only helpful for preventing potential safety hazards, optimizing maintenance plans, and reducing operating costs, but also crucial for the damage risk assessment of insurance businesses. Traditional insurance claims settlement methods are mostly post-event processing, lacking risk anticipation, resulting in high claims settlement costs and inaccurate risk assessments. Therefore, by predicting the vulnerability of facilities through advanced technologies, insurance companies can identify risks in advance, optimize resource allocation, reasonably price insurance rates, reduce claim risks, and develop more targeted insurance products.
[0003] Although traditional machine learning regression models (such as support vector machines, random forests, neural networks, etc.) have achieved certain results in predicting the damage rate of materials, they still face many limitations when applied in the insurance business. Because these methods mainly rely on the mapping relationship between static features in historical data and output targets, it is difficult to capture the dynamic change trends in time series data. The damage of power transmission and distribution facilities is affected by multiple dynamic factors, such as wind speed, temperature, humidity, equipment load changes, etc. These factors change over time, resulting in a complex non-linear degradation characteristic in the damage process of facilities. Static models are difficult to effectively simulate this dynamic process, leading to low prediction accuracy. In addition, traditional machine learning models lack a learning mechanism for environmental feedback and cannot be adjusted in real time to adapt to new working conditions, thus further affecting the prediction accuracy and reliability. These limitations significantly affect the decision-making accuracy and operational efficiency of insurance companies in core businesses such as risk assessment, dynamic rate setting, and efficient claims settlement management. Summary of the Invention
[0004] Aiming at the problem that the existing machine learning regression models lack a learning mechanism for environmental feedback, which affects the decision-making accuracy of insurance companies in core businesses such as risk assessment, dynamic rate setting, and efficient claims settlement management, the present invention provides a vulnerability insurance prediction method for power transmission and distribution facilities based on deep reinforcement learning.
[0005] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:
[0006] A vulnerability insurance prediction method for power transmission and distribution facilities based on deep reinforcement learning, including the steps of:
[0007] Collect data related to the vulnerability of power transmission and distribution facilities to form a data set;
[0008] Build a reinforcement learning network model, which includes an environment module, an estimation network, a target network, an experience pool, and a loss function;
[0009] Input the collected dataset into the reinforcement learning network model for model training;
[0010] Input the real-time collected data related to the vulnerability of power transmission and distribution facilities into the trained reinforcement learning network model for predicting the damage of power transmission and distribution facilities and outputting the prediction results.
[0011] Furthermore, the reinforcement learning network model adopts a network architecture combining Dueling DQN and Double DQN;
[0012] Double DQN includes an estimation network and a target network. The estimation network is used to select actions, and the target network is used to evaluate the value of the selected actions. The estimation network calculates the Q-value of each action according to the current state s and all possible actions a, and selects the action with the highest Q-value; the parameters of the target network are copied from the estimation network regularly but not updated frequently to ensure its stability. This method reduces the fluctuation of the target Q-value and the risk of overestimation.
[0013] Furthermore, the training process of the reinforcement learning network model includes the steps:
[0014] First, initialize the parameters of the estimation network and the target network, and initialize the experience pool;
[0015] Environment simulation: Collect the current state s and the executed action a, as well as the corresponding reward r and the next state s' from the environment, and store the data in the experience pool;
[0016] In the network training stage, randomly extract a dataset from the experience pool for training;
[0017] Use the estimation network to predict the action value Q(s,a;θ) in the current state, and at the same time use the target network to generate the target Q-value Q'(s',a';θ');
[0018] Calculate the loss function and update the parameters of the estimation network through backpropagation;
[0019] Select the optimal action a* = argmax(Q(s',a;θ)) according to the action value predicted by the estimation network, execute the selected action and interact with the environment to obtain a new state and reward.
[0020] Furthermore, the calculation formula of the target Q-value:
[0021] Target Q = r + γ × Q'(s',a*;θ'),
[0022] where r is the immediate reward, γ is the discount factor, and Q'(s', a*; θ') is the optimal action value of the next state predicted by the target network. By using two independent networks to evaluate the current policy and update the weights respectively, the instability phenomenon that may occur in the evaluation and update process of a single network is avoided, the stability of the model is improved, and the overestimation problem caused by the traditional Q-learning method is also overcome.
[0023] Furthermore, the calculation formula of the loss function:
[0024]
[0025] where N represents the number of samples used to calculate the loss function in the current batch, and each sample corresponds to a state-action pair; Target Q is the target Q value, which is calculated according to the target network and used to evaluate the value of the selected action; PredictedQ is the predicted Q value, which is calculated by the estimation network and used to select the optimal action.
[0026] The standard mean squared error is adopted as the loss function. It is simple and effective, which can directly reflect the gap between the predicted value and the true value, and prompt the model to converge to the optimal solution or approach the optimal solution faster. At the same time, it has a certain robustness to outliers, which helps to improve the generalization ability of the model.
[0027] Furthermore, the environment simulation acquisition is responsible for simulating the actual working environment, including the wind speed state wind_state (vector) and the material state material_one_hot (vector); the environment module receives the current state s and the executed action a, and then according to the action a and the current state s, the environment module simulates the environmental change, generates a new state s' and a reward r, and finally the new state s' and the reward r are returned to the estimation network and the target network.
[0028] Furthermore, the estimation network evaluates the value of each action in the current state, receives the current state s, and the shared layer extracts the state features. The merging layer then merges the state value V(s) and the action advantage A(s, a) to generate the final action value Q(s, a).
[0029] Furthermore, the target network is used to generate the target Q value, receives the next state s', and the shared layer here has the same structure as the estimation network, both of which are used to extract the state features, while the merging layer generates the target Q value Q'(s', a'; θ').
[0030] Furthermore, the experience pool stores data on historical states, actions, rewards, and next states. The experience pool receives the current state s, the executed action a, the reward r, and the next state s'. The data is stored in the experience pool, and during the training process, the experience pool randomly samples some data for training. It not only improves the stability of learning but also enhances the learning efficiency, which is crucial for building an efficient and reliable reinforcement learning system.
[0031] Furthermore, the loss function is used to optimize the network parameters and improve the prediction accuracy. The loss function receives the predicted Q value and the target Q value, calculates the loss value Loss, and the loss value Loss is used for backpropagation to update the network parameters.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] By collecting the vulnerability data of power transmission and distribution facilities, establishing a reinforcement learning network model, training the model, and inputting the real-time collected vulnerability data of power transmission and distribution facilities into the trained reinforcement learning network model for predicting the damage of power transmission and distribution facilities, a learning mechanism for environmental feedback is realized, ensuring the decision-making accuracy of insurance companies in the core businesses of risk assessment, dynamic rate setting, and efficient claims management.
[0034] The evaluation results on the test set show that using the mean square error (MSE) as the evaluation index, the present invention achieves a lower mean square error on the test set, indicating that the model can accurately predict the pole damage rate under different wind speed conditions. This performance helps the insurance business to more effectively evaluate potential losses, optimize the design of insurance products and claims decisions.
[0035] Traditional statistical methods or simple machine learning models usually have difficulty in capturing complex non-linear relationships, while the present invention utilizes the powerful representation ability of deep reinforcement learning, combines the Dueling DQN and Double DQN architectures, and effectively improves the prediction accuracy. Especially when dealing with multi-dimensional inputs (such as wind speed, material type, etc.), it shows obvious advantages.
[0036] By accurately predicting the facility damage rate, more accurate risk assessment data is provided for insurance institutions, thus supporting them to formulate more scientific and reasonable claims plans in advance, avoiding hasty responses and resource waste in claims after equipment losses, and reducing the claims cost. At the same time, the intelligent prediction system also supports insurance institutions to optimize the pricing and design of insurance products, improve the competitiveness of insurance products, and reduce operational risks. In addition, the system reduces manual intervention, avoids human errors, improves business reliability, and reduces labor costs and management costs. Description of the Drawings
[0037] Figure 1This is the overall flowchart of a vulnerability insurance prediction method for power transmission and distribution facilities based on deep reinforcement learning in an embodiment of the present invention;
[0038] Figure 2 This is the overall structural block diagram of the reinforcement learning network model in an embodiment of the present invention. Detailed implementation manners
[0039] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with embodiments and accompanying drawings. The content mentioned in the implementation manners does not limit the present invention.
[0040] As Figure 1 shown, this embodiment provides a vulnerability insurance prediction method for power transmission and distribution facilities based on deep reinforcement learning, including the steps of:
[0041] S1. Collect data related to the vulnerability of power transmission and distribution facilities to form a data set;
[0042] S2. Establish a reinforcement learning network model, where the reinforcement learning network model includes an environment module, an estimation network, a target network, an experience pool, and a loss function;
[0043] S3. Input the collected data set into the reinforcement learning network model for model training;
[0044] S4. Input the data related to the vulnerability of power transmission and distribution facilities collected in real time into the trained reinforcement learning network model for predicting the damage of power transmission and distribution facilities and outputting the prediction result.
[0045] As Figure 2 shown, the reinforcement learning network model adopts a network architecture that combines Dueling DQN and Double DQN;
[0046] Double DQN includes an estimation network and a target network. The estimation network is used to select actions, and the target network is used to evaluate the value of the selected actions. The estimation network calculates the Q value of each action according to the current state s and all possible actions a, and selects the action with the highest Q value; the parameters of the target network are copied from the estimation network regularly but not updated frequently to ensure its stability. This method reduces the fluctuation of the target Q value and the risk of overestimation.
[0047] The training process of the reinforcement learning network model includes the steps of:
[0048] First, initialize the parameters of the estimation network and the target network, and initialize the experience pool;
[0049] Environmental simulation: Collect the current state s and the executed action a, as well as the corresponding reward r and the next state s' from the environment, and store the data in the experience pool;
[0050] During the network training phase, a dataset is randomly sampled from the experience pool for training;
[0051] The estimation network is used to predict the action value Q(s,a;θ) in the current state, while the target network is used to generate the target Q value Q'(s',a';θ');
[0052] Calculate the loss function and update the parameters of the estimation network through backpropagation;
[0053] Select the optimal action a* = argmax(Q(s',a;θ)) according to the action value predicted by the estimation network, execute the selected action and interact with the environment to obtain a new state and reward.
[0054] The calculation formula for the target Q value:
[0055] Target Q = r + γ × Q'(s',a*;θ'),
[0056] where r is the immediate reward, γ is the discount factor, and Q'(s',a*;θ') is the optimal action value of the next state predicted by the target network. By using two independent networks to evaluate the current policy and update the weights respectively, the instability phenomenon that may occur in the evaluation and update process of a single network is avoided, the stability of the model is improved, and the overestimation problem caused by the traditional Q-learning method is also overcome.
[0057] The calculation formula for the loss function:
[0058]
[0059] where N represents the number of samples used to calculate the loss function in the current batch, and each sample corresponds to a state-action pair; Target Q is the target Q value, which is calculated according to the target network and used to evaluate the value of the selected action; PredictedQ is the predicted Q value, which is calculated by the estimation network and used to select the optimal action. It is simple and effective, and can directly reflect the gap between the predicted value and the true value, prompting the model to converge to the optimal solution or close to the optimal solution faster. At the same time, it has a certain robustness to outliers, which helps to improve the generalization ability of the model.
[0060] The environmental simulation acquisition is responsible for simulating the actual working environment, including the wind speed state wind_state (vector) and the material state material_one_hot (vector); the environmental module receives the current state s and the executed action a, and then according to the action a and the current state s, the environmental module simulates the environmental change, generates a new state s' and a reward r, and finally the new state s' and the reward r are returned to the estimation network and the target network.
[0061] The estimation network evaluates the value of each action in the current state, receives the current state s, where the shared layer extracts state features. The merging layer then combines the state value V(s) and the action advantage A(s,a) to generate the final action value Q(s,a).
[0062] The target network is used to generate the target Q value, receives the next state s', where the shared layer has the same structure as the estimation network and is used to extract state features, while the merging layer generates the target Q value Q'(s',a';θ').
[0063] The experience pool stores data on historical states, actions, rewards, and next states. The experience pool receives the current state s, the executed action a, the reward r, and the next state s', and the data is stored in the experience pool. During the training process, the experience pool randomly samples some data for training. It not only improves the stability of learning but also enhances the learning efficiency, which is crucial for building an efficient and reliable reinforcement learning system.
[0064] The loss function is used to optimize the network parameters and improve the prediction accuracy. The loss function receives the predicted Q value and the target Q value, calculates the loss value Loss, and the loss value Loss is used for backpropagation to update the network parameters.
[0065] Compared with the prior art, the present invention has the following beneficial effects:
[0066] Through the collection of vulnerability data of power transmission and distribution facilities, the establishment of a reinforcement learning network model, model training, and the input of real-time collected vulnerability-related data of power transmission and distribution facilities into the trained reinforcement learning network model for damage prediction of power transmission and distribution facilities, a learning mechanism for environmental feedback is realized, ensuring the decision-making accuracy of insurance companies in core operations such as risk assessment, dynamic premium rate setting, and efficient claims management.
[0067] The evaluation results on the test set show that using the mean squared error (MSE) as the evaluation index, the present invention achieves a low mean squared error on the test set, indicating that the model can accurately predict the pole damage rate under different wind speed conditions. This performance helps the insurance business to more effectively evaluate potential losses and optimize the design of insurance products and claims decisions.
[0068] Traditional statistical methods or simple machine learning models usually have difficulty capturing complex non-linear relationships, while the present invention utilizes the powerful representation ability of deep reinforcement learning, combines the Dueling DQN and Double DQN architectures, and effectively improves the prediction accuracy. Especially when dealing with multi-dimensional inputs (such as wind speed, material type, etc.), it shows obvious advantages.
[0069] By accurately predicting the facility damage rate, it provides more accurate risk assessment data for insurance institutions, thus supporting them to formulate more scientific and reasonable claims settlement plans in advance, avoiding hasty responses and resource waste in claims settlement after equipment losses, and reducing claims settlement costs. At the same time, the intelligent prediction system also supports insurance institutions to optimize the pricing and design of insurance products, improve the competitiveness of insurance products, and reduce operational risks. In addition, the system reduces manual intervention, avoids human errors, improves business reliability, and reduces labor costs and management costs.
[0070] The above has introduced in detail a vulnerability insurance prediction method for power transmission and distribution facilities based on deep reinforcement learning provided by this application. The description of specific embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning, characterized in that: Includes steps: Collect data related to the vulnerability of power transmission and distribution facilities to form a data set; Establish a reinforcement learning network model, which includes an environment module, an estimation network, a target network, an experience pool, and a loss function; Input the acquired data set into the reinforcement learning network model to perform model training; The real-time collected data related to the vulnerability of power transmission and distribution facilities are input into the trained reinforcement learning network model to predict the damage of power transmission and distribution facilities and output the prediction results.
2. According to claim 1, a method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning is characterized in that: The reinforcement learning network model adopts a network architecture that combines DuelingDQN and DoubleDQN; DoubleDQN consists of an estimation network and a target network. The estimation network is used to select actions, and the target network is used to evaluate the value of the selected actions.
3. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 2, characterized in that: The reinforcement learning network model training process includes the following steps: First, initialize the parameters of the estimation network and the target network, and initialize the experience pool; Environment simulation: collect the current state s and the executed action a from the environment, as well as the corresponding reward r and next state s', and store the data in the experience pool; During the network training phase, data sets are randomly selected from the experience pool for training; The estimation network is used to predict the action value Q(s,a;θ) in the current state, and the target network is used to generate the target Q value Q'(s',a'; θ'); Calculate the loss function and update the parameters of the estimation network through back propagation; The optimal action a*=argmax(Q(s',a;θ)) is selected based on the action value predicted by the estimation network, and the selected action is executed and interacts with the environment to obtain a new state and reward.
4. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 3, characterized in that: The calculation formula of target Q value is: Target Q=r+γ×Q'(s',a*;θ'), where r is the immediate reward, γ is the discount factor, and Q'(s',a*;θ') is the best action value for the next state predicted by the target network.
5. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 4, characterized in that: The calculation formula of the loss function is: Where N represents the number of samples used to calculate the loss function in the current batch, and each sample corresponds to a state-action pair; Target Q is the target Q value, which is calculated according to the target network and is used to evaluate the value of the selected action; Predicted Q is the predicted Q value, which is calculated by the estimation network and is used to select the optimal action.
6. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 5, characterized in that: Environmental simulation acquisition is responsible for simulating the actual working environment, including wind speed status and material status; the environmental module receives the current status and the executed action, and then based on the action and current status, the environmental module simulates environmental changes and generates new status and rewards. Finally, the new status and rewards are returned to the estimation network and the target network.
7. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 6, characterized in that: The estimation network evaluates the value of each action in the current state and receives the current state. The shared layer extracts state features, and the merging layer merges the state value and action advantage to generate the final action value.
8. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 7, characterized in that: The target network is used to generate the target Q value and receive the next state. The shared layer here shares the same structure as the estimation network, both of which are used to extract state features, while the merging layer generates the target Q value.
9. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 8, characterized in that: The experience pool stores data on historical states, actions, rewards, and the next state. The experience pool receives the current state, executed actions, rewards, and the next state. The data is stored in the experience pool. During the training process, the experience pool randomly extracts part of the data for training.
10. A method for predicting the vulnerability of power transmission and distribution facilities based on deep reinforcement learning according to claim 9, characterized in that: The loss function is used to optimize network parameters and improve prediction accuracy. The loss function receives the predicted Q value and the target Q value, calculates the loss value Loss, and the loss value Loss is used for back propagation to update network parameters.