An economical operation and maintenance method for distribution transformers based on reinforcement learning

Through a multi-dimensional feature extraction and state prediction of distribution transformers based on reinforcement learning, the maintenance strategy is optimized, and the problems of difficult to predict and cope with the state decay and fault of distribution transformers in the existing technology are solved, and efficient, economical and safe distribution transformer maintenance is achieved.

CN116128153BActive Publication Date: 2025-05-16JICHANG ELECTRIC GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310243797.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-14
Publication Date
2025-05-16
Estimated Expiration
2043-03-14

AI Technical Summary

Technical Problem

The existing distribution transformer maintenance strategies are difficult to effectively predict and deal with the complex dynamic relationship between the state decay of the distribution transformer and the fault, resulting in high maintenance costs, low reliability and latent safety hazards.

Method used

Using reinforcement learning-based method, statistical analysis of the historical abnormal operating conditions and fault records of the distribution transformer is performed, multi-dimensional features are extracted, state prediction is performed by combining LSTM and SVM models, and agents are trained through DQN algorithm to optimize maintenance strategies.

Benefits of technology

It realizes efficient mining of the complex and dynamic relationship between the state decay of the distribution transformer and the fault, optimizes the maintenance strategy, reduces maintenance costs, and improves the reliability and safety of the distribution system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116128153B_ABST
    Figure CN116128153B_ABST
Patent Text Reader

Abstract

The present invention proposes an economical operation and maintenance method for distribution transformers based on reinforcement learning, and the main steps are: 1) extracting multidimensional features of distribution transformers from online operation data and basic information of distribution transformers; 2) establishing a data-driven distribution transformer operation status evaluation and situation prediction model based on multidimensional perception of the operation status of distribution transformers; 3) training the reinforcement learning process of the interaction between intelligent agents and distribution transformers through the DQN algorithm to obtain a distribution transformer maintenance strategy optimization model; 4) constructing the predicted results of the operation status of the distribution transformer into state information and inputting it into the distribution transformer maintenance strategy optimization model to obtain a predictive maintenance decision sequence. The present invention has good versatility and applicability, is applicable to oil-immersed distribution transformers and dry-type distribution transformers, can obtain a predictive maintenance decision sequence, and provide guidance for distribution operation and maintenance personnel to implement more objective and accurate active maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart distribution networks, and in particular to an economical operation and maintenance method for distribution transformers based on reinforcement learning. Background Art

[0002] Distribution transformers are power transformers with voltage levels below 35kV. As key equipment in the distribution network, distribution transformers directly serve end users, performing voltage conversion and power transmission. They are diverse, numerous, and widely distributed, and are closely linked to the safe and stable operation of the distribution network. Distribution transformers can degrade due to wear, aging, fluctuating operating conditions, and other factors, leading to potential failures. This can reduce the reliability of the distribution network and cause widespread power outages within the distribution network, impacting residents' daily electricity needs and business operations, and even causing serious economic losses and safety incidents.

[0003] Currently, maintenance and management models for distribution transformers are primarily categorized as post-failure maintenance, scheduled maintenance, condition-based maintenance, and predictive maintenance. Post-failure maintenance involves shutting down distribution transformers for repair when a fault occurs. This maintenance strategy can result in unexpected economic losses and emergency maintenance costs. Scheduled maintenance involves regularly inspecting distribution transformers, but this can lead to unnecessary maintenance activities, resulting in increased maintenance costs and economical downtime. Post-failure and scheduled maintenance are simple maintenance strategies that fail to ensure reliable and economical power system operation. Condition-based maintenance focuses on online monitoring and assessment of distribution transformer conditions, using real-time status information as the basis for maintenance decisions. Maintenance activities are only implemented when there is evidence that the distribution transformer requires maintenance. This effectively reduces unnecessary maintenance activities and offers better economics and safety than post-failure and scheduled maintenance. However, condition-based maintenance fails to consider the changing state of the distribution transformer. A distribution transformer may experience degradation in the future without failing, posing a potential safety hazard. Predictive maintenance focuses on analyzing historical and online operating data of distribution transformers to establish a distribution transformer state prediction model. This model can proactively detect changing trends in the distribution transformer's state. The predicted results of the distribution transformer's state decay trend serve as a decision support for proactive maintenance, reducing catastrophic failures and improving distribution transformer safety. However, existing predictive maintenance methods fail to account for the uncertainty inherent in the complex dynamic relationship between distribution transformer state decay and failure. When a distribution transformer's future operating state decays, the probability of failure is uncertain. The distribution transformer may continue to operate for an extended period before failing. Premature maintenance of the distribution transformer disrupts the stable operation of the distribution system, increasing downtime and maintenance costs, and resulting in unnecessary economic losses. Failure to perform timely maintenance on the distribution transformer can lead to transformer failure, reducing the reliability of the distribution system, causing serious safety incidents, and sudden economic losses.

[0004] Reinforcement learning provides a new research idea for the exploration of intelligent maintenance of distribution transformers. Considering that the reinforcement learning model has strong learning and adaptability, it is introduced into the research of economic operation and maintenance methods of distribution transformers. Through the reinforcement learning theory, the uncertainty of the complex dynamic relationship between the state decay and failure of distribution transformers is fully explored, making the maintenance strategy efficient, intelligent and adaptable, providing guidance for distribution operation and maintenance personnel to implement more objective and accurate maintenance actions, thereby ensuring the safe, economical and stable operation of distribution transformers. Summary of the Invention

[0005] The purpose of the present invention is to solve the problems existing in the prior art.

[0006] To achieve the purpose of the present invention, the present invention proposes an economic operation and maintenance method for distribution transformers based on reinforcement learning, such as Figure 1 As shown, it mainly includes the following steps:

[0007] 1) Conduct statistical analysis on the historical abnormal operating conditions and fault records of distribution transformers, summarize the common fault types and main causes of distribution transformer faults, and extract real-time, statistical, and basic multidimensional features from the online operating data of distribution transformer voltage and current and basic information to achieve multidimensional perception of the operating status of distribution transformers, providing a data basis for distribution transformer operating status assessment and situation prediction, and distribution transformer maintenance strategy optimization.

[0008] 2) Based on the multi-dimensional perception of the operating status of distribution transformers, an improved comprehensive evaluation method is used to realize real-time evaluation of the operating status of distribution transformers. The long short-term memory neural network (LSTM) and support vector machine (SVM) model are combined to realize the prediction of the operating status of distribution transformers. This method can timely and accurately grasp the operating status and change trend of distribution transformers, laying the foundation for the optimization of distribution transformer maintenance strategies.

[0009] 3) By constructing reinforcement learning elements such as state, reward and maintenance action through data such as distribution transformer operating status characteristics, real-time operating status evaluation results and historical fault records, the DQN (Deep Q-Network) algorithm is used to train the reinforcement learning process of the interaction between the intelligent agent and the distribution transformer, and the uncertainty of the complex dynamic relationship between the state decay and fault occurrence of the distribution transformer is mined to obtain the distribution transformer maintenance strategy optimization model, such as Figure 2 shown.

[0010] Furthermore, the main steps of establishing the distribution transformer maintenance strategy optimization model are as follows:

[0011] 3.1) In the reinforcement learning task of optimizing the economic operation and maintenance strategy for distribution transformers, the specific settings of reinforcement learning elements such as state S, maintenance action A, and reward R are as follows:

[0012] 3.1.1) The state of the distribution transformer includes the operating state features and operating state evaluation results score of the distribution transformer. The state s at the i-th moment is i The expression is:

[0013] s i ={features,score},s i ∈S (1)

[0014] In formula (1), the operating state characteristics mainly include load rate, voltage deviation and three-phase load imbalance, average load rate, heavy load duration, overload duration, average three-phase load imbalance and three-phase load imbalance duration; S is an infinite state set, and the transition probability between states is unknown. The distribution transformer state is divided into the initial state s0, the general state s i and the terminal state s T There are three situations. The initial state refers to the first state in the interaction process between the intelligent agent and the distribution transformer. In this state, the distribution transformer operates well and the operating status evaluation result is within the normal range. The terminal state refers to the distribution transformer that has a fault and the interaction process between the intelligent agent and the distribution transformer ends. The general state refers to the state other than the initial state and the terminal state. In this state, the distribution transformer may operate normally or may have a fault.

[0015] 3.1.2) The agent has two maintenance action options. The first is to not perform maintenance on the distribution transformer. no The second is to repair the distribution transformer. repair , the action a performed by the agent at time i i The expression is:

[0016] a i ={a no or a repair},a i ∈A(2)

[0017] When the agent chooses to execute a repair When , it shows that the decision recommendation given by the economic operation and maintenance strategy optimization model is to repair the distribution transformer. repair The action can be further refined, that is, the targeted maintenance processing action is refined according to the current monitoring status of the distribution transformer, and the state of the distribution transformer is converted to the initial state after the maintenance.

[0018] 3.1.3) The reward r obtained by the agent at time i i The maintenance action currently being executed is a i and the next moment state s of the distribution transformer feedback i+1 Joint decision. If s i+1 is the terminal state, and the reward r obtained by the agent i If s is -1, the interaction process between the agent and the distribution transformer ends. i+1 For the general state, reward r i The calculation formula is:

[0019]

[0020] In formula (3), the reward r iThe value range is between 0 and 1; score is the state s at the next moment i+1 The distribution transformer operating status evaluation results in the ; loss is the maintenance action a performed by the agent in different scenarios i economic losses; reward r i The first half of the β1score is related to the safe operation of the distribution transformer, β1 is the safety target coefficient; the reward r i The second half, β2(100-loss), is related to the economic cost of distribution transformer maintenance, and β2 is the economic target coefficient.

[0021] 3.2) During the learning process of the DQN training agent interacting with the distribution transformer, the DQN training agent obtains factors such as the distribution transformer state s, optimizes the error between the approximate value function and the true value function through the gradient descent algorithm, and thus calculates the weight parameter w of the approximate value function. Then, the deep neural network is used to fit the true action value function to obtain the approximate action value function And use the Q-learning algorithm to update the approximate action value function to calculate the optimal action value function To obtain the optimal strategy In order to improve the efficiency and stability of the algorithm, the mechanism of experience replay and fixed Q target is introduced into the DQN training agent.

[0022] 3.3) In the DQN-based economic operation and maintenance process of distribution transformers, the DQN algorithm is used to train the reinforcement learning process of the interaction between the intelligent agent and the distribution transformer, thereby exploring the uncertainty of the complex dynamic relationship between the state decay and failure of the distribution transformer and obtaining an economic operation and maintenance strategy optimization model for the distribution transformer.

[0023] 4) The prediction results of the distribution transformer operating status are constructed into status information and input into the distribution transformer economic operation and maintenance optimization model. The economic operation and maintenance strategy optimization model outputs a predictive maintenance decision sequence.

[0024] In the context of the power Internet of Things, the present invention proposes an economical operation and maintenance method for distribution transformers based on reinforcement learning. This method can assist in outputting a predictive maintenance decision sequence based on the prediction results of the changing trend of the distribution transformer's operating status, and provide guidance for distribution operation and maintenance personnel to implement more objective and accurate proactive maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0026] Figure 1 Provide a technical framework diagram for optimizing maintenance strategies for distribution transformers;

[0027] Figure 2 Flowchart for optimizing maintenance strategy for distribution transformers based on DQN;

[0028] Figure 3 This is a diagram of the DQN training process;

[0029] Figure 4 This is a trend chart showing the operating status of a distribution transformer;

[0030] Figure 5 This is a maintenance status diagram of a distribution transformer; DETAILED DESCRIPTION

[0031] The present invention will be further described below with reference to the following examples, but it should not be understood that the scope of the present invention is limited to the following examples. Without departing from the above technical ideas of the present invention, various substitutions and modifications can be made according to common technical knowledge and customary means in the art, and all should be included in the scope of protection of the present invention.

[0032] 1) Conduct statistical analysis on the historical abnormal operating conditions and fault records of distribution transformers, summarize the common fault types and main causes of distribution transformer faults, and extract real-time, statistical, and basic multidimensional features from the online operating data of distribution transformer voltage and current and basic information to achieve multidimensional perception of the operating status of distribution transformers, providing a data basis for distribution transformer operating status assessment and situation prediction, and distribution transformer maintenance strategy optimization.

[0033] 2) Based on the multi-dimensional perception of the operating status of distribution transformers, an improved comprehensive evaluation method is adopted to realize real-time evaluation of the operating status of distribution transformers, and the long-short-term memory neural network and support vector machine model are combined to realize the prediction of the operating status of distribution transformers. It can timely and accurately grasp the operating status and change trend of distribution transformers, laying the foundation for the optimization of distribution transformer maintenance strategy.

[0034] 3) By constructing reinforcement learning elements such as state, reward and maintenance action through data such as distribution transformer operating status characteristics, real-time operating status evaluation results and historical fault records, the DQN algorithm is used to train the reinforcement learning process of the interaction between the intelligent agent and the distribution transformer, and the uncertainty of the complex dynamic relationship between the state decay and fault occurrence of the distribution transformer is mined to obtain the distribution transformer maintenance strategy optimization model, such as Figure 2 shown.

[0035] Furthermore, the main steps of establishing the distribution transformer maintenance strategy optimization model are as follows:

[0036] 3.1) In the reinforcement learning task of optimizing the economic operation and maintenance strategy for distribution transformers, the specific settings of reinforcement learning elements such as state S, maintenance action A, and reward R are as follows:

[0037] 3.1.1) The state of the distribution transformer includes the operating state features and operating state evaluation results score of the distribution transformer. The state s at the i-th moment is i The expression is:

[0038] s i ={features,score},s i ∈S(1)

[0039] In formula (1), the operating state characteristics mainly include load rate, voltage deviation and three-phase load imbalance, average load rate, heavy load duration, overload duration, average three-phase load imbalance and three-phase load imbalance duration; S is an infinite state set, and the transition probability between states is unknown. The distribution transformer state is divided into the initial state s0, the general state s i and the terminal state s T There are three situations. The initial state refers to the first state in the interaction process between the intelligent agent and the distribution transformer. In this state, the distribution transformer operates well and the operating status evaluation result is within the normal range. The terminal state refers to the distribution transformer that has a fault and the interaction process between the intelligent agent and the distribution transformer ends. The general state refers to the state other than the initial state and the terminal state. In this state, the distribution transformer may operate normally or may have a fault.

[0040] 3.1.2) The agent has two maintenance action options. The first is to not perform maintenance on the distribution transformer. no The second is to repair the distribution transformer. repair , the action a performed by the agent at time i i The expression is:

[0041] a i ={a no or a repair},a i ∈A(2)

[0042] When the agent chooses to execute a repair When , it shows that the decision recommendation given by the economic operation and maintenance strategy optimization model is to repair the distribution transformer. repair The action can be further refined, that is, the targeted maintenance processing action is refined according to the current monitoring status of the distribution transformer, and the state of the distribution transformer is converted to the initial state after the maintenance.

[0043] 3.1.3) The reward r obtained by the agent at time i i The maintenance action currently being executed is a i and the next moment state s of the distribution transformer feedback i+1 Joint decision. If si+1 is the terminal state, and the reward r obtained by the agent i If s is -1, the interaction process between the agent and the distribution transformer ends. i+1 For the general state, reward r i The calculation formula is:

[0044]

[0045] In formula (3), the reward r i The value range is between 0 and 1; score is the state s at the next moment i+1 The distribution transformer operating status evaluation results in the ; loss is the maintenance action a performed by the agent in different scenarios i economic losses; reward r i The first half of the β1score is related to the safe operation of the distribution transformer, β1 is the safety target coefficient; the reward r i The second half, β2(100-loss), is related to the economic cost of distribution transformer maintenance. β2 is the economic target coefficient. Because safety is more important than economic efficiency, β1 and β2 are set to 0.8 and 0.2, respectively.

[0046] 3.2) During the learning process of the DQN training agent interacting with the distribution transformer, the DQN training agent obtains factors such as the distribution transformer state s, optimizes the error between the approximate value function and the true value function through the gradient descent algorithm, and thus calculates the weight parameter w of the approximate value function. Then, the deep neural network is used to fit the true action value function to obtain the approximate action value function And use the Q-learning algorithm to update the approximate action value function to calculate the optimal action value function To obtain the optimal strategy In order to improve the efficiency and stability of the algorithm, the mechanism of experience replay and fixed Q target is introduced into the DQN training agent.

[0047] 3.3) The agent's maintenance decision interval is 1 hour. To expedite the reinforcement learning training process for the economical operation and maintenance of distribution transformers, in addition to considering the distribution transformer failure as the termination state, the agent's cumulative maintenance actions over a certain time scale are also considered termination states. The agent receives a reward of -1 in the former termination state and a reward of 0 in the latter. During DQN training, the agent's average reward is calculated every 10 episodes. As the number of training rounds increases, the average reward gradually converges. After 1000 episodes, the average reward converges to approximately 141, and the training process is complete. In the early stages of training, the agent's average reward is small and its variance is large, indicating that the agent jumps in and out of local optima during the initial exploration. In the middle stages of training, the average reward increases and its variance decreases, indicating that the agent is focusing on its learning focus during exploration. In the later stages of training, the average reward increases and its variance decreases, indicating that the agent is gradually approaching the global optimum during exploration. The DQN training process is as follows Figure 3 shown.

[0048] 3.4) In the DQN-based economic operation and maintenance process of distribution transformers, the DQN algorithm is used to train the reinforcement learning process of the interaction between the intelligent agent and the distribution transformer, thereby exploring the uncertainty of the complex dynamic relationship between the state decay and failure of the distribution transformer and obtaining an optimization model for the economic operation and maintenance strategy of the distribution transformer.

[0049] 4) A distribution transformer is selected for testing to obtain a predictive maintenance decision sequence.

[0050] Furthermore, the main steps of obtaining the predictive maintenance decision sequence are as follows:

[0051] 4.1) Using the prediction model of the distribution transformer operating status, we can get the trend diagram of the distribution transformer operating status, such as Figure 4 shown.

[0052] 4.2) The predicted results of the operating status of the distribution transformer are constructed into status information and input into the economic operation and maintenance optimization model of the distribution transformer. The economic operation and maintenance strategy optimization model outputs the predictive maintenance decision sequence, as shown in Table 1.

[0053] Table 1 Predictive maintenance decision sequence for a distribution transformer

[0054]

[0055] 4.3) Apply the reinforcement learning-based economic operation and maintenance method of the distribution transformer to the distribution transformer to obtain the maintenance status of the distribution transformer in winter and summer. Figure 5It can be seen that under the condition of ensuring the safe and stable operation of the distribution transformer, the economic operation and maintenance method of the distribution transformer based on reinforcement learning performs fewer maintenance treatments on the distribution transformer than the existing predictive maintenance strategy, indicating that the economic operation and maintenance method of the distribution transformer based on reinforcement learning has lower maintenance costs and better economy.

[0056] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A distribution transformer economic operation and maintenance method based on reinforcement learning, characterized in that: The main steps include: 1) Extract the multi-dimensional features of distribution transformers from online operation data and basic information of distribution transformers; 2) Based on the multi-dimensional perception of the operating status of distribution transformers, a data-driven distribution transformer operating status assessment and situation prediction model is established; 3) Through the DQN algorithm, the reinforcement learning process of the interaction between the intelligent agent and the distribution transformer is trained to obtain the distribution transformer maintenance strategy optimization model. The main steps of model construction are as follows: 3.1) In the reinforcement learning task of optimizing the economic operation and maintenance strategy of distribution transformers, the specific settings of the reinforcement learning elements of state S, maintenance action A and reward R are as follows: 3.1.1) The state of the distribution transformer includes the operating state features and operating state evaluation results score of the distribution transformer. The state s at the i-th moment is i The expression is: s i ={features,score},s i ∈S (1) In formula (1), the operating state characteristics include load rate, voltage deviation and three-phase load imbalance, average load rate, heavy load duration, overload duration, average three-phase load imbalance and three-phase load imbalance duration; S is an infinite state set, and the transition probability between states is unknown; the distribution transformer state is divided into initial state s0, general state s i and the terminal state s T There are three situations; the initial state refers to the first state in the process of interaction between the intelligent agent and the distribution transformer. In this state, the distribution transformer operates well and the operating status evaluation result is within the normal range; the terminal state refers to the distribution transformer having a fault and the interaction process between the intelligent agent and the distribution transformer ends; the general state refers to the state other than the initial state and the terminal state. In this state, the distribution transformer may operate normally or may have a fault; 3.1.2) The agent has two maintenance action options. The first is to not perform maintenance on the distribution transformer. no The second is to repair the distribution transformer. repair , the action a performed by the agent at time i i The expression is: a i ={a no or a repair },a i ∈A (2) When the agent chooses to execute a repair When , it indicates that the decision suggestion given by the economic operation and maintenance strategy optimization model is to repair the distribution transformer; a repair Action refinement, that is, refine targeted maintenance processing actions according to the current monitoring status of the distribution transformer, and the distribution transformer status is changed to the initial state after maintenance; 3.1.3) The reward r obtained by the agent at the i-th moment i The maintenance action currently being executed is a i and the next moment state s of the distribution transformer feedback i+1 Joint decision; if s i+1 is the terminal state, and the reward r obtained by the agent i is -1, the interaction process between the agent and the distribution transformer ends; if s i+1 For the general state, reward r i The calculation formula is: In formula (3), the reward r i The value range is between 0 and 1; score is the state s at the next moment i+1 The distribution transformer operating status evaluation results in the ; loss is the maintenance action a performed by the intelligent agent in different scenarios i economic loss; reward i The first half, β1score, is related to the safe operation of the distribution transformer, where β1 is the safety target coefficient; the reward r i The second half, β2(100-loss), is related to the economic cost of distribution transformer maintenance, and β2 is the economic target coefficient; 3.2) In the learning process of the interaction between the DQN algorithm training agent and the distribution transformer, the DQN training agent obtains the state s element of the distribution transformer, optimizes the error between the approximate value function and the true value function through the gradient descent algorithm, and thus calculates the weight parameter w of the approximate value function; then the real action value function is fitted through the deep neural network to obtain the approximate action value function And use the Q-learning algorithm to update the approximate action value function To obtain the optimal strategy In order to improve the efficiency and stability of the algorithm, the DQN training agent introduces the mechanism of experience replay and fixed Q target; 3.3) Setting the agent maintenance decision interval. It is not common for distribution transformers to fail during operation. In order to complete the reinforcement learning training process of economic operation and maintenance of distribution transformers as quickly as possible, in addition to taking the failure of the distribution transformer as the termination state, the cumulative execution of maintenance actions by the agent on a time scale is also regarded as the termination state; The reward obtained by the agent in the former termination state is -1, and the reward obtained in the latter termination state is 0; during the DQN training process, the average reward obtained by the agent is calculated every once in a while. As the number of training rounds increases, the average reward obtained by the agent gradually converges and the training process is completed; 3.4) In the DQN-based distribution transformer economic operation and maintenance process, the DQN algorithm is used to train the reinforcement learning process of the interaction between the intelligent agent and the distribution transformer, and the uncertainty of the complex dynamic relationship between the state decay and the occurrence of failure of the distribution transformer is mined to obtain the distribution transformer economic operation and maintenance strategy optimization model; 4) The prediction results of the distribution transformer operating status are constructed into status information and input into the distribution transformer maintenance strategy optimization model to obtain a predictive maintenance decision sequence.

2. The economic operation and maintenance method of distribution transformer based on reinforcement learning according to claim 1 is characterized in that: Extracting the multidimensional features of distribution transformers requires statistical analysis of the historical abnormal operating conditions and fault records of distribution transformers, summarizing the common fault types of distribution transformers and the main causes of the faults, and extracting real-time, statistical and basic multidimensional features from the online operating data of distribution transformer voltage and current and basic information. The real-time features include load rate, voltage deviation and three-phase load imbalance, the statistical features include average load rate, heavy load duration, overload duration, average three-phase load imbalance and three-phase load imbalance duration, and the basic features include family defect coefficient and remaining operating life coefficient.

3. The economic operation and maintenance method of distribution transformer based on reinforcement learning according to claim 2 is characterized in that: To establish a data-driven distribution transformer operating status assessment and situation prediction model, it is necessary to use an improved comprehensive evaluation method to realize real-time assessment of the distribution transformer operating status based on multi-dimensional perception of the distribution transformer operating status, and combine the long short-term memory neural network and support vector machine model to realize the prediction of the distribution transformer operating status.

4. A distribution transformer economic operation and maintenance method based on reinforcement learning according to claim 1 or claim 3, characterized in that: The distribution transformer operating status prediction results are constructed into status information and input into the distribution transformer economic operation and maintenance optimization model. The economic operation and maintenance strategy optimization model outputs a predictive maintenance decision sequence.