Intelligent transformer control system and method based on bionic autonomous architecture

By using a biomimetic autonomous architecture-based intelligent transformer control system, combined with model predictive control and deep reinforcement learning, adaptive and cooperative control of intelligent transformers is achieved. This solves the problems of model dependence and insufficient cooperation mechanisms in existing technologies, and improves the operating efficiency and fault response capability of power systems.

CN121763764APending Publication Date: 2026-03-31SHANDONG BEST NEW ENERGY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing intelligent transformer control methods rely on precise mathematical models, which are difficult to cope with changes in complex power grid environments. Furthermore, deep reinforcement learning schemes lack interpretability and security, and lack inter-device collaboration mechanisms, resulting in low system operating efficiency and insufficient fault response capabilities.

Method used

An intelligent transformer control system based on a biomimetic autonomous architecture is adopted, including a reflection control module, a learning optimization module, and a collaborative decision-making module. It achieves adaptive and collaborative control by combining model predictive control and deep reinforcement learning, optimizes the control strategy by combining lifetime degradation index, and uses an event-triggered mechanism for distributed consensus algorithm collaborative decision-making.

Benefits of technology

It improves the operating efficiency and robustness of intelligent transformer clusters in complex environments, achieves efficient fault recovery capabilities and extends equipment life, and enhances the resilience and power supply reliability of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121763764A_ABST
    Figure CN121763764A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power systems, and particularly relates to an intelligent transformer control system and method based on a bionic autonomous architecture, and the scheme achieves the self-adaption of the control in a single intelligent transformer through the construction of a three-layer bionic autonomous architecture of a reflection control module, a learning optimization module and a collaborative decision module. And the operation efficiency and control robustness of the intelligent transformer cluster in a complex power distribution network environment are remarkably improved through collaboration of group intelligence among multiple devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system technology, specifically relating to an intelligent transformer control system and method based on a biomimetic autonomous architecture. Background Technology

[0002] Currently, the control methods for intelligent transformers mainly adopt the following two categories: One type is centralized control schemes that employ traditional model predictive control, proportional-integral control, or linear quadratic optimal control. These schemes rely on accurate mathematical models of the lines and loads, typically assuming that the communication network is absolutely reliable and that the centralized controller has global information processing capabilities. Such schemes can achieve effective control under a single classic operating condition, but their performance is extremely sensitive to model accuracy and communication latency, making it difficult to cope with highly uncertain scenarios such as changing network topologies and the access of a large number of plug-and-play devices. Another type is the end-to-end adaptive control scheme based on deep reinforcement learning. This type of scheme obtains control strategies by learning from historical data, without the need for a precise physical model, and has the advantage of adaptability. However, this type of scheme usually treats the agent as a black box, which has problems such as poor interpretability, unstable convergence process and lack of real-time physical safety guarantee for control actions. At the same time, existing schemes mostly focus on the optimization of individual devices and lack intelligent collaboration mechanisms between devices.

[0003] In summary, the existing technical solutions have the following drawbacks: Traditional model predictive control and other methods rely heavily on accurate mathematical models, but the actual power grid is complex and ever-changing. Model mismatch will lead to performance degradation or even instability, and the system's ability to adapt autonomously and operate robustly is insufficient.

[0004] While data-driven methods such as single deep reinforcement learning possess adaptive capabilities, their direct generation of control commands as the "brain" may lead to dangerous actions that are not verified by physical constraints during the trial-and-error learning process. Furthermore, the black-box nature of these algorithms raises serious questions about their credibility and interpretability in critical power equipment applications.

[0005] Most existing solutions focus on "local" optimization at the level of individual devices, lacking an effective mechanism for information interaction and distributed collaborative decision-making among multiple agents at the system level. This leads to low overall operational efficiency and a lack of rapid and efficient fault support and system reconfiguration capabilities when some devices fail. Summary of the Invention

[0006] This invention provides an intelligent transformer control system and method based on a biomimetic autonomous architecture, which effectively solves the problems existing in the prior art.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A smart transformer control system based on a biomimetic autonomous architecture includes: The reflection control module is configured to, based on the real-time electrical state quantities of the smart transformer, perform calculations in microsecond cycles using a preset model predictive control algorithm to obtain the optimal switching control signal for directly driving the power switch of the smart transformer; and periodically send normal data including the real-time electrical state quantities to the learning optimization module; and when a fault is detected, send event data containing fault information to the collaborative decision-making module. The learning optimization module is configured to respond to normal data from the reflection control module by performing calculations on a second-level cycle through a pre-built deep reinforcement learning agent to obtain optimized parameters for adjusting the weight parameters of the model prediction control algorithm, and then send them to the reflection control module. The collaborative decision-making module is configured to respond to event data from the reflection control module, calculate using a pre-built distributed consensus algorithm based on an event-triggered mechanism, obtain collaborative instructions for collaborative control among multiple smart transformers, and send them to the learning optimization module. The learning optimization module is further used to optimize the decision-making process of the deep reinforcement learning agent based on the collaborative instructions sent by the collaborative decision-making module; the reflection control module is further used to optimize and adjust the objective function weight parameters in the model prediction control algorithm based on the optimization parameters sent by the learning optimization module.

[0008] Furthermore, the microsecond-level cycle of the model prediction control algorithm and the second-level cycle of the deep reinforcement learning agent constitute a heterogeneous cycle nested control architecture; wherein, the optimized parameter output of the second-level cycle is used as a slowly varying parameter and acts on the objective function of the microsecond-level cycle.

[0009] Furthermore, the deep reinforcement learning agent employs a pre-defined composite reward function to guide its decision optimization process, wherein the composite reward function is specifically expressed as follows: in, To calculate the reward, For operational efficiency optimization items, Let be the predicted junction temperature of the i-th switching device at time t. For safe temperature threshold, As an indicator of lifespan degradation, For power quality, For a binary item, , , , , These are the adjustment coefficients for different indicator items.

[0010] Furthermore, the calculation of the lifespan degradation index is specifically as follows: Real-time acquisition of physical quantities related to device aging; Using a pre-trained long short-term memory network as a lifetime prediction model, the remaining lifetime of the device is obtained based on the historical feature sequence of physical quantities related to device aging. Based on the model-predicted remaining lifetime, the lifetime degradation index of the device at the current moment is calculated, specifically expressed as: in, For design life, The remaining lifetime predicted by the model. It serves as an indicator of lifespan degradation.

[0011] Furthermore, the optimization parameters are applied to the objective function of the model predictive control algorithm to dynamically adjust the weight ratios between different control objectives within it.

[0012] Furthermore, the objective function of the model predictive control algorithm is specifically expressed as follows: Constraints: System dynamic constraints: Input constraints: ; Soft constraints: ; Safe operation constraints: ; in, For prediction in the time domain; To control the time domain; This is the output reference trajectory for the j-th future step at the current time k; To output the tracking error weight matrix; The weighted L2 norm squared of vector V is denoted as . ; To control the increment of the input; To control the incremental weight matrix; A vector of slack variables; The slack variable penalty weight vector; and To control the lower and upper bound vectors of the input; and Let them be the lower and upper bound vectors of the state variables; The predicted switch current for the j-th future step; This is the maximum allowable current function related to junction temperature; This indicates the predicted junction temperature.

[0013] Furthermore, the distributed consensus algorithm based on the event triggering mechanism is calculated, specifically including the following processes: constructing a node network topology; triggering information interaction between neighboring nodes in response to fault event data; and calculating the local coordination instructions of each node based on the interaction information.

[0014] Furthermore, the event triggering mechanism employs a dynamic threshold, which is adjusted based on system uptime or device health status to achieve a balance between communication load and collaborative accuracy.

[0015] Furthermore, the dynamic threshold is specifically represented as follows: in, , , These are adjustable parameters used to adjust the decay rate and weights. Let t be the lifetime degradation index of node i at time k, and t be the total running time or the number of events since the last major event.

[0016] The intelligent transformer control method based on a biomimetic autonomous architecture, which is based on the aforementioned intelligent transformer control system based on a biomimetic autonomous architecture, includes: The reflection control module, based on the real-time electrical state of the smart transformer, calculates the optimal switching control signal for directly driving the power switch of the smart transformer using a preset model predictive control algorithm at microsecond intervals; and periodically sends normal data including the real-time electrical state to the learning and optimization module; and when a fault is detected, it sends event data containing fault information to the collaborative decision-making module. In response to normal data from the reflection control module, the learning optimization module calculates on a second-level cycle using a pre-built deep reinforcement learning agent to obtain optimized parameters for adjusting the weight parameters of the model prediction control algorithm, and sends them to the reflection control module. In response to event data from the reflection control module, the collaborative decision-making module calculates using a pre-built distributed consensus algorithm based on an event-triggered mechanism to obtain collaborative instructions for collaborative control among multiple smart transformers, and sends them to the learning optimization module. The learning optimization module is further used to optimize the decision-making process of the deep reinforcement learning agent based on the collaborative instructions sent by the collaborative decision-making module; the reflection control module is further used to optimize and adjust the objective function weight parameters in the model prediction control algorithm based on the optimization parameters sent by the learning optimization module.

[0017] Compared with the prior art, the advantages and positive effects of the present invention are as follows: (1) The present invention provides a smart transformer control system and method based on a biomimetic autonomous architecture. The solution realizes the adaptive control within a single smart transformer and the collaborative intelligence among multiple devices by constructing a three-layer biomimetic autonomous architecture consisting of a reflection control module, a learning optimization module and a collaborative decision-making module. This significantly improves the operating efficiency and control robustness of smart transformer clusters in complex power distribution network environments.

[0018] (2) The solution described in this invention mimics the “reflection-learning” mechanism of organisms, combining the bottom-level millisecond-level high-frequency platform protection, MPC model optimization control (reflection control module) and the top-level second-level DRL strategy self-tuning (learning optimization module); wherein, the reflection control module operates independently, based on the accurate physical model combined with the real-time state, directly and quickly suppressing and constraining the parameters that may endanger the equipment, providing an absolutely safe action execution boundary for the intelligent optimization of the learning optimization module, so that the agent can explore and optimize in a safe sandbox, and achieve the unity of high robustness and adaptive capability in complex environments.

[0019] (3) The solution described in this invention learns about environmental changes online through the DRL agent and dynamically adjusts the key parameter vectors in the MPC algorithm of the reflection control module, giving the entire control system self-tuning, self-learning, and self-optimization capabilities. It can autonomously adapt to various uncertainties such as load mutations, new energy power fluctuations, and model parameter perturbations, and no longer relies on accurate offline models. At the same time, by integrating the reward function of lifetime degradation index, the control strategy is guided to operate with low stress and low loss, thereby actively extending the service life of the power electronic module and realizing the synergistic optimization of performance and life.

[0020] (4) The solution described in this invention designs an event-triggered distributed consensus algorithm, which enables multiple devices to work together to optimize system voltage and balance reactive power distribution under normal conditions, significantly reducing network losses and voltage deviations. In case of emergency events such as faults, all devices can calculate their own "remaining support capacity" and autonomously and quickly share the power and voltage support requirements of the fault node in proportion, realizing a rapid and self-organized emergency response and fault recovery capability similar to a biological community, which greatly improves the resilience and power supply reliability of the local distribution network. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below: Figure 1 This is a schematic diagram of the intelligent transformer control system based on a biomimetic autonomous architecture as described in the embodiment; Figure 2 This is a flowchart of the intelligent transformer control method based on a biomimetic autonomous architecture described in the embodiment; Figure 3 This is a flowchart illustrating the parameter transfer and closed-loop update process described in the embodiment. Figure 4 This is a flowchart of the event coordination algorithm described in the embodiment; Figure 5 This is a flowchart illustrating the decision result distribution and execution process described in the embodiment. Detailed Implementation

[0022] To better understand the above-mentioned objectives, features and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0023] Numerous specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways than those described herein, and therefore the invention is not limited to the specific embodiments disclosed in the following specification.

[0024] Example 1 The intelligent transformer control system based on a biomimetic autonomous architecture described in Embodiment 1 will be described in detail below with reference to the accompanying drawings.

[0025] First, it should be noted that the core idea of ​​the scheme described in this embodiment is to divide the control of the intelligent transformer into three mutually cooperating modules through biological simulation: a reflection control module, a learning optimization module, and a collaborative decision-making module; wherein: The reflection control module simulates the spinal reflex arc of a biological organism, and is responsible for microsecond-level rapid control and hardware protection. By adopting the model predictive control (MPC) algorithm, it performs unified and coordinated closed-loop control on all ports (AC and DC) of the smart transformer. By learning and optimizing the module to simulate the learning function of the cerebral cortex of organisms, it is responsible for the second-level adaptive parameter optimization. By adopting a lightweight deep reinforcement learning (DRL) agent, it dynamically adjusts the weight parameters of the MPC algorithm in the reflection control module according to the operating environment, load conditions and health status of the equipment, so as to achieve comprehensive optimization of multiple objectives such as efficiency, thermal stress and equipment life. The collaborative decision-making module simulates the social actions of biological groups, is responsible for distributed decision-making at the second level, and uses a distributed consensus algorithm based on an event triggering mechanism to achieve voltage and reactive power optimization among multiple smart transformers, as well as to perform collaborative support in a release-and-mutual-aid manner when a fault occurs.

[0026] In one or more embodiments, such as Figure 1 As shown, the intelligent transformer control system based on a biomimetic autonomous architecture includes: The reflection control module is configured to, based on the real-time electrical state quantities of the smart transformer, perform calculations in microsecond cycles using a preset model predictive control algorithm to obtain the optimal switching control signal for directly driving the power switch of the smart transformer; and periodically send normal data including the real-time electrical state quantities to the learning optimization module; and when a fault is detected, send event data containing fault information to the collaborative decision-making module. The learning optimization module is configured to respond to normal data from the reflection control module by performing calculations on a second-level cycle through a pre-built deep reinforcement learning agent to obtain optimized parameters for adjusting the weight parameters of the model prediction control algorithm, and then send them to the reflection control module. The collaborative decision-making module is configured to respond to event data from the reflection control module, calculate using a pre-built distributed consensus algorithm based on an event-triggered mechanism, obtain collaborative instructions for collaborative control among multiple smart transformers, and send them to the learning optimization module. The learning optimization module is further used to optimize the decision-making process of the deep reinforcement learning agent based on the collaborative instructions sent by the collaborative decision-making module; the reflection control module is further used to optimize and adjust the objective function weight parameters in the model prediction control algorithm based on the optimization parameters sent by the learning optimization module.

[0027] In specific implementation, the microsecond-level cycle of the model prediction control algorithm and the second-level cycle of the deep reinforcement learning agent constitute a heterogeneous cycle nested control architecture; wherein, the optimized parameter output of the second-level cycle is used as a slowly varying parameter and acts on the objective function of the microsecond-level cycle.

[0028] In practice, the deep reinforcement learning agent uses a preset composite reward function to guide its decision optimization process.

[0029] In practice, the optimization parameters are applied to the objective function of the model predictive control algorithm to dynamically adjust the weight ratios between different control objectives within it.

[0030] In specific implementation, the distributed consensus algorithm based on the event triggering mechanism is calculated, which includes the following processing steps: constructing the node network topology; triggering information interaction between neighboring nodes in response to fault event data; and calculating the local coordination instructions of each node based on the interaction information.

[0031] In specific implementation, the reflection control module includes the following processing steps: (1) System modeling and discretization First, taking a smart transformer using a cascaded H-bridge or modular multilevel topology as an example, a mathematical model of the smart transformer is performed. The modeling objects include: the AC side grid connection port, the DC side load or energy storage port, and the DC bus capacitor and switching devices of the internal power module. The specific modeling process is as follows: Based on Kirchhoff's laws and the switching function averaging method, a continuous-time averaged state-space model of the system is established: in, State vector The first derivative with respect to time, i.e., the rate of change of state, Let be the continuous-time state vector of the system. To control the input vector, For the measurable disturbance vector, For the system output vector, Let be the continuous-time state matrix of the system. For continuous-time input matrices, The perturbation matrix is ​​a continuous-time perturbation matrix. This is the output matrix.

[0032] The state vector This includes, but is not limited to, AC inductor current, DC bus capacitor voltage, and estimated junction temperature of switching devices; The control input vector This mainly refers to the modulation signal (such as duty cycle) of each bridge arm. The measurable disturbance vector Mainly refers to grid voltage; The output vector This mainly refers to the physical quantities to be controlled, such as grid connection point voltage and DC bus voltage.

[0033] The continuous-time state matrix It is determined by parameters such as resistance, inductance, and capacitance of the circuit topology; The continuous-time input matrix It describes how control inputs affect changes in state; The continuous-time perturbation matrix This describes how external disturbances affect changes in the state; The output matrix , used to select or combine output vectors from state vectors.

[0034] To facilitate mathematical calculations, the scheme described in this embodiment uses the forward Euler method or a zero-order hold to discretize the above continuous model, obtaining a discrete-time model suitable for online computation: in, This represents the k-th discrete control cycle. In this embodiment, the cycle length is set to 50 milliseconds. This is the discretized state vector at the start of the k-th cycle. The control input vector is constant and applied during the k-th period. Let be the perturbation vector measured or predicted during the k-th period. Let be the perturbation vector measured or predicted during the k-th period. The discrete-time state matrix, The input matrix is ​​in discrete time. The perturbation matrix is ​​in discrete time. The output matrix in discrete time is usually the same as the output matrix mentioned above. same.

[0035] (2) Data processing and online optimization solution Taking the k-th control cycle as an example, the reflection control module performs the following processing procedure to obtain the optimal control quantity. The specific steps are as follows: 1) Data Acquisition and State Estimation: The system collects electrical quantities of the intelligent transformer in real time through preset sensors (e.g., AC inductor current, DC bus capacitor voltage, and estimated junction temperature of switching devices), and uses a Kalman filter to analyze the system status. Perform optimal estimation to suppress measurement noise and obtain the estimated state. ; 2) State and output prediction: State estimated based on the current moment and known / predicted perturbation sequences { Using discrete-time models to predict the future The state trajectory at each time step and output trajectory , where (k+j|k) represents the prediction of time k+j at time k.

[0036] 3) Construct and solve the finite-time optimization problem: The objective of the scheme described in this embodiment is to find a control input sequence that optimizes the system's performance over a future period of time. The objective function is specifically expressed as follows: Constraints: A. System dynamic constraints: B. Input Constraints: ; C. Soft constraints: ; D. Safety operation constraints: ; in, For the prediction time domain, i.e., the number of steps to predict forward (e.g.: =10); To control the time domain, i.e., the length of the future control input sequence to be optimized (e.g.: ,and ),exist After the step, the control input is usually assumed to remain unchanged; The output reference trajectory for the j-th future step at the current time k is usually set to a constant value; The output tracking error weight matrix is ​​typically a positive definite diagonal matrix; The weighted L2 norm squared of vector V is denoted as . ; To control the increment of the input, it is defined as: ,in, The control quantity actually applied in the previous cycle; To control the incremental weight matrix, it is usually a positive definite diagonal matrix; This is a vector of slack variables. Introducing this vector can avoid misunderstandings in the optimization problem caused by the infeasibility of hard constraints; that is, it allows the state to... Violation of composite constraints within the scope; The slack variable penalty weight vector consists of large positive numbers that are used to penalize the use of the corresponding slack variable, ensuring that constraints are violated only when necessary. and The lower and upper bound vectors of the input are determined by the physical limits of the hardware (such as a duty cycle between 0 and 1); and These are the lower and upper bound vectors for the state variables, specifically set according to the safety operation requirements; The predicted switch current for the j-th future step; This is the maximum allowable current function related to junction temperature. This relationship is provided in the device datasheet and indicates how current increases with the predicted junction temperature. As the current increases, the maximum allowable current value must be reduced accordingly to prevent thermal breakdown.

[0037] 4) Problem optimization solution The optimization problem described above is usually transformed into a quadratic programming (QP) problem: in, For the decision variable vector, Let f be the Hessian matrix, f be the gradient vector, and G and h define linear inequality constraints.

[0038] To meet microsecond-level real-time requirements, a dedicated solver with high efficiency QP solvers (such as those based on interior-point or active-set methods) is employed and deployed on an FPGA or high-performance DSP to ensure the solution is completed within a single control cycle. The solution result is the optimal control input sequence: 5) Extract and apply real-time control quantities A rolling time-domain control strategy is adopted, taking only the first element in the optimal sequence as the actual control quantity applied in the current period: This control quantity After modulation, a PWM signal is generated to drive the power switching devices, thereby realizing closed-loop control of the smart transformer.

[0039] 6) Periodic rolling updates At the next sampling time k+1, repeat steps 1) to 5) to achieve continuous control optimization.

[0040] In one or more embodiments, to cope with nanosecond-level faults (such as direct short circuits), the reflection control module is also equipped with an ultra-high-speed hardware protection channel independent of the MPC software algorithm, and its processing logic is as follows: Real-time monitoring of switching transistor current using a high-speed comparator ,Voltage And temperature sensor signals; When any signal exceeds a preset safety threshold, the channel immediately generates and executes protection actions, including but not limited to: a) Block all drive signals directly through hardware logic; b) Send a termination signal to the MPC controller to pause and enter a safe state; c) Send a fault event message to the coordination layer to trigger a group system response.

[0041] In specific implementation, the learning optimization module serves as the intelligent hub for achieving long-term performance adaptation and multi-objective optimization of the entire system. Its main concept is as follows: First, the control parameter optimization problem is transformed into a DRL (Deep Reinforcement Learning) problem; then, multi-data from the reflection control module, the environment (e.g., operating condition information, physical parameters, and external commands), and the lifetime model (in this embodiment, an LSTM network is used for lifetime prediction) is collected and processed; next, through the learning and decision-making process of the DRL agent, the optimal parameter adjustment strategy for the MPC algorithm is calculated; finally, this strategy is output and applied to the reflection control module. Specifically, the following processing steps are included: (1) Problem modeling To achieve adaptive optimization, the problem of optimizing the optimal parameters of the reflection control module's MPC algorithm is first modeled as a Markov decision process (MDP) that can be solved by deep reinforcement learning (DRL): 1) State Space definition: At any moment t The state observed by the DRL agent A high-dimensional information vector is constructed as follows: in, These are estimated key state quantities from the reflection control module, such as current, voltage, and temperature. For environmental and operating condition monitoring, such as current load power, ambient temperature, and grid frequency deviation, Historical information summaries, such as the changing trends of control inputs over a period of time and historical losses, can be represented by the hidden state of a recursive network (such as LSTM). This state design ensures that the agent can make decisions based on the full-dimensional dynamics of the system.

[0042] 2) Definition of action space A: Actions output by the intelligent agent The parameter set directly corresponds to the MPC algorithm of the reflection control module, which requires dynamic adjustment. : in, For the weight vector, These are the weighting coefficients for optimization objectives such as efficiency, thermal stress, and lifespan.

[0043] 3) Reward function design: reward function This is a quantitative manifestation of multi-objective trade-offs, specifically expressed as follows: in, For efficiency optimization items, Let be the predicted junction temperature of the i-th switching device at time t. For safe temperature threshold, As an indicator of lifespan degradation, For power quality, For a binary item, , , , , These are the adjustment coefficients for different indicator items.

[0044] It should be noted here that: Efficiency optimization items This refers to the total system loss, including switching loss and conduction loss, which can be set according to actual needs. Power quality items The total harmonic distortion rate (THD) of the AC side port output voltage is the instantaneous value or its average value over one learning cycle.

[0045] binary item It is an exponential function, with a value of 1 when all hard safety constraints are met, and 0 otherwise. The hard safety constraints include, but are not limited to, DC bus voltage not exceeding limits, grid-side current not exceeding limits, etc.

[0046] (2) Multi-source information fusion The learning optimization module needs to process and effectively integrate data from different sources and time scales, specifically including: 1) Data stream of the reflection control module It periodically receives initial normal data from the reflection control module, mainly including the state vector. The sampled values ​​or statistical characteristics (such as mean or variance); 2) Data processing for life prediction models A lifetime prediction model based on a Long Short-Term Memory (LSTM) network is constructed, and its data processing flow is as follows: Input feature vector Real-time acquisition of physical quantities related to device aging, such as the voltage ripple amplitude of the DC bus capacitor, junction temperature of switching devices, effective current value, and cumulative running time; Model prediction: Utilizing a pre-trained LSTM model, based on historical feature sequences... Predicting future remaining useful life .

[0047] 3) Calculation of lifespan degradation index Predicting remaining lifetime based on LSTM model Calculate the lifespan degradation index at the current moment. , as input to the reward function: in, For design life, The larger the value, the worse the device's health status, resulting in stronger negative incentives in the reward function, prompting the DRL agent to adopt a more "gentle" control strategy—extending its lifespan.

[0048] (3) Training and decision-making of DRL agents 1) Construction of DRL agents The Actor-Critic framework is adopted, and the specific implementation uses Proximal Policy Optimization (PPO). Its main architecture includes: Policy Network (Actor) According to the status Output Action (i.e., parameter set) ); Value Network (Critic) : Assess the value of the current state and know when to update the policy.

[0049] 2) Offline training Large-scale offline training is performed in a simulation environment to learn a general optimization strategy: a) The agent interacts with the environment to generate experience trajectories. ; b) Update the policy network parameters using the collected empirical data and the PPO algorithm. and value network parameters ; c) Repeat the interaction and update until the policy converges and stabilizes on the simulation test set.

[0050] 3) Online decision-making and fine-tuning The trained policy network is deployed on edge computing units. In each learning cycle, the agent adjusts its behavior based on the current state. The optimal action is calculated directly through forward propagation. ; At the same time, it is necessary to set an appropriate update cycle and use real data collected during online operation to perform small, safe incremental updates to the policy network in order to adapt to the uniqueness of the actual operating environment.

[0051] (4) Parameter transfer and closed-loop update like Figure 3As shown, the specific processing steps include the following: 1) Generate and output optimization parameters; The action output by the agent at each decision time t in the learning optimization module. This is the parameter set that the MPC algorithm in the reflection control module should use in the next time period. .

[0052] 2) Pass parameters to the reflection control module; The learning optimization module transmits the parameter set through a predefined internal communication interface. This data is sent as the second data to the reflection control module.

[0053] 3) New parameters are applied to the reflection control module; The reflection control module received the parameter set. Then, immediately or at the beginning of the next control cycle, it is applied to the weight matrix and objective function in the MPC algorithm optimization process to complete the closed-loop update.

[0054] 4) Receive instructions from the collaborative decision-making module for optimization; When receiving third data (such as new voltage / reactive power reference values) sent by the collaborative decision-making module The learning optimization module uses this data as a state. This is part of the output MPC algorithm parameter set, thus enabling the next round of decision-making. It enables the reflection control module to track the new target of the collaborative decision-making module.

[0055] In practical implementation, the collaborative decision-making module mainly performs the following processing steps: It should be noted that the collaborative decision-making module is the core of realizing the collective intelligence and biomimetic mutual assistance of multi-agent systems. In this embodiment, it mainly follows the following processing logic: First, clarify the collaborative goal and system topology; then design an efficient event-triggered communication mechanism; next, execute distributed system algorithms based on interactive information to reach consensus or optimal decision; finally, distribute the decision results to the learning layer of each agent to know its local optimization.

[0056] (1) Definition of collaborative goals and system modeling 1) System topology and node definition Multiple smart transformers deployed in the same power distribution network area are regarded as a set of nodes. These nodes form a communication network via power line carrier communication, wireless mesh networks, or wired Ethernet, and its topology is represented by an undirected graph. Let E be the set of edges. For node i, its neighbor set is defined as follows: .

[0057] 2) Collaborative Target Classification Collaborative objectives are divided into two categories based on their operational status: a) Routine collaborative goals: When the system is running smoothly, the focus should be on global performance optimization, for example: Voltage uniformity: ensuring that the voltages at all nodes converge to a common optimal value; Reactive power balancing: Under the premise of meeting voltage requirements, the reactive power undertaken by each node is proportional to its capacity; Load balancing: Based on the capacity and health status of each node, the total load of the region is distributed proportionally.

[0058] b) Event Coordination Objectives In the event of a system malfunction or failure, the focus should be on rapid recovery and support, for example: Bionic cooperative fault support: Faulty nodes receive power or voltage support from healthy neighboring nodes. De-rated operation coordination: Multiple devices work together to reduce total power and avoid local overload or overheating; Fault isolation and reconstruction: collaboratively isolate faulty regions and transfer the load to healthy nodes.

[0059] (2) Information exchange mechanism: event-triggered communication To reduce communication overhead and improve response speed, an event-triggered mechanism is adopted to replace traditional periodic communication, and the following processing procedure is executed. 1) Triggering condition design Node i calculates its local state error vector at discrete time k: in, This is the key operating state vector of node i (such as voltage, current, power, etc.). Let $i$ be the moment when node $i$ was last successfully triggered and transmitted data. When the norm of this error vector exceeds a dynamic threshold, $i$ is considered to be in the state of $i$. At that time, that is Node i triggers a communication session.

[0060] 2) Dynamic threshold calculation In practical implementation, the dynamic threshold Adaptive adjustment based on system status: in, , , These are adjustable parameters used to adjust the decay rate and weights. This is a lifespan degradation indicator, where t is the total operating time or the number of events since the last major event.

[0061] It should be noted here that the first item To make the system more sensitive to errors during the initial commissioning phase, frequent communication is used to accelerate the initialization learning, parameter tuning, and convergence process; the second item When the health status of device i deteriorates (i.e., the lifespan degradation index increases), the threshold will decrease accordingly. This means that during periods when the device is vulnerable or requires higher reliability, the system will increase the communication frequency and improve the collaborative control precision and status awareness to provide more refined protection.

[0062] 3) Event Information Packet Construction When the triggering condition is met, node i constructs and broadcasts (or multicasts) a standardized event information packet. The event information packet includes the following information: node identifier, event type, event severity level, data timestamp, key state vector snapshot at the time of triggering, and the remaining support capacity currently available to node i.

[0063] (3) Collaborative algorithms and decision-making: Distributed computing Based on the received time information packets from neighboring nodes, each node independently runs a lightweight distributed algorithm to achieve a globally or locally optimal decision.

[0064] 1) Normal Cooperative Algorithm Taking voltage consistency as an example, node i asynchronously updates its local voltage reference value according to the following iterative formula: in, Let k be the reference voltage value at node i at time k. The measured or estimated voltage value of node i at time k; The voltage value of neighbor j obtained from the most recent event packet at time k; For communication weights; This is the iteration step size (i.e., the learning rate).

[0065] 2) Event Cooperative Algorithm When node i fails (e.g., power module failure triggers the FAULT_DETECTED event), such as Figure 4 As shown, perform the following fast distributed decision: Step 1: Fault Information Broadcast The faulty node i immediately broadcasts a high-priority event information packet. ,in: Event type = SUPPORT_REQUEST; Severity level of the incident = 9; The key state vector snapshot at the trigger moment includes its current voltage drop and reactive power gap.

[0066] Step 2: Neighbor Node Reactive Power Support Decision Healthy neighbor node j received Then, the reactive power support it can provide is calculated independently and in parallel.

[0067] Step 3: Voltage Support Decision Neighbor node j simultaneously calculates the temporary offset of the local voltage reference value required to support the voltage of the faulty node and compensate for the line voltage drop.

[0068] Understandably, similar decisions will be made in the event of other types of failures, which will not be elaborated here.

[0069] (4) Distribution and execution of decision results: closed-loop collaboration The ultimate goal of the collaborative decision-making module is to guide the local behavior of each agent, forming a closed loop from group decision-making to physical execution, such as... Figure 5 As shown, the specific processing steps include the following: Step 1: Generate cooperative instruction vectors Node j (whether after a normal update or an event response) packages the local decision result calculated by the distributed algorithm into a standard cooperative instruction vector. , wherein the cooperative instruction vector This includes updated local voltage reference values ​​and local reactive power reference values.

[0070] Step 2: Pass instructions to the learning optimization module The collaborative decision-making module of node j uses an internally defined interface to... As third data, it is sent to the intelligent agent of its corresponding learning and optimization module; Step 3: Learn and optimize module integration instructions. The DRL agent in the learning optimization module will receive As its state As part of this, at the next decision time t+1, the AI ​​will output a new set of MPC algorithm parameters. This set of parameters will guide the MPC algorithm of the launch control module, enabling its output control actions to effectively track the new targets issued by the collaborative decision-making module.

[0071] Step 4: The reflection control module executes the optimized control. The launch control module adopts a new MPC algorithm parameter set. It solves the MPC optimization problem in real time within a microsecond-level control cycle, generates PWM drive signals, and ultimately achieves coordinated goals such as voltage regulation, reactive power compensation or power support in physical terms, completing a complete closed loop of "sensing-communication-decision-execution".

[0072] Example 2 like Figure 2 As shown, the intelligent transformer control method based on a biomimetic autonomous architecture, which is based on the aforementioned intelligent transformer control system based on a biomimetic autonomous architecture, includes: The reflection control module, based on the real-time electrical state of the smart transformer, calculates the optimal switching control signal for directly driving the power switch of the smart transformer using a preset model predictive control algorithm at microsecond intervals; and periodically sends normal data including the real-time electrical state to the learning and optimization module; and when a fault is detected, it sends event data containing fault information to the collaborative decision-making module. In response to normal data from the reflection control module, the learning optimization module calculates on a second-level cycle using a pre-built deep reinforcement learning agent to obtain optimized parameters for adjusting the weight parameters of the model prediction control algorithm, and sends them to the reflection control module. In response to event data from the reflection control module, the collaborative decision-making module calculates using a pre-built distributed consensus algorithm based on an event-triggered mechanism to obtain collaborative instructions for collaborative control among multiple smart transformers, and sends them to the learning optimization module. The learning optimization module is further used to optimize the decision-making process of the deep reinforcement learning agent based on the collaborative instructions sent by the collaborative decision-making module; the reflection control module is further used to optimize and adjust the objective function weight parameters in the model prediction control algorithm based on the optimization parameters sent by the learning optimization module.

[0073] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications or equivalent changes made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. An intelligent transformer control system based on a biomimetic autonomous architecture, characterized in that, include: The reflection control module is configured to calculate the optimal switching control signal for directly driving the power switch of the smart transformer based on the real-time electrical state of the smart transformer through a preset model predictive control algorithm in microsecond-level cycles. In addition, it periodically sends normal data including the real-time electrical state quantities to the learning and optimization module; and when a fault is detected, it sends event data containing fault information to the collaborative decision-making module. The learning optimization module is configured to respond to normal data from the reflection control module by performing calculations on a second-level cycle through a pre-built deep reinforcement learning agent to obtain optimized parameters for adjusting the weight parameters of the model prediction control algorithm, and then send them to the reflection control module. The collaborative decision-making module is configured to respond to event data from the reflection control module, calculate using a pre-built distributed consensus algorithm based on an event-triggered mechanism, obtain collaborative instructions for collaborative control among multiple smart transformers, and send them to the learning optimization module. The learning optimization module is further used to optimize the decision-making process of the deep reinforcement learning agent based on the collaborative instructions sent by the collaborative decision-making module; the reflection control module is further used to optimize and adjust the objective function weight parameters in the model prediction control algorithm based on the optimization parameters sent by the learning optimization module.

2. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 1, characterized in that, The microsecond-level cycle of the model prediction control algorithm and the second-level cycle of the deep reinforcement learning agent constitute a heterogeneous cycle nested control architecture; wherein, the optimized parameter output of the second-level cycle is used as a slowly varying parameter and acts on the objective function of the microsecond-level cycle.

3. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 1, characterized in that, The deep reinforcement learning agent uses a pre-defined composite reward function to guide its decision optimization process, wherein the composite reward function is specifically represented as follows: in, To calculate the reward, For operational efficiency optimization items, Let be the predicted junction temperature of the i-th switching device at time t. For safe temperature threshold, As an indicator of lifespan degradation, For power quality, For a binary item, , , , , These are the adjustment coefficients for different indicator items.

4. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 3, characterized in that, The calculation of the lifespan degradation index is as follows: Real-time acquisition of physical quantities related to device aging; Using a pre-trained long short-term memory network as a lifetime prediction model, the remaining lifetime of the device is obtained based on the historical feature sequence of physical quantities related to device aging. Based on the model-predicted remaining lifetime, the lifetime degradation index of the device at the current moment is calculated, specifically expressed as: in, For design life, The remaining lifetime predicted by the model. It serves as an indicator of lifespan degradation.

5. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 1, characterized in that, The optimization parameters are applied to the objective function of the model predictive control algorithm to dynamically adjust the weight ratios between different control objectives within it.

6. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 5, characterized in that, The objective function of the model predictive control algorithm is specifically expressed as follows: Constraints: System dynamic constraints: Input constraints: ; Soft constraints: ; Safe operation constraints: ; in, For prediction in the time domain; To control the time domain; This is the output reference trajectory for the j-th future step at the current time k; To output the tracking error weight matrix; The weighted L2 norm squared of vector V is denoted as . ; To control the increment of the input; To control the incremental weight matrix; A vector of slack variables; The slack variable penalty weight vector; and To control the lower and upper bound vectors of the input; and Let them be the lower and upper bound vectors of the state variables; The predicted switch current for the j-th future step; This is the maximum allowable current function related to junction temperature; This indicates the predicted junction temperature.

7. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 1, characterized in that, The distributed consensus algorithm based on the event triggering mechanism is calculated, specifically including the following processes: constructing a node network topology; triggering information interaction between neighboring nodes in response to fault event data; and calculating local coordination instructions for each node based on the interaction information.

8. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 1, characterized in that, The event triggering mechanism uses a dynamic threshold, which is adjusted according to the system running time or the health status of the equipment to achieve a balance between communication load and collaborative accuracy.

9. The intelligent transformer control system based on a biomimetic autonomous architecture as described in claim 8, characterized in that, The dynamic threshold is specifically represented as follows: in, , , These are adjustable parameters used to adjust the decay rate and weights. Let t be the lifetime degradation index of node i at time k, and t be the total running time or the number of events since the last major event.

10. A method for controlling an intelligent transformer based on a biomimetic autonomous architecture, wherein the method is based on the intelligent transformer control system based on a biomimetic autonomous architecture as described in any one of claims 1-9, characterized in that, include: The reflection control module, based on the real-time electrical state of the smart transformer, calculates the optimal switching control signal for directly driving the power switch of the smart transformer in microsecond-level cycles using a preset model predictive control algorithm. In addition, it periodically sends normal data including the real-time electrical state quantities to the learning and optimization module; and when a fault is detected, it sends event data containing fault information to the collaborative decision-making module. In response to normal data from the reflection control module, the learning optimization module calculates on a second-level cycle using a pre-built deep reinforcement learning agent to obtain optimized parameters for adjusting the weight parameters of the model prediction control algorithm, and sends them to the reflection control module. In response to event data from the reflection control module, the collaborative decision-making module calculates using a pre-built distributed consensus algorithm based on an event-triggered mechanism to obtain collaborative instructions for collaborative control among multiple smart transformers, and sends them to the learning optimization module. The learning optimization module is further used to optimize the decision-making process of the deep reinforcement learning agent based on the collaborative instructions sent by the collaborative decision-making module; the reflection control module is further used to optimize and adjust the objective function weight parameters in the model prediction control algorithm based on the optimization parameters sent by the learning optimization module.

Citation Information

Patent Citations

  • Transformer on-load voltage regulation control method

    CN120879626A

  • Smart energy system based on power electronic transformer and monitoring method

    CN120879732A

  • Transformer optimization design method and system based on bionic swarm intelligence algorithm

    CN121031375A

  • Intelligent scheduling and control method and device for integrated energy system

    CN121032158A