A transformer voltage regulation collaborative control system based on Internet of Things (IoT) data acquisition
By implementing a collaborative control system that quantifies grid regulation resilience, identifies data attacks, and switches control, the vulnerability of AI controllers to data attacks in the Internet of Things (IoT) environment has been addressed. This system enables highly sensitive identification of data attacks and accurate assessment of their physical consequences, and dynamically switches control strategies, thereby enhancing the resilience and operational safety of the power grid.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 石家庄广运变压器有限公司
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies struggle to effectively identify covert data attacks in the Internet of Things (IoT) environment and quantify the physical consequences of such attacks in real time, making AI controllers vulnerable to data attacks and causing devices to fail more quickly.
By employing a power grid regulation resilience quantification unit, a data attack identification unit, a control switching decision unit, and a heterogeneous control strategy execution unit, and through spectrum correlation analysis and cumulative resilience decay index evaluation, the system achieves highly sensitive identification of data attacks and accurate assessment of their physical consequences, and dynamically switches control strategies to ensure a balance between system safety and efficiency.
It improves the accuracy of threat perception against data attacks such as high-frequency ghost loads, avoids equipment failure due to excessive fatigue, achieves a dynamic balance between the resilience and safety of power grid operation, and ensures that the system automatically resumes efficient operation after the threat is eliminated.
Smart Images

Figure CN121417246B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of transformer voltage regulation collaborative control technology, specifically a transformer voltage regulation collaborative control system based on Internet of Things (IoT) data acquisition. Background Technology
[0002] In the current Internet of Things (IoT) environment, power grid control systems are widely adopting AI controllers such as Deep Reinforcement Learning (DRL) to pursue the highest short-term operating efficiency, such as renewable energy absorption rate and voltage qualification rate. These AI controllers rely heavily on real-time measurement data from external IoT sensors to perform high-frequency, fine-grained voltage regulation control.
[0003] However, this architecture exposes the system to data attack risks, particularly high-frequency ghost load attacks. These attacks contaminate sensor data, inducing DRL controllers, which prioritize short-term efficiency, to misjudge the grid status and execute unnecessary high-frequency voltage regulation. This abnormal high-frequency regulation accelerates the physical fatigue wear of critical regulating equipment such as on-load tap changers of transformers, leading to a shortened equipment lifespan. Existing technologies struggle to simultaneously and effectively identify concealed data attacks and quantify the physical consequences of such attacks in real time, making AI controllers vulnerable to data attacks and causing accelerated equipment damage. Therefore, ensuring the efficient operation of AI controllers while simultaneously enabling them to perceive data attacks, assess physical consequences, and intelligently switch between efficient and secure strategies to enhance the resilience of the power grid in the IoT environment is a pressing technical challenge. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a transformer voltage regulation collaborative control system based on Internet of Things (IoT) data acquisition. Specifically, the technical solution of this invention includes:
[0005] The system includes a power grid regulation resilience quantification unit, a data attack identification unit, a control handover decision-making unit, and a heterogeneous control strategy execution unit.
[0006] The heterogeneous control policy execution unit is used to define deep reinforcement learning policies and classical robust control policies, and to generate expected control command signals;
[0007] The power grid regulation resilience quantification unit is used to collect real-time control command data of the transformer voltage regulation equipment to be controlled, and calculate the cumulative resilience attenuation index based on the real-time control command data.
[0008] The data attack identification unit is used to simultaneously acquire raw voltage measurement signals from external IoT sensors and expected control command signals generated by the heterogeneous control strategy execution unit; and to perform spectral correlation analysis on the raw voltage measurement signals and expected control command signals to obtain the attack correlation index.
[0009] The control handover decision unit is used to collect the cumulative resilience decay index and attack-related index to calculate the handover trigger value; and to generate the final control output based on the comparison result between the handover trigger value and the preset handover threshold.
[0010] The heterogeneous control policy execution unit is also used to switch between deep reinforcement learning policies and classical robust control policies in response to the final control output.
[0011] Preferably, the process of calculating the cumulative resilience attenuation index by the power grid regulation resilience quantification unit includes:
[0012] Based on real-time control command data, the adjustment frequency and adjustment amplitude of the transformer voltage regulating equipment to be controlled are obtained;
[0013] The toughness attenuation increment is calculated based on the adjustment frequency, adjustment amplitude, preset frequency reference, and preset weight parameters;
[0014] Among them, the toughness attenuation increment is used to quantify the physical fatigue loss caused by adjustment actions and adjustment amplitudes that exceed the frequency reference.
[0015] The cumulative toughness decay index is obtained by iterating the incremental toughness decay over time.
[0016] Preferably, the process of the data attack identification unit performing spectral correlation analysis includes:
[0017] The original voltage measurement signal and the expected control command signal are processed by time-frequency transformation respectively to extract the power spectral density of the signal in the preset ghost attack characteristic frequency band;
[0018] The power spectral density includes the voltage signal power spectrum and the control command power spectrum.
[0019] Preferably, the process by which the data attack identification unit obtains the attack-related index includes:
[0020] Calculate the total power of the voltage signal power spectrum within the ghost attack characteristic frequency band, and the total power of the control command power spectrum within the ghost attack characteristic frequency band, respectively.
[0021] Multiplying the total power of the voltage signal by the total power of the control command yields the attack correlation index;
[0022] The attack correlation index is used to quantify the abnormal response of the expected control command signal to high-frequency noise in the original voltage measurement signal.
[0023] Preferably, the process by which the control handover decision unit calculates the handover trigger value includes:
[0024] Obtain the cumulative resilience decay index and attack-related index;
[0025] Calculate the rate of change of the cumulative toughness decay index to obtain the toughness decay rate;
[0026] Divide the attack-related index by the preset attack threshold to obtain the normalized attack index.
[0027] Divide the cumulative resilience decay index by a preset decay threshold and apply a first weight to obtain the weighted cumulative loss;
[0028] Divide the toughness decay rate by a preset rate threshold and apply a second weight to obtain the weighted instantaneous impact.
[0029] Adding the weighted cumulative loss to the weighted instantaneous impact yields the normalized physical consequences;
[0030] Multiply the normalized attack exponent by the normalized physical consequence to obtain the switching trigger value.
[0031] Preferably, the process by which the control switching decision unit generates the final control output includes:
[0032] The switching trigger value is compared and analyzed with the preset switching threshold;
[0033] When the switching trigger value is greater than the preset switching threshold, it is determined to be a high-risk attack state, and the final control output switches to the classic robust control strategy.
[0034] When the switching trigger value is not greater than the preset switching threshold, it is determined to be a normal or low-risk state, and the final control output executes the deep reinforcement learning strategy.
[0035] Preferred deep reinforcement learning strategies are defined as follows:
[0036] Construct a reward function that aims to maximize short-term operating efficiency;
[0037] Short-term operating efficiency includes renewable energy absorption rate and voltage qualification rate;
[0038] The actions generated by deep reinforcement learning strategies aim to maximize the reward function.
[0039] The preferred, classic robust control strategy is defined as:
[0040] Construct a cost function that minimizes security costs;
[0041] The cost function includes a penalty term for voltage drop below a preset safe voltage lower limit, and an action penalty term applied to the adjustment range performed by the transformer voltage regulating equipment under control.
[0042] Classic robust control strategies generate actions that aim to minimize the cost function;
[0043] The effect of minimizing the cost function is to prioritize ensuring that the voltage is not lower than the lower limit of the safe voltage and minimize equipment operation, while accepting a decrease in the renewable energy consumption rate.
[0044] Preferably, it further includes recovery logic, the recovery logic including:
[0045] When the final control output is in the classic robust control strategy, continuously monitor the attack-related index and the cumulative resilience decay index.
[0046] Determine whether the attack-related index is below the attack threshold and continuously preset the recovery time;
[0047] Determine whether the cumulative resilience degradation index has fallen below the preset safe recovery threshold;
[0048] When both decisions are true, the trigger value will be reset so that the final control output is automatically returned to the deep reinforcement learning policy.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] 1. This system extracts power characteristics within a specific frequency band by performing spectral correlation analysis on sensor voltage signals and expected control command signals, enabling it to highly sensitively identify data attacks such as high-frequency ghost loads and improving the accuracy of threat perception;
[0051] 2. This system collects real-time control commands from transformer voltage regulating equipment, quantifies the physical fatigue losses caused by its regulation frequency and amplitude, and achieves accurate assessment of the physical consequences of attacks, thus preventing equipment failure due to excessive fatigue.
[0052] 3. This system has constructed a decision-making mechanism that integrates attack confirmation and physical consequence assessment. Switching is only triggered when the attack is confirmed and the physical consequences are severe, avoiding unnecessary rigid switching and achieving a dynamic balance between security and efficiency.
[0053] 4. This system provides a heterogeneous control strategy that pursues short-term efficiency and ensures long-term security. Through automatic recovery logic, it ensures that the system automatically returns to the high-efficiency mode after the threat is eliminated and the equipment is restored, thus constructing a complete anti-fragile closed loop and significantly improving the resilience of the power grid operation. Attached Figure Description
[0054] The present invention will be further explained below with reference to the accompanying drawings and embodiments:
[0055] Figure 1 This is a structural diagram of the system of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0057] Example 1:
[0058] Please see Figure 1 A transformer voltage regulation collaborative control system based on Internet of Things (IoT) data acquisition includes:
[0059] The system includes a power grid regulation resilience quantification unit, a data attack identification unit, a control handover decision-making unit, and a heterogeneous control strategy execution unit.
[0060] The heterogeneous control policy execution unit is used to define deep reinforcement learning policies and classical robust control policies, and to generate expected control command signals;
[0061] The power grid regulation resilience quantification unit is used to collect real-time control command data of the transformer voltage regulation equipment to be controlled, and calculate the cumulative resilience attenuation index based on the real-time control command data.
[0062] The data attack identification unit is used to simultaneously acquire raw voltage measurement signals from external IoT sensors and expected control command signals generated by the heterogeneous control strategy execution unit; and to perform spectral correlation analysis on the raw voltage measurement signals and expected control command signals to obtain the attack correlation index.
[0063] The control handover decision unit is used to collect the cumulative resilience decay index and attack-related index to calculate the handover trigger value; and to generate the final control output based on the comparison result between the handover trigger value and the preset handover threshold.
[0064] The heterogeneous control policy execution unit is also used to switch between deep reinforcement learning policies and classical robust control policies in response to the final control output.
[0065] This embodiment provides a transformer voltage regulation collaborative control system based on Internet of Things data acquisition. The system includes: a power grid regulation resilience quantification unit, a data attack identification unit, a control power switching decision unit, and a heterogeneous control strategy execution unit.
[0066] The purpose of the power grid regulation resilience quantification unit is to quantify in real time the physical fatigue losses caused by high-frequency or large-amplitude regulation operations of key power grid regulation equipment such as transformer on-load tap changers (OLTC) and capacitor banks. In this embodiment, the unit is used to collect real-time control command data of the transformer tap changer to be controlled, and calculate the cumulative resilience attenuation index based on the real-time control command data. This index is the key basis for the subsequent decision-making unit to evaluate the physical consequences.
[0067] The data attack identification unit aims to identify data injection attacks targeting the control system, particularly high-frequency ghost load attacks, i.e., high-frequency noise injected into sensors. In this embodiment, the unit is used to simultaneously acquire raw voltage measurement signals from external IoT sensors and expected control command signals generated by a heterogeneous control strategy execution unit, where the expected control command signals specifically refer to signals generated by the DRL strategy. The unit performs spectral correlation analysis on these two signals to obtain an attack correlation index, which is used to quantify whether an attack has occurred.
[0068] The heterogeneous control strategy execution unit aims to provide two control strategies with distinct functions and objectives, serving as the main executor of the system. In this embodiment, this unit defines a deep reinforcement learning (DRL) strategy, which aims for maximum short-term efficiency, and a classical robust control (RC) strategy, which aims to ensure long-term security, and generates expected control command signals for use by the data attack identification unit. Furthermore, this unit is also used to switch between the deep reinforcement learning strategy and the classical robust control strategy in response to the final control output generated by the control switching decision unit. To ensure the continuity of input to the data attack identification unit, the heterogeneous control strategy execution unit computes the deep reinforcement learning strategy signal in parallel at any given time. Regardless of whether the signal is ultimately executed, it is transmitted to the data attack identification unit to calculate the attack-related index.
[0069] The control switching decision unit is intended to act as the decision-making core of the system, determining the appropriate control strategy based on risk assessment. In this embodiment, the unit aggregates the cumulative resilience decay index from the power grid regulation resilience quantification unit and the attack correlation index from the data attack identification unit. Based on these two inputs, the unit calculates the switching trigger value. The unit generates the final control output, i.e., the decision command, based on the comparison between the switching trigger value and the preset switching threshold, and sends it to the heterogeneous control strategy execution unit.
[0070] This embodiment constructs a closed-loop control system of perception-decision-execution through the collaborative work of the above four units. It can not only detect hidden data attacks (i.e., perceive threats) through the data attack identification unit, but also assess the physical consequences (i.e., perceived state) caused by the attack through the power grid regulation resilience quantification unit. Based on the dual judgment of attack confirmation and physical consequences, the control switching decision unit intelligently switches between the DRL strategy that pursues short-term efficiency and the RC strategy that ensures long-term security. This solves the problem in the prior art that AI controllers are vulnerable to data attacks, leading to accelerated equipment damage, and significantly improves the antifragility and operational resilience of the power grid in the Internet of Things environment.
[0071] Example 2:
[0072] The process of calculating the cumulative resilience attenuation index by the power grid regulation resilience quantification unit includes:
[0073] Based on real-time control command data, the adjustment frequency and adjustment amplitude of the transformer voltage regulating equipment to be controlled are obtained;
[0074] The toughness attenuation increment is calculated based on the adjustment frequency, adjustment amplitude, preset frequency reference, and preset weight parameters;
[0075] Among them, the toughness attenuation increment is used to quantify the physical fatigue loss caused by adjustment actions and adjustment amplitudes that exceed the frequency reference.
[0076] The cumulative toughness decay index is obtained by iterating the incremental toughness decay over time.
[0077] Based on Example 1, the specific implementation process of the power grid regulation resilience quantification unit for calculating the cumulative resilience attenuation index is as follows:
[0078] Based on the collected real-time control command data, this unit obtains the adjustment frequency of the transformer voltage regulating equipment to be controlled. and adjustment range ;
[0079] Based on the adjustment frequency Adjustment range Preset frequency reference and preset weight parameters , Calculate the toughness attenuation increment The increase in toughness attenuation The calculation model used to quantify the physical fatigue loss caused by adjustment actions and amplitudes exceeding the frequency reference can be specifically implemented as a nonlinear fatigue accumulation model in this embodiment, for example... ;
[0080] in: This represents the physical loss within a single calculation cycle; To adjust the frequency, it represents the number of actions per unit time, which is obtained from real-time control command data acquisition; The adjustment range, such as the change in tap position, is also obtained from real-time control command data acquisition. The preset frequency reference is a calibration parameter that represents the normal adjustment frequency that the equipment can withstand. Its value is determined based on the physical design life of the voltage regulating equipment, such as the maximum number of operations allowed per day. Here, the Herveside step function ensures that the frequency adjustment only occurs when the frequency is adjusted. Greater than the frequency reference Only then is the first loss factor taken into account, thus distinguishing between normal adjustment and high-frequency harmful adjustment;
[0081] and These are preset weighting parameters, calibration parameters, used to balance the different contributions of frequency and amplitude to loss; to determine and The value of this parameter can be determined by those skilled in the art through accelerated aging experiments or finite element fatigue simulation; for example, the value can be obtained through simulation at different adjustment frequencies. and adjustment range Equipment fatigue loss data under combination ,in , and To calibrate the independent variables in the dataset, regression analysis techniques such as least squares were used to analyze the above-mentioned variables. The model is fitted to calibrate. and The value;
[0082] This formula quantifies, through a weighted summation, the risk of an exponential surge in device lifespan caused by high-frequency AI regulation, regardless of whether such regulation is caused by an attack.
[0083] By adjusting the toughness attenuation increment By performing cumulative iterations over time, the cumulative resilience decay index is obtained. This cumulative iterative model is used to track the real-time health status of devices, and can be specifically implemented as follows: ;in, Indicates equipment exist Accumulated physical wear and tear over time; This is the value from the previous time step, and its initial value. It can be set to 0, representing the initial health of the device; This represents the toughness attenuation increment calculated using the formula from the previous step.
[0084] This embodiment implements a quantifiable method for assessing equipment physical fatigue by defining a toughness attenuation increment and a cumulative iterative model. It enables the system to accurately track the implicit physical costs caused by high-frequency control behaviors, providing a crucial physical consequence assessment basis, i.e., an indicator, for subsequent control switching decisions. This prevents the system from reaching the failure boundary due to excessive equipment fatigue.
[0085] Example 3:
[0086] The process of spectral correlation analysis performed by the data attack identification unit includes:
[0087] The original voltage measurement signal and the expected control command signal are processed by time-frequency transformation respectively to extract the power spectral density of the signal in the preset ghost attack characteristic frequency band;
[0088] The power spectral density includes the voltage signal power spectrum and the control command power spectrum.
[0089] The process by which the data attack identification unit obtains attack-related indices includes:
[0090] Calculate the total power of the voltage signal power spectrum within the ghost attack characteristic frequency band, and the total power of the control command power spectrum within the ghost attack characteristic frequency band, respectively.
[0091] Multiplying the total power of the voltage signal by the total power of the control command yields the attack correlation index;
[0092] The attack correlation index is used to quantify the abnormal response of the expected control command signal to high-frequency noise in the original voltage measurement signal.
[0093] Based on Example 1, the process of the data attack identification unit performing spectral correlation analysis and calculating attack correlation index is specifically implemented in this example as follows:
[0094] During the spectral correlation analysis, this unit analyzes the synchronously acquired raw voltage measurement signals. and expected control command signal That is, the signals generated by the DRL strategy are subjected to time-frequency transformation processing, such as Short Time Fourier Transform (STFT) or Wavelet Transform.
[0095] By using time-frequency transformation, the two signals are extracted into the preset ghost attack characteristic frequency band. The power spectral density (PSD) within; here It is a specific high-frequency band that corresponds to the characteristic frequency of spoof noise injected by spectral payload attacks. Its value is determined in advance by performing spectral analysis on historical attack data or simulated injected spectral payload signals.
[0096] The power spectral density includes the power spectrum of the voltage signal. , its origin The transformed power spectrum and the control command power spectrum , its origin Obtained by transformation;
[0097] In obtaining the attack-related indices, this unit calculates the power spectrum of the voltage signal. In the ghost attack signature frequency band Total power within, and control command power spectrum Total power within the same frequency band;
[0098] This unit multiplies the total power of the voltage signal by the total power of the control command to obtain the attack correlation index. The attack-related index The calculation model used to quantify the abnormal response of the expected control command signal to high-frequency noise in the original voltage measurement signal can be specifically implemented as a band power product model in this embodiment. ;
[0099] in: The attack-related index, whose physical meaning is the product of frequency band power, has dimensions such as... ; The total power of the voltage signal represents the energy of the voltage signal within the attack characteristic frequency band, and its source is the... Obtained by performing time-frequency transformation and integration; The total power of the control command represents the energy of the DRL control action within the same frequency band, and its source is the energy from the control command. Obtained by performing time-frequency transformation and integration;
[0100] The technical motivation behind this formula lies in capturing a specific pattern of spectral payload attacks: that is, the attacker injects high-frequency spoof noise, leading to... The voltage increased, and the DRL controller was tricked into performing unnecessary high-frequency voltage regulation, resulting in... It also increases; only when both have high energy within the characteristic frequency band, their product increases. Only then will it increase significantly, indicating that an attack is taking place;
[0101] This embodiment introduces a frequency band power product model to provide a highly sensitive data attack identification mechanism; it can accurately identify abnormal resonance between the high-frequency noise of the DRL controller and the sensor, thus quantifying the occurrence of the attack, which is the indicator. This provides crucial attack confirmation signals for subsequent decision-making units, preventing the system from misjudging normal power grid fluctuations as attacks and improving the accuracy of decision-making.
[0102] Example 4:
[0103] The process by which the control handover decision unit calculates the handover trigger value includes:
[0104] Obtain the cumulative resilience decay index and attack-related index;
[0105] Calculate the rate of change of the cumulative toughness decay index to obtain the toughness decay rate;
[0106] Divide the attack-related index by the preset attack threshold to obtain the normalized attack index.
[0107] Divide the cumulative resilience decay index by a preset decay threshold and apply a first weight to obtain the weighted cumulative loss;
[0108] Divide the toughness decay rate by a preset rate threshold and apply a second weight to obtain the weighted instantaneous impact.
[0109] Adding the weighted cumulative loss to the weighted instantaneous impact yields the normalized physical consequences;
[0110] Multiply the normalized attack exponent by the normalized physical consequence to obtain the switching trigger value.
[0111] Based on Example 1, the control handover decision unit calculates the handover trigger value. The process is specifically implemented as follows in this embodiment:
[0112] This unit acquires the cumulative resilience degradation index from the power grid regulation resilience quantification unit. and attack-related indices from the data attack identification unit ;based on The unit calculates its rate of change to obtain the toughness decay rate. ; The toughness decay rate reflects the instantaneous impact of equipment loss, and its implementation can be achieved through... The first-order difference approximation, i.e. ,in To calculate the step size;
[0113] Attack-related indices Divide by the preset attack threshold Obtain the normalized attack index; The preset attack threshold is determined by testing on normal operating data. The baseline value is set to ensure a low false alarm rate;
[0114] To calculate the normalized physical consequences, the cumulative toughness degradation index is... Divide by the preset attenuation threshold and apply the first weight. The weighted cumulative loss is obtained; the toughness decay rate is then calculated. Divide by the preset rate threshold and apply a second weight. We obtain the weighted instantaneous impact; by adding the weighted cumulative loss to the weighted instantaneous impact, we obtain the normalized physical consequences. and These are preset attenuation thresholds and preset rate thresholds, respectively, whose values are determined based on hard constraints set according to equipment safety margins and mechanical fatigue limits. and These are the first and second weights, used to balance the importance of accumulated losses and instantaneous impacts in decision-making. Their values can be set by the power grid dispatch manager based on their risk preferences; for example, they can satisfy... Multiply the normalized attack exponent by the normalized physical consequence to obtain the switching trigger value. ;
[0115] This switching trigger value This is the core decision-making logic of the system, which integrates the outputs of the attack confirmation and physical consequence assessment modules. Its calculation model, in this embodiment, can be specifically implemented as follows: ;
[0116] This formula employs a multiplication-gated logic; the first term It is the gating factor confirmed by the attack; the weighted sum in parentheses in the second term This is a quantification of physical consequences; the technical motivation of this design is that the product of the two terms is only obtained when the attack is confirmed (i.e., the first term is greater than 1) and the physical consequences are severe (i.e., the weighted sum within the parentheses of the second term is also greater than 1), indicating that the cumulative loss or instantaneous impact has exceeded the threshold. Only then will it be significantly greater than 1, thus triggering a high-cost control mode switch;
[0117] This embodiment achieves a risk-based intelligent decision-making approach by designing a multiplication-gated switching trigger value calculation logic. It avoids the rigid logic of switching immediately upon detecting an attack in existing technologies, and instead realizes a strategic choice to switch only when the attack is confirmed and the physical consequences are severe. This ensures power grid security while maximizing the efficient operation time of the DRL strategy, achieving a dynamic balance between security and efficiency.
[0118] Example 5:
[0119] The process by which the control handover decision unit generates the final control output includes:
[0120] The switching trigger value is compared and analyzed with the preset switching threshold;
[0121] When the switching trigger value is greater than the preset switching threshold, it is determined to be a high-risk attack state, and the final control output switches to the classic robust control strategy.
[0122] When the switching trigger value is not greater than the preset switching threshold, it is determined to be a normal or low-risk state, and the final control output executes the deep reinforcement learning strategy.
[0123] Based on Example 5, the control switching decision unit generates the final control output. The process is specifically implemented as follows in this embodiment:
[0124] This unit will use the switching trigger value calculated in the previous step. A comparison analysis is performed with a preset switching threshold; here, the preset switching threshold is the decision execution threshold, and its value can be preset to 1 in this embodiment, compared with the previous embodiment. Corresponding to the normalized design;
[0125] When switching trigger values Greater than the preset switching threshold (i.e.) When this occurs, the system determines it to be a high-risk attack state; at this time, the final control output... Switch to classic robust control strategy ;
[0126] When switching trigger values Not greater than the preset switching threshold (i.e. When this condition is met, the system determines the state to be normal or low-risk; at this time, the final control output is... Execute deep reinforcement learning strategies ;
[0127] The generated final control output (i.e., the decision result is) or The data is sent to the heterogeneous control strategy execution unit, which then switches and executes it.
[0128] This embodiment achieves closed-loop switching execution of the control strategy by establishing explicit comparison logic between the switching trigger value and the preset threshold; it provides the system with a clear and unambiguous decision exit, ensuring that calculation is performed immediately upon perception. Switch to action immediately or maintain The reliable execution of the instruction chain is the final link in achieving antifragile control of the system.
[0129] Example 6:
[0130] Deep reinforcement learning strategies are defined as follows:
[0131] Construct a reward function that aims to maximize short-term operating efficiency;
[0132] Short-term operating efficiency includes renewable energy absorption rate and voltage qualification rate;
[0133] Deep reinforcement learning strategies generate actions that aim to maximize the reward function;
[0134] Classic robust control strategies are defined as follows:
[0135] Construct a cost function that minimizes security costs;
[0136] The cost function includes a penalty term for voltage drop below a preset safe voltage lower limit, and an action penalty term applied to the adjustment range performed by the transformer voltage regulating equipment under control.
[0137] Classic robust control strategies generate actions that aim to minimize the cost function;
[0138] The effect of minimizing the cost function is to prioritize ensuring that the voltage is not lower than the lower limit of the safe voltage and minimize equipment operation, while accepting a decrease in the renewable energy consumption rate.
[0139] The two strategies defined by the heterogeneous control strategy execution unit are specifically implemented as follows in this embodiment:
[0140] Deep Reinforcement Learning Strategy (DRL): The deep reinforcement learning strategy Defined as an optimal solution strategy that pursues extreme efficiency, it operates under normal or low-risk conditions (i.e., Execute under )
[0141] This strategy constructs a reward function that aims to maximize short-term operating efficiency. The reward function Aiming to maximize renewable energy absorption and voltage qualification rate, it can be specifically achieved as follows: ;
[0142] in: The objective of DRL optimization is the reward function; The renewable energy absorption rate is the source of renewable energy output, which is monitored or predicted in real time. The voltage is monitored and collected by an Internet of Things (IoT) sensor. This is a reference voltage, and its value is based on the standard parameters of power grid operation. The maximum permissible voltage deviation is determined based on the standard parameters for power grid operation. The weight is determined based on the economic dispatch objectives of the power grid, such as the carbon neutrality objective. The item is the voltage pass rate, which is used to quantify the degree to which the voltage deviates from the reference value;
[0143] Actions generated by deep reinforcement learning strategies The aim is to maximize this reward function Its effect is to actively and frequently adjust the transformer to maximize the absorption of fluctuating renewable energy sources. At the same time, we should try our best to maintain the voltage at a qualified level;
[0144] Classical Robust Control Strategy (RC): The classic robust control strategy Defined as a conservative, suboptimal security solution strategy that is insensitive to data noise, it is suitable for high-risk attack states (i.e., ) is activated;
[0145] This strategy constructs a cost function that aims to minimize the security cost. The cost function Aimed at sacrificing efficiency to ensure voltage safety and equipment health, it can be specifically implemented as follows: ;
[0146] in: The cost function is RC. The set of control actions to be optimized; To monitor voltage in real time; The lower limit of the preset safe voltage is determined based on the physical red line set according to the transient stability boundary of the power grid; For the Hervisside step function, ensure that the first penalty term applies only at voltage. Falling below the safe lower limit It is only activated at certain times; For the specific adjustment range of device i, it specifically refers to the adjustment range made by... The range of motion to be performed; This is the action penalty coefficient; to ensure dimensional consistency, the first term has the following dimensions. ,therefore The dimensions of the quantity must be specified as ,in for The quantization unit, thus making The dimensions of the term are also ; The value is set to a very large value to ensure that the device's actions take higher priority than voltage fluctuations;
[0147] Actions generated by classic robust control strategies The aim is to minimize this cost function ;because The value is extremely large, and the system is minimized. It will become extremely slow, for The system is insensitive to high-frequency noise (which may be caused by an attack) because any adjustment action... All of these will bring about huge punishments;
[0148] This embodiment defines two diametrically opposed objectives, namely... and The control strategy provides the system with clear aggressive and conservative operating modes; the DRL strategy ensures efficient and economical operation of the system under normal conditions; the robust control strategy ensures that, under high-risk conditions, the system will prioritize ensuring that the voltage does not fall below the safe lower limit. And minimize device actions To protect equipment while accepting renewable energy integration rates The reduced cost; this heterogeneous strategy is the foundation for implementing antifragile control.
[0149] Example 7:
[0150] This system also includes recovery logic, which includes:
[0151] When the final control output is in the classic robust control strategy, continuously monitor the attack-related index and the cumulative resilience decay index.
[0152] Determine whether the attack-related index is below the attack threshold and continuously preset the recovery time;
[0153] Determine whether the cumulative resilience degradation index has fallen below the preset safe recovery threshold;
[0154] When both decisions are true, the trigger value will be reset so that the final control output is automatically returned to the deep reinforcement learning policy.
[0155] To enable the automatic transfer of control, this embodiment also includes a set of recovery logic, which is implemented as follows:
[0156] When the final control output Classic robust control strategy In this mode, the system initiates recovery monitoring;
[0157] The system continuously monitors attack-related indices from the data attack identification unit. and the cumulative resilience degradation index from the grid regulation resilience quantification unit ;
[0158] Determine the attack-related index Is it below the attack threshold? And whether this state lasted for a preset recovery time (denoted as...). ); here It is a time window used to ensure that the attack has truly ended and is not just a brief interruption, preventing control from being jittered at the end of the attack;
[0159] Determining the cumulative toughness degradation index Has it fallen back to the preset safe recovery threshold? The following; here It is below the decay threshold The threshold (i.e.) The value is determined based on ensuring that the accumulated physical fatigue of the equipment has been adequately cooled or recovered before the DRL high-frequency control is restored, thus providing a safety margin.
[0160] The system will switch the trigger value only if both of the above conditions are met. Reset (e.g., force reset to 0) to make the final control output... Automatically return to deep reinforcement learning strategy ;
[0161] This embodiment introduces a dual-judgment recovery logic based on the end of the attack and device recovery, realizing a complete control loop of perception-decision-execution-recovery. It ensures that after the system switches to the secure mode (RC strategy), it will not remain in an inefficient state indefinitely. Instead, after confirming that the threat has been eliminated and the physical system (device) has sufficient tolerance, it will automatically recover to the efficient operating mode (DRL strategy), realizing dynamic recovery of system resilience and complete anti-fragility capabilities.
[0162] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A transformer voltage regulation collaborative control system based on Internet of Things (IoT) data acquisition, characterized in that, include: The system includes a power grid regulation resilience quantification unit, a data attack identification unit, a control handover decision-making unit, and a heterogeneous control strategy execution unit. The heterogeneous control policy execution unit is used to define deep reinforcement learning policies and classical robust control policies, and to generate expected control command signals; The power grid regulation resilience quantification unit is used to collect real-time control command data of the transformer voltage regulation equipment to be controlled, and calculate the cumulative resilience attenuation index based on the real-time control command data. The data attack identification unit is used to simultaneously acquire raw voltage measurement signals from external IoT sensors and expected control command signals generated by the heterogeneous control strategy execution unit; and to perform spectral correlation analysis on the raw voltage measurement signals and expected control command signals to obtain the attack correlation index. The control switching decision unit is used to aggregate the cumulative resilience decay index and attack-related index to calculate the switching trigger value. It is used to generate the final control output based on the comparison result between the switching trigger value and the preset switching threshold; The heterogeneous control policy execution unit is also used to switch between deep reinforcement learning policies and classical robust control policies in response to the final control output.
2. The transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 1, characterized in that, The process by which the power grid regulation resilience quantification unit calculates the cumulative resilience attenuation index includes: Based on real-time control command data, the adjustment frequency and adjustment amplitude of the transformer voltage regulating equipment to be controlled are obtained; The toughness attenuation increment is calculated based on the adjustment frequency, adjustment amplitude, preset frequency reference, and preset weight parameters; Among them, the toughness attenuation increment is used to quantify the physical fatigue loss caused by adjustment actions and adjustment amplitudes that exceed the frequency reference. The cumulative toughness decay index is obtained by iterating the incremental toughness decay over time.
3. The transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 1, characterized in that, The process of spectral correlation analysis performed by the data attack identification unit includes: The original voltage measurement signal and the expected control command signal are processed by time-frequency transformation respectively to extract the power spectral density of the signal in the preset ghost attack characteristic frequency band; The power spectral density includes the voltage signal power spectrum and the control command power spectrum.
4. A transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 3, characterized in that, The process by which the data attack identification unit obtains the attack-related index includes: Calculate the total power of the voltage signal power spectrum within the ghost attack characteristic frequency band, and the total power of the control command power spectrum within the ghost attack characteristic frequency band, respectively. Multiplying the total power of the voltage signal by the total power of the control command yields the attack correlation index; The attack correlation index is used to quantify the abnormal response of the expected control command signal to high-frequency noise in the original voltage measurement signal.
5. A transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 1, characterized in that, The process by which the control handover decision unit calculates the handover trigger value includes: Obtain the cumulative resilience decay index and attack-related index; Calculate the rate of change of the cumulative toughness decay index to obtain the toughness decay rate; Divide the attack-related index by the preset attack threshold to obtain the normalized attack index. Divide the cumulative resilience decay index by a preset decay threshold and apply a first weight to obtain the weighted cumulative loss; Divide the toughness decay rate by a preset rate threshold and apply a second weight to obtain the weighted instantaneous impact. Adding the weighted cumulative loss to the weighted instantaneous impact yields the normalized physical consequences; Multiply the normalized attack exponent by the normalized physical consequence to obtain the switching trigger value.
6. A transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 5, characterized in that, The process by which the control switching decision unit generates the final control output includes: The switching trigger value is compared and analyzed with the preset switching threshold; When the switching trigger value is greater than the preset switching threshold, it is determined to be a high-risk attack state, and the final control output switches to the classic robust control strategy. When the switching trigger value is not greater than the preset switching threshold, it is determined to be a normal or low-risk state, and the final control output executes the deep reinforcement learning strategy.
7. A transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 6, characterized in that, The deep reinforcement learning strategy is defined as follows: Construct a reward function that aims to maximize short-term operating efficiency; Short-term operating efficiency includes renewable energy absorption rate and voltage qualification rate; The actions generated by deep reinforcement learning strategies aim to maximize the reward function.
8. A transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 6, characterized in that, The classic robust control strategy is defined as follows: Construct a cost function that minimizes security costs; The cost function includes a penalty term for voltage drop below a preset safe voltage lower limit, and an action penalty term applied to the adjustment range performed by the transformer voltage regulating equipment under control. Classic robust control strategies generate actions that aim to minimize the cost function; The effect of minimizing the cost function is to prioritize ensuring that the voltage is not lower than the lower limit of the safe voltage and minimize equipment operation, while accepting a decrease in the renewable energy consumption rate.
9. A transformer voltage regulation collaborative control system based on Internet of Things data acquisition according to claim 6, characterized in that, It also includes recovery logic, which includes: When the final control output is in the classic robust control strategy, continuously monitor the attack-related index and the cumulative resilience decay index. Determine whether the attack-related index is below the attack threshold and continuously preset the recovery time; Determine whether the cumulative resilience degradation index has fallen below the preset safe recovery threshold; When both decisions are true, the trigger value will be reset so that the final control output is automatically returned to the deep reinforcement learning policy.