Method and system for dynamic response scheduling of virtual power plant based on cognitive spectrum network
By optimizing communication through cognitive spectrum network self-organization and reinforcement learning algorithms, and combining Lagrange relaxation method and risk transmission model, the problems of unreliable communication, high computational complexity and weak risk management of virtual power plants are solved, realizing an adaptive and robust scheduling system and improving system performance and risk control capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 四川电力设计咨询有限责任公司
- Filing Date
- 2026-02-13
- Publication Date
- 2026-04-24
AI Technical Summary
Existing virtual power plant technology suffers from poor communication reliability, high computational complexity, and weak adaptability in complex electromagnetic environments. Furthermore, it lacks an effective risk management mechanism, resulting in limited system performance and susceptibility to uncertain events.
A cognitive spectrum network self-organizing network is adopted to sense interference in real time through sensor nodes, dynamically select the optimal channel, optimize communication strategy using reinforcement learning algorithm, decompose optimization problem by combining Lagrange relaxation method, and use asynchronous iterative parallel algorithm and risk transmission model for resource grouping and market risk hedging to build an adaptive and robust scheduling system.
It improves the communication reliability and computing efficiency of virtual power plants, achieves efficient risk management, and builds a new generation of scheduling system with high adaptability and intelligence, adapting to the large-scale utilization of distributed resources.
Smart Images

Figure CN121710223B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual power plant technology, specifically a dynamic response scheduling method and system for virtual power plants based on cognitive spectrum networks. Background Technology
[0002] With the large-scale integration of distributed renewable energy and the deepening of new power system and power market reforms, virtual power plants (VPS) have attracted widespread attention as an important technological means to aggregate decentralized resources and participate in the power market. Existing VPS technologies mainly rely on fixed-frequency wireless communication networks for resource coordination, but they face serious communication reliability issues in complex electromagnetic environments. Frequent channel interference and network congestion lead to large data transmission delays and high packet loss rates, affecting the system's real-time scheduling performance and stable operation. Furthermore, the existing communication architecture lacks an adaptive mechanism and cannot dynamically adjust communication strategies according to the spectrum environment, resulting in a sharp deterioration in communication quality under high-density deployment scenarios.
[0003] In terms of optimization algorithms, traditional virtual power plant scheduling often employs centralized optimization methods, requiring all resource information to be processed centrally and allocated uniformly. This results in high computational complexity and slow convergence speed, making it difficult to meet the real-time scheduling needs of large-scale distributed resources. While existing distributed optimization algorithms alleviate computational pressure to some extent, most are based on synchronous iteration mechanisms, requiring strict synchronous updates from each node. This ignores the impact of differences in the computing capabilities of resource nodes and network communication latency in the actual system, leading to overall system performance being limited by the slowest node. Furthermore, existing technologies generally lack effective risk management mechanisms, failing to identify and prevent risk transmission between resources. Under the impact of uncertain events such as market price fluctuations and equipment failures, the system is prone to chain reactions and large-scale losses. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for dynamic response scheduling of virtual power plants based on cognitive spectrum networks, which can improve the reliability, adaptability, computational efficiency and risk management capabilities of virtual power plant communication systems.
[0005] The technical solution adopted by this invention to solve its technical problem is: a dynamic response scheduling method for virtual power plants based on cognitive spectrum networks, comprising:
[0006] S1. Establish a cognitive spectrum self-organizing network to connect various power resources under the virtual power plant. The cognitive spectrum self-organizing network perceives the interference status of the communication frequency band in real time through sensor nodes, dynamically selects the optimal channel, and learns the state, action and reward function with the help of reinforcement learning algorithm, so that each node can autonomously optimize the communication strategy and realize the automatic reconstruction of the network topology according to the communication environment.
[0007] S2. Collect current operating status data of each power resource under the virtual power plant and historical operating status data under different operating conditions; the operating status data includes target power, actual power, rated power, output power, input power, power loss, voltage, current and temperature parameters;
[0008] S3. Using the collected current operating status data of each power resource and the historical operating status data under different operating conditions, calculate the resource response characteristics of each power resource.
[0009] S4. Based on the obtained resource response characteristics of each power resource, construct an objective function and constraints that consider resource energy efficiency ratio, response time, adjustment range and reliability. Use the Lagrange relaxation method to decompose the global optimization problem of the virtual power plant into multiple local sub-problems, and solve them in parallel by each resource node using an asynchronous iterative parallel algorithm, and output the preliminary optimization results.
[0010] S5. For the asynchronous iterative parallel algorithm, the convergence and stability of its asynchronous iterative process are verified online using the Lyapunov method to obtain reliable optimization results.
[0011] S6. Based on the aforementioned credible optimization results, construct a graph theory-based risk transmission model, abstract the virtual power plant into a directed graph, calculate the edge weights to represent the risk transmission intensity by calculating the electrical correlation, market correlation, geographical correlation and technical correlation between resources, and use spectral clustering algorithm and Laplace matrix eigenvalue decomposition to dynamically group resources into risk isolation islands and generate a system risk partitioning topology report.
[0012] S7. Based on the reliable optimization results and the resource response characteristics of each power resource, a virtual power plant resource value option model is constructed using Black-Scholes option pricing theory, and a GARCH model incorporating leverage is used to estimate real-time electricity price volatility. Furthermore, combining the obtained resource response characteristics of each power resource, the optimal dynamic hedging ratio considering the interaction risk between resources is calculated. Then, based on this, a multi-dimensional risk hedging strategy is constructed with risk value, conditional risk value, and transaction costs as constraints. Finally, based on the set dynamic rebalancing trigger conditions and real-time hedging cost optimization algorithm, an automatically executable market risk hedging strategy is generated. The hedging strategy includes the hedging ratio and reserved capacity for each resource.
[0013] S8. Integrate the trusted optimization results, the system risk partitioning topology report, and the market risk hedging strategy to generate a final scheduling instruction that includes the power setpoints, execution times, hedging operations, and risk constraints for each resource, and issue it for execution through the cognitive spectrum self-organizing network.
[0014] Furthermore, the cognitive spectrum self-organizing network senses the interference situation of the communication frequency band in real time through sensor nodes and dynamically selects the optimal channel, including the following steps:
[0015] The network sensor node adopts an adaptive detection threshold calculation method, which dynamically adjusts the detection threshold based on the real-time estimated noise and signal power, and outputs the optimal adaptive threshold that achieves the best balance between false alarm and missed detection probabilities.
[0016] Based on the optimal adaptive threshold of the output, a local perception algorithm that integrates energy detection and feature value detection is used to make a decision, and the local decision result of each node is output.
[0017] Based on the local decision results of multiple nodes, the signal-to-noise ratio and historical reliability are weighted and fused to output a global channel occupancy state whose reliability meets the preset requirements.
[0018] Based on the output global channel occupancy status, a global spectrum map reflecting the spatiotemporal dynamics of the entire network spectrum is generated using spatial interpolation and time prediction algorithms.
[0019] Using the generated real-time spectrum map as input, a multi-dimensional evaluation model integrating signal-to-noise ratio, interference level, stability and availability is constructed, and the comprehensive quality score of all available channels is calculated and output.
[0020] The algorithm takes the overall quality score of all available channels, the primary user status, and the channel utility difference as input, and uses a spectrum switching decision algorithm to determine the optimal channel selection command.
[0021] Furthermore, the reinforcement learning algorithm is the Q-learning algorithm.
[0022] Further, step S3 includes:
[0023] S31. Based on the historical operating status data of each power resource collected, the parameters are identified in real time, and a multi-dimensional feature vector containing response time, adjustment range, adjustment accuracy, energy efficiency ratio and reliability indicators is constructed to form a dynamic response fingerprint database for each power resource.
[0024] S32. Based on the dynamic response fingerprint database of each power resource, the current operating status data of each power resource is collected, and the best historical pattern is retrieved from the dynamic response fingerprint database using a similarity matching algorithm to obtain the resource response characteristics of each power resource.
[0025] Furthermore, based on the collected historical operating status data, the parameters are identified in real time using the recursive least squares method.
[0026] Furthermore, the formula for constructing the multidimensional feature vector is as follows:
[0027] ;
[0028] ;
[0029] ;
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] In the formula, For resources eigenvectors; For resources Response time metrics; and For resources The vertical adjustment range; For resources The adjustment precision; For resources The energy efficiency ratio; For resources Reliability indicators; , Resources Target power, actual power, rated power; , Resources Target power and actual power at the j-th time step in history; This represents the number of historical data sampling points. For the first One time step; The attenuation coefficient; This represents the total number of state variables. For the first Weighting factors for each state variable; For resources The Deviation of each state variable; For resources ; output power; For resources Input power; For resources Power loss.
[0035] Furthermore, based on the obtained resource response characteristics of each power resource, an objective function and constraints considering resource energy efficiency ratio, response time, regulation range, and reliability are constructed. The objective function is:
[0036] ;
[0037] In the formula, For resources Decision variables, including resources Power setting value and execution time ; Total number of resources; For resources exerting effort Operating costs below; For resources The energy efficiency ratio; For resources Response time; For resources The change in power; For resources The adjustment precision; For resources Target power; For resources Reliability; For resources The level of risk; These are the weighting coefficients;
[0038] Constraints:
[0039] (1) Power balance constraint:
[0040] ;
[0041] In the formula, This represents the total active power load demand of the entire virtual power plant system at a specific time t.
[0042] (2) Power constraints based on adjustment range:
[0043] ;
[0044] In the formula, For resources Rated power; and For resources The vertical adjustment range;
[0045] (3) Timing constraints based on response time:
[0046] ;
[0047] In the formula, For resources Execution time, For response time, This constraint, set as the scheduling deadline, ensures that resources can respond in a timely manner.
[0048] (4) Reliability-based standby capacity constraints:
[0049] ;
[0050] In the formula, For the first One resource group; For resources The spare capacity; For the first Group's reserve capacity factor; For resources The reserve capacity adjustment coefficient;
[0051] (5) Power deviation constraint based on adjustment accuracy:
[0052] ;
[0053] In the formula, For resources The maximum permissible power deviation;
[0054] (6) Slope rate constraint:
[0055] ;
[0056] In the formula, For resources Maximum gradeability; For resources The execution time interval.
[0057] Furthermore, the similarity matching algorithm is as follows:
[0058] ;
[0059] ;
[0060] In the formula, For similarity functions; This is the current feature vector; The first in the historical fingerprint database 1 eigenvector; This is a similarity adjustment parameter; It is the Euclidean norm; For the best matching index; This refers to the fingerprint database capacity.
[0061] Furthermore, the asynchronous iterative parallel algorithm is as follows:
[0062] ;
[0063] ;
[0064] ;
[0065] ;
[0066] In the formula, For resources In the Decision variables for the next iteration; For the first The equality constraints on the 1st Lagrange multipliers in the next iteration; For the first The inequality constraint on the th inequality constraint Lagrange multipliers in the next iteration; For the first The step size of the equality constraint in the next iteration; For the first The step size of the inequality constraints in the next iteration; For the first The inequality constraint on the th inequality constraint Lagrange multipliers in the next iteration; For the first The equality constraints on the 1st Lagrange multipliers in the next iteration; For resource nodes Local iteration counter; For resources The feasible domain; For resources The local Lagrangian function; For the first The equality constraint function on the th... The value of the next iteration; For the first The inequality constraint function is in the th inequality constraint function at the th . The value of the next iteration; For the first The step size of the equality constraint in the next iteration; For the first The step size of the inequality constraints in the next iteration; This indicates projection onto the non-negative quadrant; The step size decay parameter for equality constraints; is the step size decay parameter for inequality constraints.
[0067] A cognitive spectrum network-based virtual power plant dynamic response scheduling system is used to execute the aforementioned cognitive spectrum network-based virtual power plant dynamic response scheduling method, including:
[0068] A cognitive spectrum self-organizing network construction module is configured to execute step S1.
[0069] The resource data acquisition module is configured to execute step S2.
[0070] The resource cognitive modeling module is configured to execute the S3 step;
[0071] A collaborative optimization solution module is configured to execute step S4.
[0072] An algorithm convergence guarantee module is configured to execute step S5.
[0073] The system risk quantification and isolation module is configured to execute step S6.
[0074] The market risk hedging module is configured to execute step S7; and,
[0075] The scheduling instruction generation and execution module is configured to execute the S8 step.
[0076] The beneficial effects of this invention are as follows: Through deep integration and innovation in multiple dimensions such as communication, computing, and risk management, this invention not only solves the pain points of unreliable communication, low computing efficiency, and weak risk management in traditional virtual power plant scheduling, but also constructs a new generation of virtual power plant scheduling system with high adaptability, strong robustness, and intelligence, providing a practical and feasible technical path for the large-scale and efficient utilization of distributed resources in the context of the energy internet. Attached Figure Description
[0077] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0078] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0079] like Figure 1 As shown, the present invention provides a dynamic response scheduling method for virtual power plants based on cognitive spectrum networks, comprising:
[0080] S1. Establish a cognitive spectrum self-organizing network to connect various power resources under the virtual power plant. The cognitive spectrum self-organizing network perceives the interference status of the communication frequency band in real time through sensor nodes, dynamically selects the optimal channel, and learns the state, action and reward function with the help of reinforcement learning algorithm, so that each node can autonomously optimize the communication strategy and realize the automatic reconstruction of the network topology according to the communication environment.
[0081] S2. Collect current operating status data of each power resource under the virtual power plant and historical operating status data under different operating conditions; the operating status data includes target power, actual power, rated power, output power, input power, power loss, voltage, current and temperature parameters; target power, actual power and rated power are used for the calculation of response time, adjustment range and adjustment accuracy in subsequent step S3, output power, input power and power loss are used for the calculation of energy efficiency ratio in subsequent step S3, and voltage, current and temperature state variables are used for the calculation of reliability indicators in subsequent step S3;
[0082] S3. Using the current operating status data of each power resource collected in step S2 and the historical operating status data under different operating conditions, calculate the resource response characteristics of each power resource.
[0083] S4. Based on the resource response characteristics of each power resource obtained in step S3, construct an objective function and constraints that consider resource energy efficiency ratio, response time, adjustment range and reliability. Use the Lagrange relaxation method to decompose the global optimization problem of the virtual power plant into multiple local sub-problems, and solve them in parallel by each resource node using an asynchronous iterative parallel algorithm, and output the preliminary optimization results.
[0084] S5. For the asynchronous iterative parallel algorithm described in step S4, the convergence and stability of its asynchronous iterative process are verified online using the Lyapunov method. If the verification is successful, the preliminary optimization result is considered a reliable optimization result, thus obtaining a reliable optimization result.
[0085] S6. Based on the credible optimization results obtained in step S5, construct a risk transmission model based on graph theory, abstract the virtual power plant into a directed graph, obtain edge weights by calculating the electrical correlation, market correlation, geographical correlation and technical correlation between resources to represent the risk transmission intensity, and use spectral clustering algorithm and Laplace matrix eigenvalue decomposition to dynamically group resources into risk isolation islands and generate a system risk partitioning topology report.
[0086] S7. Based on the reliable optimization results obtained in step S5 and the resource response characteristics of each power resource obtained in step S3, a virtual power plant resource value option model is constructed using Black-Scholes option pricing theory, and a GARCH model incorporating leverage is used to estimate real-time electricity price volatility. Furthermore, combining the obtained resource response characteristics of each power resource, the optimal dynamic hedging ratio considering the interaction risk between resources is calculated. Then, based on this, a multi-dimensional risk hedging strategy is constructed with risk value, conditional risk value, and transaction costs as constraints. Finally, based on the set dynamic rebalancing trigger conditions and real-time hedging cost optimization algorithm, an automatically executable market risk hedging strategy is generated. The hedging strategy includes the hedging ratio and reserved capacity for each resource.
[0087] S8. By integrating the trusted optimization results obtained in step S5, the system risk partitioning topology report generated in step S6, and the market risk hedging strategy generated in step S7, a final scheduling instruction is generated, which includes the power setpoints, execution times, hedging operations, and risk constraints of each resource. This instruction is then issued and executed through the cognitive spectrum self-organizing network.
[0088] Specifically: the safety factor is adjusted using the trusted optimization results, the power coordination constraint of resources within the same risk isolation island is applied using the system risk partition topology report, the capacity reservation is adjusted using the market risk hedging strategy, and the final scheduling power value of each resource is obtained by combining the results, thereby generating a final scheduling instruction that includes the power setting value, execution time, hedging operation and risk constraint conditions of each resource.
[0089] The following example illustrates step S8: Assume the reliable optimization result gives the initial scheduling of resource A as follows: output 10kW, execution time... +5min; The system risk zoning topology report indicates that the isolation island where resource A is located is a high-risk area, requiring the output limit to be tightened to 8kW and the ramp rate to not exceed 2kW / min; therefore, the power setpoint of resource A is first corrected to 8kW. Simultaneously, the market risk hedging strategy gives the current risk budget requirement to reduce the total exposure by 20%, so the power shortfall of 2kW is allocated to low-risk resource B for compensation; based on the resource response characteristics (response time) of resource B... =3min), the execution time of resource B is determined to be +2 minutes (start early to ensure) (Target output reached in 5 minutes); based on the hedging ratio calculated using the market risk hedging strategy. =0.6, generating the corresponding hedging operation: purchase an option hedging position of 2kW × 0.6 = 1.2kW, executed at time... Simultaneously, based on risk constraints, power constraints are set for resource A. Slope rate constraint Set a spare capacity constraint for resource B. (Meets VaR risk budget).
[0090] The scheduling method of this invention identifies the dynamic characteristic parameters of each distributed power resource in a virtual power plant under different operating conditions in real time, constructs a multi-dimensional feature vector including response time, adjustment range, adjustment accuracy, energy efficiency ratio, and reliability indicators, thereby forming a dynamic response fingerprint database. Then, a similarity matching algorithm is used to quickly retrieve historical response patterns from the database, achieving accurate prediction of resource adjustment capabilities and response times. This effectively solves the problems of static resource models and low prediction accuracy in traditional methods, providing high-quality input information for subsequent optimized scheduling.
[0091] Meanwhile, this invention establishes a cognitive spectrum self-organizing network to dynamically select the communication channel with the least interference; each sensor node acts as a reinforcement learning agent, learning the optimal communication strategy through reinforcement learning algorithms, enabling the network topology to be automatically reconstructed according to the spectrum environment and communication quality. This design achieves a high degree of adaptability in the communication link, significantly improving the system's reliability and anti-interference capability in complex electromagnetic environments.
[0092] At the optimization solution level, this invention employs the Lagrange relaxation method to decompose the global optimization problem of the virtual power plant into multiple local subproblems, allowing each resource node to asynchronously iterate and update decision variables at its own pace, and gradually converge through information exchange between neighboring nodes. Furthermore, by using the Lyapunov method to construct the energy function, the convergence and stability of the asynchronous iterative algorithm are rigorously guaranteed from a mathematical theory perspective. This joint strategy overcomes the performance bottleneck of synchronous optimization and significantly improves the collaborative optimization efficiency of large-scale distributed resources.
[0093] In terms of risk management, this invention first establishes a graph-based risk transmission model, abstracting the virtual power plant into a directed graph (where the edge weights represent the risk transmission strength). Then, a spectral clustering algorithm is used to dynamically group resources into risk isolation islands, forming a risk isolation architecture characterized by "strong coupling within groups and weak coupling between groups." Building upon this, a dynamic hedging strategy is further designed based on option pricing theory. The optimal hedging ratio is calculated according to real-time electricity price volatility and resource response characteristics, ultimately achieving proactive risk management and refined control.
[0094] In this embodiment of the invention, the cognitive spectrum self-organizing network senses communication frequency bands in real time through sensor nodes, dynamically selects the optimal channel, and uses reinforcement learning algorithms to enable each node to autonomously optimize its communication strategy, thereby achieving automatic reconfiguration of the network topology based on the communication environment. Specifically, the steps include:
[0095] a. The network sensor nodes adopt an adaptive detection threshold calculation method, which dynamically adjusts the detection threshold based on the real-time estimated noise and signal power, and outputs the optimal adaptive threshold that achieves the best balance between false alarm and missed detection probabilities.
[0096] b. Based on the optimal adaptive threshold of the output, a local perception algorithm that integrates energy detection and feature value detection is used to make a decision and output the local decision result of each node.
[0097] c. Based on the local decision results of multiple output nodes, the signal-to-noise ratio and historical reliability are weighted and fused to output a high-reliability global channel occupancy state that meets the preset reliability requirements.
[0098] d. Based on the output global channel occupancy status, a global spectrum map reflecting the spatiotemporal dynamics of the entire network spectrum is generated using spatial interpolation and time prediction algorithms.
[0099] e. Using the generated real-time spectrum map as input, construct a multi-dimensional evaluation model that integrates signal-to-noise ratio, interference level, stability, and availability, and calculate and output the comprehensive quality score of all available channels;
[0100] f. Take the overall quality score of all available channels, the primary user status, and the channel utility difference as input, and use the spectrum switching decision algorithm to make a judgment and output the optimal channel selection command.
[0101] The adaptive detection threshold calculation method dynamically adjusts the detection threshold based on real-time estimated noise and signal power, achieving an optimal balance between false alarm probability and detection probability. While fixed thresholds suffer performance degradation with noise power fluctuations, the adaptive threshold automatically adapts to environmental changes by minimizing the weighted costs of false alarms and missed detections. This method receives statistics from local sensing detection and historical power estimates as input, outputs the optimal detection threshold, and feeds it back to the local sensing algorithm to update the decision criteria, thereby providing more accurate channel occupancy status information for subsequent channel quality assessment.
[0102] In this embodiment of the invention, the adaptive detection threshold calculation method is as follows:
[0103] False alarm probability Calculation formula: ;
[0104] Detection probability Calculation formula: ;
[0105] Optimal threshold optimization objective function : ;
[0106] Noise statistics parameters: ; ;
[0107] Signal statistical parameters : ; ;
[0108] In the formula, The energy detection threshold; This represents the average noise power. The standard deviation of the noise power; This represents the average signal power. The standard deviation of the signal power; For the false alarm-false miss balance parameter, ; This represents the probability of a false alarm. For detection probability; This is the Q function (complementary cumulative distribution function of the standard normal distribution); The power of additive white Gaussian noise; The power of the primary user signal; This represents the number of sampling points.
[0109] In this invention, the local perception algorithm is specifically as follows:
[0110] ; ;
[0111] ;
[0112] ;
[0113] ;
[0114] Fusion Judgment Criteria: ;
[0115] In the formula, There is no assumption that the primary user is the main user. The assumption exists regarding the primary user; For sampling point index, ; For the first The received signal at each sampling point; Primary user signal; This is the channel impulse response; It is additive white Gaussian noise; This is an energy detection statistic; This is the eigenvalue detection statistic; The received signal covariance matrix; express Hermitian transpose; The energy detection threshold; The threshold for feature value detection; This represents the total number of sampling points; Represents the largest eigenvalue of the matrix; This represents the convolution operation.
[0116] In the weighted fusion step of the local decision results of multiple output nodes in this invention, based on their signal-to-noise ratio and historical reliability, the specific fusion algorithm is as follows:
[0117] ;
[0118] ;
[0119] ;
[0120] ;
[0121] ;
[0122] ;
[0123] ;
[0124] In the formula, For nodes Local detection statistics; For nodes The local detection threshold; For nodes The local judgment result is 1, indicating that the main user was detected, and 0, indicating that it was not detected. For nodes At any moment The fusion weight; For nodes At any moment Signal-to-noise ratio; The total number of nodes participating in the integration; The result is determined by "or" fusion decision; if any node detects the primary user, the result is set to 1. For the AND-based fusion judgment result, all nodes must detect the primary user before the result is judged as 1; The result is a weighted and integrated judgment. This is the global fusion threshold; This represents the global false alarm probability. This represents the global false negative probability. The parameter is used to weigh the costs of false alarms and missed detections. For nodes At any moment Reliability; For nodes Historical detection accuracy; For nodes At any moment Information age; This is the reliability decay time constant.
[0125] The spatial interpolation algorithm in this invention is as follows:
[0126] ;
[0127] in, For nodes In frequency ,time The power measurement value; For nodes to interpolation point The Euclidean distance; This is a spatial decay parameter that controls the rate at which the influence decays with distance. For dynamic spatial weights;
[0128] ;
[0129] This weight is determined by the node's historical detection reliability. and instantaneous signal-to-noise ratio Together, we decided to ensure that data from highly reliable and high-quality nodes would have a higher weight in the interpolation process.
[0130] The time prediction algorithm is as follows:
[0131] ;
[0132] ;
[0133] ;
[0134] In the formula, Indicates the first Each cognitive node at time For frequency The predicted value of the channel occupancy probability (or channel power / occupancy intensity index); This represents the predicted value for the same frequency point at the previous moment; Indicates time For frequency The actual measured value at the location; This is a time smoothing coefficient used to balance the weights of historical forecasts and current measurements. The larger the value, the more it relies on historical trends; the smaller the value, the more it emphasizes current observations, thus achieving an exponentially weighted update of the spectrum occupancy over time. This represents the reciprocal of the update frequency or update period of the Spectrum Map. and These represent the minimum and maximum allowable update frequency boundaries, respectively, used to limit the update frequency to be no less than the system's minimum refresh requirement and no more than the upper limit of communication and computing capabilities; To update the frequency adjustment coefficient, which is used to map the fluctuation level of the spectrum map to the updated frequency; It represents the variance or volatility measure of the current spectrum map, used to characterize the rate of environmental change. The larger the variance, the more drastic the change in spectrum occupancy, and the higher the update frequency should be. Indicates the prediction step size / prediction time domain, through candidate step sizes Select from the set smallest To determine, among which For a true spectrum map in the future The state at any given moment, For the future Prediction results of the time-frequency spectrum map. The mean squared error is used to quantify the prediction error and adaptively select the optimal prediction time domain accordingly.
[0135] Master user status: If the master user exists, the channel is occupied, and the value is 1; if the master user does not exist, the channel is idle, and the value is 0.
[0136] Poor channel efficiency: ;
[0137] In the formula: The channel with the highest utility among all currently available channels, calculated by the dynamic channel selection strategy, at time [time value missing]. The utility value; The channel currently in use At any moment The utility value.
[0138] and All originate from utility functions The calculation results.
[0139] ;
[0140] ;
[0141] ;
[0142] ;
[0143] ;
[0144] In the formula, For at any time The optimal channel selected; The set of available channels; For channel At any moment Utility function; For the weight parameters, satisfying and ; For channel At any moment The overall quality score; For channel At any moment Expected throughput; This represents the total number of channel states. For channel At any moment In state The probability of; For channel In state Below the time Data rate; To switch to channel At any moment Switching costs; The channel currently in use; Fixed switching cost coefficient; This is the variable switching cost coefficient; For channel At any moment Channel diversity entropy; For the set of neighboring nodes; Neighboring nodes At any moment Use channels The probability of.
[0145] The present invention constructs a multi-dimensional evaluation model that integrates signal-to-noise ratio, interference level, stability, and availability, specifically as follows:
[0146] ;
[0147] ;
[0148] ;
[0149] ;
[0150] ;
[0151] In the formula, For channel index; The current moment; For channel At any moment The overall quality score; For channel At any moment Normalized signal-to-noise ratio quality score; For channel At any moment The anti-interference quality score; For channel At any moment Stability mass fraction; For channel At any moment Usability quality score; The weighting coefficients and , ; For channel At any moment Signal-to-interference-plus-noise ratio; and These are the minimum and maximum signal-to-interference-plus-noise ratio thresholds used for normalization, respectively; For channel For the channel At any moment The intensity of interference; Interference threshold; Indicates time window Internal channel signal-to-noise ratio quality The variance, where The time variable within the time window; The length of the time window for stability assessment; This is a stability normalization parameter; For channel At any moment The probability of being busy (the probability of being occupied by the main user); For channel At any moment The fading probability.
[0152] In this invention, the spectrum switching decision algorithm is as follows:
[0153] ;
[0154] Condition 1: The current channel quality is below the threshold.
[0155] ;
[0156] Condition 2: Utility difference exceeds the switching threshold
[0157] ;
[0158] Condition 3: Primary user detected
[0159] ;
[0160] Condition 4: Prolonged period of low quality
[0161] ;
[0162] In the formula, For at any time The handover decision is set to 1 to trigger a handover and 0 to maintain the current channel. This is the quality threshold; The switching threshold; For at any time For the currently used channel The main user detection flag, 1 indicates that a main user has been detected, and 0 indicates that no main user has been detected; Duration of continuous low quality; Maximum tolerance time; This is the index of the currently used channel.
[0163] In this invention, the Q-learning algorithm is specifically used to learn the state-action-reward function, enabling each node to autonomously optimize its communication strategy.
[0164] Specifically, the state space of the Q-learning algorithm is defined as follows:
[0165] ;
[0166] , ;
[0167] , ;
[0168] ;
[0169] ;
[0170] ;
[0171] In the formula, This represents the state space of the Q-learning algorithm. For the first The system state at any given moment; This is the channel occupancy vector. This represents the total number of available channels. Indicates channel At any moment The occupancy status is indicated by 1 for occupied and 0 for idle. It is a network connectivity vector. The total number of neighboring nodes. Indicates the relationship with neighboring nodes At any moment The connection quality, value range The closer the value is to 1, the better the connection quality. The communication quality vector includes the signal-to-noise ratio. Delay Packet loss rate ; The energy consumption state vector includes the remaining energy. and energy consumption rate ; The load state vector includes CPU load. Memory load and cache load .
[0172] The action space of the Q-learning algorithm is defined as follows:
[0173] ;
[0174] ;
[0175] ;
[0176] ;
[0177] The reward function of the Q-learning algorithm is defined as follows:
[0178] ;
[0179] ;
[0180] ;
[0181] ;
[0182] ;
[0183] In the formula, This represents the action space of the Q-learning algorithm. For the first Momentary action selection; In the first The communication channel selected at any time ; This represents the total number of available channels. For the first The transmit power level at any given time; Minimum transmission power; This is the maximum transmission power; For the first Multi-hop routing path selection at any time; This represents the total number of available routes. In the state Next action The comprehensive reward function; The weighting coefficients and ; As a reward for communication quality, the channel is used directly. At any moment Overall quality score ; As a reward for energy efficiency; Rewards for resistance to interference; For channel For the selected channel At any moment The intensity of interference; Interference threshold; Rewards for network topology; This represents the total number of neighboring nodes. In order to be with the first The neighboring nodes at time... Connection quality; Neighboring nodes The topological weight coefficients.
[0184] The Q-function update formula of the Q-learning algorithm is as follows:
[0185] ;
[0186] ;
[0187] ;
[0188] In the formula, The state-action value function represents the state... Next action Expected cumulative reward; For learning rate, Control the speed at which new information is updated; As a discount factor, This is used to balance the importance of immediate rewards and future rewards; For instant reward functions; To perform the action The state transitioned to in the next moment; For the next state The following are candidate actions; For action space; The maximum Q value for the next state; For the policy function, adopt - Greedy exploration strategy; For the first The exploration rate at the next iteration decays over time. Minimum exploration rate; The initial maximum exploration rate; For exploration rate decay coefficient; This represents the number of iterations or time steps.
[0189] In this embodiment of the invention, step S3 includes:
[0190] S31. Based on the historical operating status data of each power resource collected, the parameters are identified in real time, and a multi-dimensional feature vector containing response time, adjustment range, adjustment accuracy, energy efficiency ratio and reliability indicators is constructed to form a dynamic response fingerprint database for each power resource.
[0191] S32. Based on the dynamic response fingerprint database of each power resource, the current operating status data of each power resource is collected, and the best historical pattern is retrieved from the dynamic response fingerprint database using a similarity matching algorithm to obtain the resource response characteristics of each power resource.
[0192] The formula for the mapping relationship between identification parameters and feature vectors:
[0193] ;
[0194] in, For resources The autoregressive coefficient, For input coefficients, , These represent the model order. Based on the identification... The dynamic response of a resource can be expressed as:
[0195]
[0196] Furthermore, response time metrics With system poles (Depend on The relationship (determined) is:
[0197]
[0198] in, As the dominant pole of the system, This indicates taking the real part.
[0199] In this embodiment of the invention, based on the collected historical operating status data, the parameters are identified in real time using the recursive least squares method. Real-time identification using the recursive least squares method provides accurate and dynamically updated internal model parameters of the resource under the current operating conditions. These parameters are the foundation for understanding and predicting resource behavior. The multidimensional feature vector formula, based on these identification results and the original data, further processes the data to generate standardized, quantifiable, and directly usable resource characteristic "fingerprints" for subsequent matching and optimization. In other words, real-time identification is about "understanding" the resource, while feature vector construction is about "describing" the resource; the former is the prerequisite and input source for the latter. The final feature vectors are stored in a "dynamic response fingerprint database" for subsequent similarity matching to achieve accurate resource recognition.
[0200] The formula for parameter identification using recursive least squares is:
[0201] ;
[0202] ;
[0203] ;
[0204] In the formula, For the first The parameter estimation vector at time step; For the first The parameter estimation vector at time step; For the first The system output at that moment; For the first The regression vector at time step; For regression vectors Transpose of; For the first Gain matrix at time step; For the first The covariance matrix at time t; For the first The covariance matrix at time t; Forgetting factor, ; It is an identity matrix.
[0205] In this embodiment of the invention, the formula for constructing the multidimensional feature vector is specifically as follows:
[0206] ;
[0207] ;
[0208] ;
[0209] ;
[0210] ;
[0211] ;
[0212] ;
[0213] In the formula, For resources eigenvectors; For resources Response time metrics; and For resources The vertical adjustment range; For resources The adjustment precision; For resources The energy efficiency ratio; For resources Reliability indicators; , Resources Target power, actual power, rated power; , Resources Target power and actual power at the j-th time step in history; This represents the number of historical data sampling points. For the first One time step; The attenuation coefficient; This represents the total number of state variables. For the first Weighting factors for each state variable; For resources The Deviation of each state variable; For resources ; output power; For resources Input power; For resources Power loss.
[0214] In this embodiment of the invention, the similarity matching algorithm is:
[0215] ;
[0216] ;
[0217] In the formula, For similarity functions; This is the current feature vector; The first in the historical fingerprint database 1 eigenvector; This is a similarity adjustment parameter; It is the Euclidean norm; For the best matching index; This refers to the fingerprint database capacity.
[0218] The Lagrange relaxation method introduces Lagrange multipliers to relax the coupled global constraints into the local objective functions of each resource node, thereby decomposing the global optimization problem of the virtual power plant into multiple independently solvable local subproblems. Each resource node only needs to make optimization decisions based on local information and a small amount of neighbor information, which greatly reduces computational complexity and communication overhead.
[0219] Based on the obtained resource response characteristics of each power resource, an objective function and constraints are constructed considering resource energy efficiency ratio, response time, regulation range, and reliability. The objective function is:
[0220] ;
[0221] In the formula, For resources Decision variables, including resources Power setting value and execution time ; Total number of resources; For resources exerting effort Operating costs below; For resources The energy efficiency ratio; For resources Response time; For resources The change in power; For resources The adjustment precision; For resources Target power; For resources Reliability; For resources The level of risk; These are the weighting coefficients;
[0222] Constraints:
[0223] (1) Power balance constraint:
[0224] ;
[0225] In the formula, This represents the total active power load demand of the entire virtual power plant system at a specific time t.
[0226] (2) Power constraints based on adjustment range:
[0227] ;
[0228] In the formula, For resources Rated power; and For resources The vertical adjustment range;
[0229] (3) Timing constraints based on response time:
[0230] ;
[0231] In the formula, For resources Execution time, For response time, This constraint, set as the scheduling deadline, ensures that resources can respond in a timely manner.
[0232] (4) Reliability-based standby capacity constraints:
[0233] ;
[0234] In the formula, For the first One resource group; For resources The spare capacity; For the first Group's reserve capacity factor; For resources The reserve capacity adjustment coefficient;
[0235] (5) Power deviation constraint based on adjustment accuracy:
[0236] ;
[0237] In the formula, For resources The maximum permissible power deviation;
[0238] (6) Slope rate constraint:
[0239] ;
[0240] In the formula, For resources Maximum gradeability; For resources The execution time interval.
[0241] The objective function optimizes scheduling costs from multiple dimensions, including energy efficiency ratio, response time, adjustment accuracy, and reliability. The constraints limit resource output, timing, and reserve capacity by adjusting range, response time, reliability, and adjustment accuracy, thereby achieving refined scheduling based on resource response characteristics.
[0242] In this invention, the formula for the Lagrange relaxation method is:
[0243] ;
[0244] ;
[0245] ;
[0246] In the formula, It is a Lagrange function; This is a vector of global decision variables; For resources Decision variables; For resources The local objective function; and These are the Lagrange multiplier vectors for equality constraints and inequality constraints, respectively; For the first Lagrange multipliers with equality constraints; For the first Lagrange multipliers constrained by inequality; and For global equality and inequality constraint functions; For resources The local Lagrangian function; and For resources The corresponding Lagrange multiplier vectors; and They are respectively and Transpose of; , This is the constraint coefficient matrix; , To constrain the right-hand side; Total number of resources; The number of equality constraints; Let be the inequality constraint number.
[0247] In this invention, the Lyapunov method is used to verify the convergence and stability of its asynchronous iterative process online. The specific formula for verifying the convergence and stability of the Lyapunov method is as follows:
[0248] ;
[0249] ;
[0250] ;
[0251] Convergence condition:
[0252] ,when and ;
[0253] In the formula, It is a Lyapunov function; , , The first The decision variable vector, equality constraint Lagrange multiplier vector, and inequality constraint Lagrange multiplier vector for each iteration; , , This is the optimal solution; , These are the step sizes for equality constraints and inequality constraints, respectively. For expectation operators; and The first Second and third The Lyapunov function value in the next iteration; The convergence rate; The gradient of the Lagrange function; For the first The asynchronous error term of the next iteration; This is a global iteration counter; For nodes Local iteration counter; This indicates that for all resource nodes Iteration delay Take the maximum value; and These are positive constants related to system parameters; For network communication delay; For the first The step size of the next iteration; It is the Euclidean norm.
[0254] In this embodiment of the invention, the asynchronous iterative parallel algorithm is:
[0255] ;
[0256] ;
[0257] ;
[0258] ;
[0259] In the formula, For resources In the Decision variables for the next iteration; For the first The equality constraints on the 1st Lagrange multipliers in the next iteration; For the first The inequality constraint on the th inequality constraint Lagrange multipliers in the next iteration; For the first The step size of the equality constraint in the next iteration; For the first The step size of the inequality constraints in the next iteration; For the first The inequality constraint on the th inequality constraint Lagrange multipliers in the next iteration; For the first The equality constraints on the 1st Lagrange multipliers in the next iteration; For resource nodes Local iteration counter; For resources The feasible domain; For resources The local Lagrangian function; For the first The equality constraint function on the th... The value of the next iteration; For the first The inequality constraint function is in the th inequality constraint function at the th . The value of the next iteration; For the first The step size of the equality constraint in the next iteration; For the first The step size of the inequality constraints in the next iteration; This indicates projection onto the non-negative quadrant; The step size decay parameter for equality constraints; is the step size decay parameter for inequality constraints.
[0260] The risk transmission model mainly includes a risk transmission diagram model, risk correlation calculation, and dynamic updating of risk transmission intensity. In this invention, the risk transmission diagram model is as follows:
[0261] ;
[0262] ;
[0263] ;
[0264] ;
[0265] The formula for calculating the risk correlation is:
[0266] ;
[0267] ;
[0268] ;
[0269] ;
[0270] In the formula, This is a risk transmission diagram; This is a set of nodes, representing various power resources; This represents the total number of resource nodes. Indicates the first One resource node; Let be the set of edges, representing the risk transmission path; Represents a node To the node Risk transmission edge; Represents a node and nodes Risk correlation between them; This is the risk correlation threshold; a propagation edge is only established when the correlation exceeds this threshold. This is the weight matrix; For nodes To the node The intensity of risk transmission; The weighting coefficients and ; Electrical relevance; Pearson correlation coefficient; For nodes At any moment The actual power; For nodes Electrical distance (busbar number); For electrical distance attenuation parameters; For market relevance; For nodes At any moment Real-time electricity price; For resources Type encoding; The maximum type difference; Geographical relevance; For nodes Geographic coordinates; The square of the Euclidean distance; Geographical attenuation parameter; For technical relevance; This is a technology type indicator function; it is 1 when two nodes have the same technology type, and 0 otherwise. For nodes The age of the equipment; This represents the maximum equipment age difference.
[0271] The dynamic update formula for the risk transmission intensity is:
[0272] ;
[0273] ;
[0274] ;
[0275] ;
[0276] In the formula, For the first Time Node To the node The intensity of risk transmission; For the first The intensity of risk transmission at any given moment; For time smoothing factor, Control the weights of historical values and new values; The newly calculated risk transmission strength; The strength of basic risk transmission (calculated from the aforementioned correlation); For the first Time Node and Risk change interaction items; For nodes In the The level of risk at any given moment; For risk sensitivity parameters; This is the time decay coefficient; The decay time constant; For nodes In the Volatility at any given moment; For nodes In the Actual power at any given moment; For nodes The predicted power; For nodes In the The probability of failure at any given moment; The weighting coefficients calculated for the risk level satisfy the following conditions: .
[0277] In this embodiment of the invention, the spectral clustering algorithm is specifically as follows:
[0278] ;
[0279] ;
[0280] , ;
[0281] ;
[0282] , ;
[0283] In the formula, The graph is a Laplace matrix; For degree matrix, ; This is the weight matrix (i.e., the risk transmission strength matrix in the aforementioned risk transmission diagram model). For the normalized Laplace matrix; Degree matrix of The power (i.e.) The inverse matrix of the square root); and For the first Each eigenvalue and its corresponding eigenvector; This represents the total number of resource nodes. The eigenvector matrix is formed by the previous... Minimum eigenvalues Corresponding feature vector Composed by columns; The number of clusters (i.e., the number of risk isolation islands that need to be divided); This is the normalized embedding matrix; For matrix The Line number Column elements; For matrix The Row vectors; for Norm (Euclidean norm).
[0284] The formula for determining the optimal number of clusters is:
[0285] ;
[0286] ;
[0287] In the formula, This is a modularity function; The total weight of the edges; For nodes The degree; For nodes Clustering labels; The Kronecker function; For the profile coefficient; The average distance within the cluster; The average distance between the nearest clusters; and These are the weight parameters.
[0288] In this invention, based on the reliable optimization results and the obtained resource response characteristics of each power resource, a virtual power plant resource value option model is constructed using Black-Scholes option pricing theory. Specifically, the virtual power plant resource value option model is as follows:
[0289] ;
[0290] ;
[0291] ;
[0292] ;
[0293] ;
[0294] In the formula, Total value of the virtual power plant; This is a vector of all resource prices; The current time; For resource indexing; Total number of resources; and Resources The value of call and put options; For resources The current price; For resources The exercise price (strike price); The risk-free rate; The expiration date; The remaining due date; For resources volatility; The cumulative distribution function of the standard normal distribution; and The intermediate parameter in the Black-Scholes option pricing formula (and the distance parameter mentioned above) Constraining the right-hand term (These are different concepts) is the base of the natural logarithm; It is the natural logarithm function; This represents the value of interaction between resources.
[0295] The specific formula for estimating real-time electricity price volatility using a GARCH model incorporating leverage effects is as follows:
[0296] ;
[0297] ;
[0298] ;
[0299] ;
[0300] In the formula, For resources At any moment The time-varying conditional variance (squared volatility); For GARCH model parameters; For residual terms; for A set of information at any given moment; For resources The long-term mean; These are autoregressive parameters; It is an exogenous variable; For market state variables; This is a lever effect indicator function.
[0301] In this invention, the formula for calculating the optimal dynamic hedging ratio is:
[0302] ;
[0303] In the formula, For resources At any moment The optimal dynamic hedging ratio; The hedging ratio is the decision variable; For variance operators; For changes in the value of virtual power plants; For resources Price changes.
[0304] The multidimensional risk hedging strategy is as follows:
[0305] ;
[0306] ;
[0307] ;
[0308] In the formula, For at any time The hedging strategy vector; For resources At any moment The hedging ratio; Total number of resources; For confidence level Changes in the value of the virtual power plant The risk value; Indicates the infimum; For probability; hedging strategy At any moment Transaction costs; For resources Fixed transaction cost coefficient; For resources The variable transaction cost coefficient.
[0309] The dynamic rebalancing triggering condition in this invention is:
[0310] , ;
[0311] ;
[0312] ;
[0313] In the formula, The hedging ratio deviates from the threshold; To reduce tolerance for loss; For the portfolio at time of Value (second-order risk sensitivity); For resources At any moment of value; For resources At any moment The weights; This is the maximum rebalancing interval threshold.
[0314] In this invention, the real-time hedging cost optimization algorithm is as follows:
[0315] ;
[0316] ;
[0317] ;
[0318] ;
[0319] ;
[0320] In the formula, For at any time Total hedging costs; Total number of resources; For resources At any moment Option costs; For resources At any moment Commission fees; For resources At any moment Slippage cost; For resources At any moment The hedging ratio; For resources At any moment The hedging ratio; For resources The option value function is the resource price. ,time and volatility The function; For resources At any moment The price; For resources At any moment volatility; For resources Transaction fee rates; This is the slippage coefficient; For at any time The optimal timing for execution; Candidate values for the execution time; The allowed execution time window length; Indicates at time Information set Expectation operator under given conditions.
[0321] This invention also provides a virtual power plant dynamic response scheduling system based on cognitive spectrum networks, used to execute the above-described virtual power plant dynamic response scheduling method based on cognitive spectrum networks, comprising:
[0322] A cognitive spectrum self-organizing network module is configured to perform step S1;
[0323] The resource data acquisition module is configured to execute step S2.
[0324] The resource cognitive modeling module is configured to execute the S3 step;
[0325] A collaborative optimization solution module is configured to execute step S4.
[0326] An algorithm convergence guarantee module is configured to execute step S5.
[0327] The system risk quantification and isolation module is configured to execute step S6.
[0328] The market risk hedging module is configured to execute step S7; and,
[0329] The scheduling instruction generation and execution module is configured to execute the S8 step.
Claims
1. A method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks, characterized in that, Includes the following steps: S1. Establish a cognitive spectrum self-organizing network to connect various power resources under the virtual power plant. The cognitive spectrum self-organizing network perceives the interference status of the communication frequency band in real time through sensor nodes, dynamically selects the optimal channel, and learns the state, action and reward function with the help of reinforcement learning algorithm, so that each node can autonomously optimize the communication strategy and realize the automatic reconstruction of the network topology according to the communication environment. S2. Collect current operating status data of each power resource under the virtual power plant and historical operating status data under different operating conditions; the operating status data includes target power, actual power, rated power, output power, input power, power loss, voltage, current and temperature parameters; S3. Using the collected current operating status data of each power resource and the historical operating status data under different operating conditions, calculate the resource response characteristics of each power resource. S4. Based on the resource response characteristics of each power resource, construct an objective function and constraints that consider resource energy efficiency ratio, response time, adjustment range and reliability. Use the Lagrange relaxation method to decompose the global optimization problem of the virtual power plant into multiple local sub-problems, and solve them in parallel by each resource node using an asynchronous iterative parallel algorithm, and output the preliminary optimization results. S5. For the asynchronous iterative parallel algorithm, the convergence and stability of its asynchronous iterative process are verified online using the Lyapunov method to obtain reliable optimization results. S6. Based on the aforementioned credible optimization results, construct a graph theory-based risk transmission model, abstract the virtual power plant into a directed graph, calculate the edge weights to represent the risk transmission intensity by calculating the electrical correlation, market correlation, geographical correlation and technical correlation between resources, and use spectral clustering algorithm and Laplace matrix eigenvalue decomposition to dynamically group resources into risk isolation islands and generate a system risk partitioning topology report. S7. Based on the reliable optimization results and the resource response characteristics of each power resource, a virtual power plant resource value option model is constructed using Black-Scholes option pricing theory, and a GARCH model incorporating leverage is used to estimate real-time electricity price volatility. Furthermore, combining the obtained resource response characteristics of each power resource, the optimal dynamic hedging ratio considering the interaction risk between resources is calculated. Then, based on this, a multi-dimensional risk hedging strategy is constructed with risk value, conditional risk value, and transaction costs as constraints. Finally, based on the set dynamic rebalancing trigger conditions and real-time hedging cost optimization algorithm, an automatically executable market risk hedging strategy is generated. The hedging strategy includes the hedging ratio and reserved capacity for each resource. S8. Integrate the trusted optimization results, the system risk partitioning topology report, and the market risk hedging strategy to generate a final scheduling instruction that includes the power setpoints, execution times, hedging operations, and risk constraints for each resource, and issue it for execution through the cognitive spectrum self-organizing network.
2. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 1, characterized in that, The cognitive spectrum self-organizing network senses the interference situation of communication frequency bands in real time through sensor nodes and dynamically selects the optimal channel, including the following steps: The network sensor node adopts an adaptive detection threshold calculation method, which dynamically adjusts the detection threshold based on the real-time estimated noise and signal power, and outputs the optimal adaptive threshold that achieves the best balance between false alarm and missed detection probabilities. Based on the optimal adaptive threshold of the output, a local perception algorithm that integrates energy detection and feature value detection is used to make a decision, and the local decision result of each node is output. Based on the local decision results of multiple nodes, the signal-to-noise ratio and historical reliability are weighted and fused to output a global channel occupancy state whose reliability meets the preset requirements. Based on the output global channel occupancy status, a global spectrum map reflecting the spatiotemporal dynamics of the entire network spectrum is generated using spatial interpolation and time prediction algorithms. Using the generated real-time spectrum map as input, a multi-dimensional evaluation model integrating signal-to-noise ratio, interference level, stability and availability is constructed, and the comprehensive quality score of all available channels is calculated and output. The algorithm takes the overall quality score of all available channels, the primary user status, and the channel utility difference as input, and then uses a spectrum switching decision algorithm to determine the optimal channel selection command.
3. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 1, characterized in that, The reinforcement learning algorithm is the Q-learning algorithm.
4. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 1, characterized in that, Step S3 includes: S31. Based on the historical operating status data of each power resource collected, the parameters are identified in real time, and a multi-dimensional feature vector containing response time, adjustment range, adjustment accuracy, energy efficiency ratio and reliability indicators is constructed to form a dynamic response fingerprint database for each power resource. S32. Based on the dynamic response fingerprint database of each power resource, the current operating status data of each power resource is collected, and the best historical pattern is retrieved from the dynamic response fingerprint database using a similarity matching algorithm to obtain the resource response characteristics of each power resource.
5. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 4, characterized in that, Based on the collected historical operating status data, the parameters are identified in real time using the recursive least squares method.
6. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 4, characterized in that, The formula for constructing the multidimensional feature vector is as follows: ; ; ; ; ; ; ; In the formula, For resources eigenvectors; For resources Response time metrics; and For resources The vertical adjustment range; For resources The adjustment precision; For resources The energy efficiency ratio; For resources Reliability indicators; , Resources Target power, actual power, rated power; , Resources Target power and actual power at the j-th time step in history; This represents the number of historical data sampling points. For the first One time step; The attenuation coefficient; This represents the total number of state variables. For the first Weighting factors for each state variable; For resources The Deviation of each state variable; For resources ; output power; For resources The input power; For resources Power loss.
7. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 1, characterized in that, Based on the obtained resource response characteristics of each power resource, an objective function and constraints are constructed considering resource energy efficiency ratio, response time, regulation range, and reliability. The objective function is: ; In the formula, For resources Decision variables, including resources Power setting value and execution time ; Total number of resources; For resources exerting effort Operating costs below; For resources The energy efficiency ratio; For resources Response time; For resources The change in power; For resources The adjustment precision; For resources Target power; For resources Reliability; For resources The level of risk; These are the weighting coefficients; Constraints: (1) Power balance constraint: ; In the formula, This represents the total active power load demand of the entire virtual power plant system at a specific time t. (2) Power constraints based on adjustment range: ; In the formula, For resources Rated power; and For resources The vertical adjustment range; (3) Timing constraints based on response time: ; In the formula, For resources Execution time, For response time, This constraint, set as the scheduling deadline, ensures that resources can respond in a timely manner. (4) Reliability-based standby capacity constraints: ; In the formula, For the first One resource group; For resources The spare capacity; For the first Group's reserve capacity factor; For resources The reserve capacity adjustment coefficient; (5) Power deviation constraint based on adjustment accuracy: ; In the formula, For resources The maximum permissible power deviation; (6) Slope rate constraint: ; In the formula, For resources Maximum gradeability; For resources The execution time interval.
8. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 4, characterized in that, The similarity matching algorithm is as follows: ; ; In the formula, It is a similarity function; This is the current feature vector; The first in the historical fingerprint database 1 eigenvector; This is a similarity adjustment parameter; It is the Euclidean norm; For the best matching index; This refers to the fingerprint database capacity.
9. The method for dynamic response scheduling of virtual power plants based on cognitive spectrum networks as described in claim 1, characterized in that, The asynchronous iterative parallel algorithm is as follows: ; ; ; ; In the formula, For resources In the Decision variables for the next iteration; For the first The equality constraints on the th Lagrange multipliers in the next iteration; For the first The inequality constraint on the th inequality constraint Lagrange multipliers in the next iteration; For the first The step size of the equality constraint in the next iteration; For the first The step size of the inequality constraints in the next iteration; For the first The inequality constraint on the th inequality constraint Lagrange multipliers in the next iteration; For the first The equality constraints on the th Lagrange multipliers in the next iteration; For resource nodes Local iteration counter; For resources The feasible domain; For resources The local Lagrangian function; For the first The equality constraint function on the th... The value of the next iteration; For the first The inequality constraint function is in the th inequality constraint function at the th . The value of the next iteration; For the first The step size of the equality constraint in the next iteration; For the first The step size of the inequality constraints in the next iteration; This indicates projection onto the non-negative quadrant; The step size decay parameter for equality constraints; is the step size decay parameter for inequality constraints.
10. A virtual power plant dynamic response scheduling system based on cognitive spectrum networks, used to execute the virtual power plant dynamic response scheduling method based on cognitive spectrum networks as described in any one of claims 1 to 9, characterized in that, include: A cognitive spectrum self-organizing network module is configured to perform step S1; The resource data acquisition module is configured to execute step S2. The resource cognitive modeling module is configured to execute the S3 step; A collaborative optimization solution module is configured to execute step S4. An algorithm convergence guarantee module is configured to execute step S5. The system risk quantification and isolation module is configured to execute step S6. The market risk hedging module is configured to execute the S7 step; as well as, The scheduling instruction generation and execution module is configured to execute the S8 step.
Citation Information
Patent Citations
Multi-time-scale comprehensive energy optimization scheduling method based on federated learning framework
CN118627817A
Virtual power plant collaborative optimization scheduling method, system and device based on multiple spatial-temporal scales and storage medium
CN120999696A