A reinforcement learning-based energy complex dynamic state aggregation method and system

By constructing a continuous-time dynamic system model and an aggregation strategy optimization model based on reinforcement learning, the problems of dynamic characteristic modeling and high-dimensional data processing of complex multi-dimensional heterogeneous energy systems are solved, and the precise scheduling and efficient aggregation of flexible resources within the energy complex are realized.

CN120996381BActive Publication Date: 2026-03-17GUIZHOU ANRONG TECH DEV CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing energy aggregation technologies struggle to accurately model the dynamic characteristics of complex, multi-element, heterogeneous energy systems. In particular, when faced with nonlinear, time-varying, and uncertain factors, traditional modeling methods lack precision and adaptability, and are unable to effectively process high-dimensional, non-equidistant time series data. This results in inaccurate assessment of resource response characteristics and affects the optimization effect of aggregation decisions.

Method used

A reinforcement learning-based approach is adopted to construct a continuous-time dynamic system model through the process modeling technique of neural ordinary differential equations. The aggregation problem is identified and classified by combining static spiking neural networks and adaptive dynamic spiking neural networks. The multi-marginal stochastic flow matching method is used to perform alignment analysis on energy data measured at non-equidistant time points. The twin-delay deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model.

Benefits of technology

It enables accurate modeling and efficient dynamic scheduling of complex, multi-element, heterogeneous energy systems, improves the accuracy of resource response characteristic assessment and the efficiency of aggregated decision-making, and can continuously adapt to system changes to achieve adaptive dynamic optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996381B_ABST
    Figure CN120996381B_ABST
Patent Text Reader

Abstract

The application provides a kind of energy complex flexible dynamic aggregation method and system based on reinforcement learning, which collects real-time operating parameters of electric, gas, heat and multi-energy coupling devices in the energy complex through industrial Ethernet protocol, uses neural differential equation process modeling technology to build a continuous-time dynamic system model, and accurately describes the nonlinear dynamic characteristics of various energy resources. The aggregation process is modeled as a Markov decision process, and an improved twin-delay deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation strategy model. The system realizes self-adaptive dynamic optimization through an online learning update mechanism, significantly improving the renewable energy consumption rate and system operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of energy management technology, specifically to a method and system for flexible and dynamic aggregation of energy complexes based on reinforcement learning. Background Technology

[0002] As complex systems integrating multiple energy forms, energy complexes have become an important development direction for modern energy management. With the continuous increase in the penetration rate of renewable energy and the diversification of energy demand, how to effectively aggregate and dispatch various flexible resources within energy complexes to achieve a coordinated balance between the system's economy, reliability, and environmental protection has become a core technological challenge in the current energy field.

[0003] Existing energy aggregation technologies mainly include static aggregation methods based on optimization algorithms and dynamic aggregation methods based on predictive control. Static aggregation methods typically employ linear programming or quadratic programming algorithms to determine the optimal combination of resources based on a preset objective function and constraints, but they are difficult to adapt to real-time changes in the system's operating state. While dynamic aggregation methods based on predictive control can consider future changes in the system's state, their prediction accuracy is limited by the model's accuracy, and their computational complexity is high, making it difficult to meet the requirements of real-time control.

[0004] Furthermore, the energy aggregation method based on model predictive control (MPC) is currently the primary approach. This technology establishes a mathematical model of the energy system, predicts the system state for future periods, and solves an optimization problem within each control cycle to determine the optimal control strategy. Its basic principle is to minimize system operating costs or maximize benefits within a finite time window, while simultaneously satisfying the operational constraints of various devices and system balance constraints, achieving dynamic scheduling through rolling optimization.

[0005] However, existing technologies have two key problems: First, they are difficult to accurately model the dynamic characteristics of complex, multi-element, heterogeneous energy systems, especially when faced with nonlinear, time-varying, and uncertain factors, where traditional modeling methods lack accuracy and adaptability. Second, they lack the ability to effectively process high-dimensional, non-equidistant time series data, leading to inaccurate assessment of resource response characteristics and affecting the optimization effect of aggregated decision-making. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a flexible dynamic aggregation method and system for energy complexes based on reinforcement learning. This method can accurately model the dynamic characteristics of complex multi-dimensional heterogeneous energy systems, effectively process high-dimensional, non-equidistant time series data, and realize the dynamic aggregation and optimized scheduling of flexible resources within the energy complex.

[0007] To achieve the above objectives, the technical solution provided by the present invention is as follows:

[0008] A flexible and dynamic aggregation method for energy complexes based on reinforcement learning, comprising:

[0009] The system acquires operational data of flexible resources within the energy complex, collects real-time operational parameters of electrical, gas, thermal, and multi-energy coupling equipment, and constructs a continuous-time dynamic system model using the neural network differential equation process modeling technique.

[0010] Based on the continuous-time dynamic system model, a static spiking neural network is used to initially identify energy aggregation problems. After identifying potential aggregation needs, an adaptive dynamic spiking neural network classifier is activated to accurately classify specific energy aggregation problems.

[0011] The different aggregation demand types in the accurate classification results of the specific energy aggregation problem are obtained. The multi-marginal stochastic flow matching method is used to perform alignment analysis on the energy data measured at non-equidistant time points. The response characteristics of each flexibility resource are quantitatively evaluated by measuring value spline enhancement technology and fraction matching. The response characteristics of each flexibility resource include the randomness of state transitions and the sequentiality of decision-making in the energy aggregation process.

[0012] The energy aggregation process is modeled as a Markov decision process, and the twin-delay deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model.

[0013] The reinforcement learning aggregation strategy optimization model receives state data in real time and outputs the optimal aggregation strategy. Using the optimal aggregation strategy, the distributed control system sends adjustment commands to each flexible resource and collects actual response data for online learning and updating, thereby achieving adaptive dynamic optimization.

[0014] Preferably, operational data of flexible resources within the energy complex are acquired, real-time operational parameters of electrical, gas, thermal, and multi-energy coupling equipment are collected, and a continuous-time dynamic system model is constructed using neural network differential equation process modeling techniques, including:

[0015] The operational data of each flexibility resource is obtained, and data points that exceed the normal range are removed using the three sigma rule. The difference in units is eliminated by Z-score standardization, and preprocessed operational data is generated.

[0016] Based on the preprocessed operating data, the energy system state evolution process is represented as a differential equation. A neural network with a residual network architecture is used to approximate the differential equation function and the gradient is calculated by the adjoint sensitivity method to establish a battery energy storage model, a gas turbine model and a thermal energy storage model.

[0017] A variable step-size ordinary differential equation solver was used to train the battery energy storage model, gas turbine model, and thermal energy storage model respectively, so as to obtain the continuous-time dynamic system model.

[0018] Preferably, based on the continuous-time dynamic system model, a static spiking neural network is used to initially identify the energy aggregation problem. After identifying potential aggregation demands, an adaptive dynamic spiking neural network classifier is activated to accurately classify specific energy aggregation problems, including:

[0019] Normalized state data is obtained, and pulse coding is performed using the leak-integral discharge neuron model through the static pulse neural network to identify four types of aggregated demands: load peak-valley regulation, frequency regulation, voltage support, and emergency response.

[0020] Based on the four types of aggregation requirements, an adaptive dynamic spiking neural network classifier is activated, and the network structure is dynamically adjusted according to the novelty of the input data through the adaptive dynamic spiking neural network classifier, and neurons are automatically added to expand the network capacity.

[0021] The connection weights of the dynamic spiking neural network classifier are updated by an adaptive peak timing-dependent plasticity rule, and the dynamic spiking neural network classifier is used to accurately classify specific energy aggregation problems.

[0022] Preferably, based on the four types of aggregation requirements, an adaptive dynamic spiking neural network classifier is activated, and the network structure is dynamically adjusted according to the novelty of the input data through the adaptive dynamic spiking neural network classifier, automatically adding neurons to expand the network capacity, including:

[0023] The feature vectors of the four types of aggregation requirements are obtained, the Euclidean distance between the input data and the existing neuron prototype vectors is calculated, and when the minimum distance exceeds the preset novelty threshold, it is determined to be a new mode, triggering the network structure adjustment mechanism.

[0024] Based on the network structure adjustment mechanism, a new hidden layer neuron is created in the corresponding output category. The feature vector of the current input data is used as the prototype vector of the new neuron, and the connection weights between the new neuron and the input and output layers are initialized.

[0025] The activation intensity of the new neurons in the network is updated through a competitive learning mechanism to determine the optimal response neuron and ultimately determine the expanded network capacity.

[0026] Preferably, a multilateral stochastic flow matching method is used to perform alignment analysis on energy data measured at non-equidistant time points. Through measured value spline augmentation and fractional matching, the response characteristics of each flexibility resource are quantitatively evaluated, including:

[0027] The data distribution characteristics of different types of energy equipment are obtained, and multiple marginal distributions of electric energy storage resources, thermal energy storage resources and gas energy storage resources are constructed. Nonparametric modeling is performed by kernel density estimation method to generate a smoothed distribution model.

[0028] Based on the smoothed distribution model, cubic B-spline basis functions are constructed using measured value spline technology to process the observations at irregular time points and generate continuous time series.

[0029] The fractional matching algorithm is used to calculate the difference between the minimum true data distribution of the continuous time series and the fractional function of the distribution model, thereby obtaining the response characteristics of each flexibility resource.

[0030] Preferably, a fractional matching algorithm is used to calculate the difference between the minimum true data distribution of the continuous time series and the fractional function of the distribution model to obtain the response characteristics of each flexibility resource, including:

[0031] Obtain the true data distribution of the continuous time series, calculate the log density gradient of the true data distribution and use the log density gradient of the true data distribution as the true fractional function, and at the same time calculate the log density gradient of the distribution model and use the log density gradient of the distribution model as the fractional function of the distribution model.

[0032] Based on the true fraction function and the fraction function of the distribution model, a fraction matching objective function is constructed. The least squares method is used to calculate the squared error between the true fraction function and the fraction function of the distribution model, and the calculation results are integrated and summed.

[0033] The model parameters of the score matching objective function are optimized by using the gradient descent algorithm, so that the score function of the distribution model approximates the true score function, and the response characteristics of each flexibility resource, including response time constant, adjustment accuracy and stability index, are obtained.

[0034] Preferably, based on the smoothed distribution model, a cubic B-spline basis function is constructed using measured value spline technology to process the observations at irregular time points, generating a continuous time series, including:

[0035] Obtain the irregular time point observations in the smoothed distribution model, determine the spline node sequence, set the node density according to the time interval distribution of the observation data, and increase the number of nodes in the data-dense region.

[0036] Based on the node sequence, cubic B-spline basis functions are constructed. Each B-spline basis function is a cubic polynomial in four consecutive node intervals and zero in other intervals.

[0037] The coefficients of the cubic B-spline basis function are determined by the least squares fitting method, so that the spline curve passes through or approximates all observation points, thus obtaining the continuous time series with time continuity and smoothness.

[0038] Preferably, a twin-delay deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model, including:

[0039] The energy aggregation process is modeled as a Markov decision process. A state space containing load demand, renewable energy output, and real-time resource status is constructed, as well as an action space containing resource combination and output allocation ratio. A reward function that integrates response speed, adjustment accuracy, and operating cost is designed. Based on the Markov decision process framework, state space, action space, and reward function, a twin-delay deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model.

[0040] Preferably, based on the Markov decision process framework, state space, action space, and reward function, a Siamese-delayed deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregated policy optimization model, including:

[0041] Acquire multidimensional characteristic data, construct a state space including load demand dimension, renewable energy dimension, energy storage status dimension, power quality dimension and economic signal dimension, and an action space including resource participation allocation vector and aggregation mode selection, and establish a Markov decision process framework.

[0042] Within the Markov decision process framework, a multi-objective reward function is designed, which includes a response speed reward, a regulation accuracy reward, an operating cost reward, a system stability reward, and an environmental benefit reward, thereby generating a comprehensive performance evaluation index.

[0043] Based on the comprehensive performance evaluation index, the actor network and the critic network are constructed using the improved twin-delay deep deterministic policy gradient algorithm, resulting in the optimized model of the reinforcement learning aggregation policy.

[0044] Preferably, the reinforcement learning aggregation strategy optimization model receives state data in real time and outputs the optimal aggregation strategy. Using the optimal aggregation strategy, adjustment commands are sent to each flexible resource through a distributed control system, and actual response data is collected for online learning and updating to achieve adaptive dynamic optimization, including:

[0045] The reinforcement learning aggregation strategy optimization model is used to obtain real-time state data. The feasibility of the strategy is verified by the constraint checking module. If the strategy violates the constraints, the action vector is projected into the feasible region using a projection algorithm to generate the optimal aggregation strategy.

[0046] Based on the optimal aggregation strategy, adjustment instructions are sent to each flexible resource through a communication protocol, and actual execution effect data is collected to obtain actual response data.

[0047] For the actual response data, the actual reward value is calculated and stored in the experience replay buffer, and online learning updates are initiated to achieve adaptive dynamic optimization of the system.

[0048] This invention also provides a flexible and dynamic aggregation system for energy complexes based on reinforcement learning, comprising:

[0049] The acquisition module is used to acquire operational data of flexible resources within the energy complex. It collects real-time operational parameters of electrical, gas, thermal, and multi-energy coupling equipment via the industrial Ethernet protocol and constructs a continuous-time dynamic system model using the neural network differential equation process modeling technique.

[0050] The identification module is used to perform preliminary identification of energy aggregation problems based on the continuous-time dynamic system model using a static spiking neural network. After identifying potential aggregation needs, it activates an adaptive dynamic spiking neural network classifier to accurately classify specific energy aggregation problems.

[0051] The analysis module is used to determine the corresponding data processing strategy and time window parameters by utilizing the different aggregation demand types in the accurate classification results of the specific energy aggregation problem, and to perform alignment analysis on energy data measured at non-equidistant time points using the multilateral stochastic flow matching method. Through measurement value spline enhancement technology and fraction matching, the response characteristics of each flexibility resource are quantitatively evaluated. The response characteristics of each flexibility resource include the randomness of state transitions and the sequentiality of decision-making in the energy aggregation process.

[0052] The module is used to model the energy aggregation process as a Markov decision process, and to train and optimize it using the twin-delay deep deterministic policy gradient algorithm to generate a reinforcement learning aggregation policy optimization model.

[0053] The output module is used to receive state data in real time and output the optimal aggregation strategy using the reinforcement learning aggregation strategy optimization model. Using the optimal aggregation strategy, it sends adjustment instructions to each flexible resource through the distributed control system and collects actual response data for online learning and updating to achieve adaptive dynamic optimization.

[0054] The technical solution provided by this invention has the following beneficial effects:

[0055] 1. This invention uses the process modeling technique of the divine constant differential equation to construct a continuous-time dynamic system model, which can accurately describe the dynamic characteristics of complex multi-element heterogeneous energy systems and effectively solve the problem of insufficient accuracy and adaptability of traditional modeling methods when facing nonlinear, time-varying and uncertain factors.

[0056] 2. This invention employs a two-layer spiking neural network for the identification and classification of aggregation problems. The static spiking neural network is responsible for the initial identification, while the adaptive dynamic spiking neural network is responsible for the precise classification. This structure features high computational efficiency and strong adaptability, and can accurately identify various energy aggregation demands.

[0057] 3. This invention employs a multilateral random flow matching method to perform alignment analysis on high-dimensional energy data measured at non-equidistant time points. Through measurement value spline enhancement technology and fractional matching algorithm, it effectively solves the problems of high-dimensional data processing and non-equidistant time series analysis, and improves the accuracy of resource response characteristic assessment.

[0058] 4. This invention models the energy aggregation process as a Markov decision process and uses an improved twin-delay deep deterministic policy gradient algorithm for training and optimization. This algorithm effectively improves the stability and convergence speed of policy learning through techniques such as dual-path network design, priority sampling, and adaptive soft update.

[0059] 5. This invention enables the reinforcement learning model to continuously adapt to system changes through an online learning and update mechanism, achieving adaptive dynamic optimization and significantly improving the efficiency and accuracy of flexible resource aggregation in energy complexes. Attached Figure Description

[0060] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0061] Figure 1 This is a flowchart of the flexible dynamic aggregation method for energy complexes based on reinforcement learning, as described in this invention.

[0062] Figure 2 This is a structural diagram of the flexible dynamic aggregation system for energy complexes based on reinforcement learning, as described in this invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0064] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0065] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0066] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0067] like Figure 1 As shown, this invention provides a flexible and dynamic aggregation method for energy complexes based on reinforcement learning, comprising the following steps:

[0068] S1: Acquire operational data of flexible resources within the energy complex, collect real-time operational parameters of electrical, gas, thermal, and multi-energy coupling equipment, and construct a continuous-time dynamic system model using the neural network differential equation process modeling technique.

[0069] An energy complex refers to a complex system that integrates the conversion, storage, transmission, and utilization of multiple energy forms, achieving coordinated and optimized operation of various energy sources such as electricity, heat, and natural gas through a unified energy management platform. This integrated system typically includes diverse energy production equipment, storage facilities, conversion devices, and end-use energy equipment, forming an interconnected and collaborative organic whole. The core characteristic of an energy complex lies in its multi-energy coupling nature; that is, different forms of energy can be mutually converted through various conversion devices, thereby improving overall energy utilization efficiency and enhancing the flexibility and reliability of system operation.

[0070] Within energy complexes, energy storage stations serve as crucial energy buffers, undertaking multiple functions such as peak shaving, frequency regulation, and voltage support. These energy storage facilities typically employ large-capacity lithium-ion battery systems or other advanced energy storage technologies, possessing rapid response and precise control capabilities. Distributed photovoltaic power stations utilize dispersed resources such as building rooftops and open spaces to convert solar energy into electricity and connect it to the grid nearby. Their output characteristics are significantly affected by weather conditions, exhibiting marked intermittency and fluctuation. Electric vehicle charging stations are not only energy consumption facilities; equipped with bidirectional charging and discharging technology, they can also participate in grid regulation as mobile energy storage resources, and their aggregation effect is becoming increasingly important in urban energy systems.

[0071] Flexible resources refer to various devices and systems that can rapidly adjust their energy production, consumption, or storage status according to system needs. These resources possess adjustable and controllable characteristics, enabling them to respond to system scheduling commands at different time scales. Among electrical flexible resources, battery energy storage systems, with their millisecond-level response speed and high-precision power control capabilities, have become the most important regulation resource. Although supercapacitors have relatively small energy storage capacity, their ultra-fast charging and discharging characteristics make them particularly suitable for handling instantaneous power fluctuations and frequency regulation tasks. Interruptible loads include energy-consuming equipment such as industrial production lines and large refrigeration equipment that can temporarily stop operating in emergencies, providing emergency regulation capabilities for the system.

[0072] Flexibility in gas resources primarily includes various gas storage and regulation facilities. Underground gas storage facilities leverage the natural advantages of geological structures to provide large-scale, long-term natural gas storage capacity, playing a crucial role in supply and demand balance and emergency preparedness. While surface gas storage facilities such as high-pressure spherical tanks have relatively smaller capacities, their regulation response is more flexible, allowing for rapid responses to short-term gas volume adjustments. The regulation characteristics of these gas resources are mainly reflected in pressure regulation, flow control, and storage management.

[0073] Thermal flexibility resources participate in system regulation through the storage and release of thermal energy. Thermal storage tanks utilize the sensible heat properties of water or other media to store thermal energy, playing a role in peak shaving and stable operation in heating systems. Thermal storage materials, especially phase change materials, can absorb or release large amounts of latent heat during phase change processes, achieving high-density thermal energy storage. The response characteristics of these thermal resources are affected by heat conduction and convection heat transfer processes, typically exhibiting a certain degree of thermal inertia, with relatively long response times but strong sustained response capabilities.

[0074] Multi-energy coupling devices are key hubs connecting different energy forms, enabling the conversion between them. Combined heat and power (CHP) units simultaneously generate electricity and heat, and by rationally adjusting their operating modes, they can be optimally allocated between the electricity and heat markets. Electric boilers convert electricity into heat, increasing electricity consumption during periods of renewable energy surplus while simultaneously providing a heat source for heating systems. The operating parameters of these coupling devices include multiple dimensions such as conversion efficiency, response time, and operating constraints.

[0075] The operational data for various flexible resources encompasses a wealth of technical parameters and status information. Real-time power reflects the instantaneous operating status of equipment and serves as the direct basis for system scheduling. Capacity parameters, including rated capacity, available capacity, and remaining capacity, determine the resource's adjustment potential. Response time describes the time required for equipment to reach its target state from receiving a command, directly affecting its applicability in different adjustment services. Charging and discharging efficiency reflects the loss characteristics during energy conversion, impacting the overall economic efficiency of the system. Operational constraints include technical boundary conditions such as maximum and minimum output limits, ramp rate limits, and start-stop time constraints. These constraints ensure that equipment operates within a safe and reliable range and are also crucial factors that must be considered when formulating scheduling strategies.

[0076] In this step, the flexible dynamic aggregation system of the energy complex (hereinafter referred to as the system) collects operational data of various flexible resources within the energy complex in real time through a high-speed data channel established by the industrial Ethernet protocol. For electrical resources, the system collects electrical parameters including voltage, current, active power, reactive power, and power factor, as well as characteristic parameters of battery energy storage devices such as state of charge, charge-discharge cycle count, and health status. The energy storage devices employ a high-frequency sampling rate of 10Hz to capture rapid dynamic response characteristics. For thermal resources, the system collects thermal parameters such as temperature, pressure, flow rate, and thermal power, as well as characteristic parameters of thermal energy storage devices such as heat storage capacity, heat loss rate, and heat transfer coefficient, using a standard sampling rate of 1Hz. For gaseous resources, the system collects gas parameters such as pressure, flow rate, composition, and calorific value, as well as characteristic parameters of gas energy storage devices such as gas storage capacity, compression ratio, and gas release rate. For multi-energy coupling devices, such as combined heat and power units, electric refrigeration / electric heating equipment, and hydrogen power generation equipment, the system simultaneously collects input and output parameters and conversion efficiencies of multiple energy forms.

[0077] The collected real-time operating parameters (raw data) undergo preprocessing, starting with outlier detection and removal. The system employs the three sigma (3σ) rule to calculate the mean μ and standard deviation σ for each data category, marking data points falling outside the range [μ-3σ, μ+3σ] as outliers and removing them. Next, data standardization is performed using the Z-score standardization method, converting data of different dimensions and scales into a standard distribution with a mean of 0 and a standard deviation of 1, eliminating the impact of dimensional differences on subsequent modeling. Finally, time series alignment is performed to address inconsistencies in sampling frequencies across different devices and parameters, establishing a unified time baseline.

[0078] Based on the preprocessed high-quality data, the system uses the neural network ordinary differential equation process modeling technique to construct a continuous-time dynamic system model. This technique combines traditional ordinary differential equations with deep neural networks, enabling it to accurately describe the dynamic characteristics of complex nonlinear systems.

[0079] First, the dynamic process of the energy equipment is expressed in the form of differential equations: Here, x(t) represents the system state vector, u(t) represents the control input vector, and θ represents the parameters to be learned. A deep neural network with a residual network architecture is then used to approximate the function f, and the gradient is calculated using the adjoint sensitivity method to achieve efficient parameter optimization.

[0080] Representing the dynamic process of energy equipment as differential equations is a fundamental step in establishing an accurate mathematical model. This representation stems from the continuous-time characteristics of physical systems, accurately describing the evolution of the system state over time. Within this differential equation framework, the system state vector contains all the key variables describing the current operating state of the energy equipment, such as the energy storage device's state of charge, temperature, internal resistance, and other internal state parameters. The control input vector represents the control signals applied to the system externally, such as controllable inputs like charging / discharging power setpoints and valve opening adjustments. The set of parameters to be learned contains all the parameters in the system model that need to be determined through data learning; these parameters may include the equipment's physical characteristics, conversion efficiency coefficients, time constants, etc.

[0081] Traditional physical modeling methods typically construct differential equations based on first principles and empirical formulas. However, this approach often struggles to accurately capture all nonlinear characteristics and coupling relationships in complex energy systems. The limitations of pure physical models become even more apparent when considering factors such as equipment aging, environmental changes, and diverse operating conditions. Therefore, this invention innovatively employs deep neural networks to approximate the state evolution functions in the differential equations. This method combines the interpretability of physical modeling with the powerful fitting capabilities of machine learning.

[0082] The introduction of the residual network architecture is the key innovation of this method. Through a skip connection mechanism, the residual network enables the network to learn state changes rather than absolute state values, which is highly consistent with the idea of ​​using differential equations to describe the derivatives of states. Specifically, each basic block of the residual network contains a shortcut connection with an identity mapping, making the network output the result of adding a residual function to the input. This design not only alleviates the gradient vanishing problem in deep network training, but more importantly, it matches the essence of differential equations in a physical sense, because differential equations themselves describe incremental changes in the system state.

[0083] The specific architecture of the neural network fully considers the characteristics of the energy system. The input layer of the network simultaneously receives the current state vector and the control input vector, forming a complete system input. The hidden layers adopt a multi-layer fully connected structure, with each layer equipped with an appropriate activation function and regularization mechanism. The choice of activation function is particularly important, requiring it to both express nonlinear characteristics and maintain good gradient propagation properties. The output layer of the network generates a state derivative vector, whose dimension is the same as the state vector, directly corresponding to the right-hand side function of the differential equation.

[0084] The adjoint sensitivity method is a core technique for training such neuromorphic differential equation models. This method originates from the adjoint equation theory in optimal control theory. By constructing an adjoint equation dual to the original differential equation, the gradient of the objective function with respect to the model parameters can be efficiently calculated. Specifically, the adjoint sensitivity method first solves the original differential equation forward to obtain the state trajectory, and then solves the adjoint equation backward to calculate the gradient information. The greatest advantage of this method is that its computational complexity is independent of the number of time steps; regardless of the integration time, the complexity of gradient calculation remains constant, which is extremely important for processing long-term energy system data.

[0085] The parameter optimization process is conducted iteratively, with each iteration consisting of two phases: forward propagation and back propagation. In the forward propagation phase, the system solves the differential equation using the current parameters to obtain the predicted state trajectory. In the back propagation phase, the gradient of the loss function with respect to all model parameters is calculated using the adjoint sensitivity method, and then the parameters are updated using gradient descent algorithms. This process continues until the model performance reaches a satisfactory level or the convergence condition is met. The entire optimization process maintains the continuity and interpretability of the physical model while fully leveraging the advantages of deep learning in handling complex nonlinear relationships, providing a new technical approach for energy system modeling.

[0086] Furthermore, in S1, for battery energy storage systems, the first model considers changes in state of charge, variations in internal resistance with temperature and aging, nonlinear characteristics of power conversion efficiency, and capacity decay effects. For gas turbines, the second model includes start-stop characteristics, operational inertia, output regulation rate limitations, and the characteristics of thermoelectric conversion efficiency varying with load. For thermal energy storage systems, the third model integrates the dynamic characteristics of heat transfer processes, heat loss mechanisms, and the nonlinear characteristics of phase change materials. These sub-models (the first, second, and third models) are trained and validated using a variable-step-size ordinary differential equation solver, ultimately integrating into a continuous-time system model capable of accurately describing the dynamic characteristics of various resources, providing a solid foundation for subsequent aggregation problem identification and strategy optimization.

[0087] S2: Based on the continuous-time dynamic system model, a static spiking neural network is used to initially identify the energy aggregation problem. After identifying the potential aggregation demand, an adaptive dynamic spiking neural network classifier is activated to accurately classify the specific energy aggregation problem.

[0088] Step S2 implements the energy aggregation problem identification and classification using a two-layer spiking neural network. This step innovatively employs biomimetic neural computing principles, simulating the information encoding and processing mechanisms of biological neurons to construct an efficient and robust framework for identifying energy aggregation problems. First, in S2, the state data generated from the continuous-time dynamic system model in step S1 is normalized to ensure that all input feature dimensions are unified within the [0,1] interval, thereby improving the network's learning efficiency and stability.

[0089] A static spiking neural network serves as the first layer of the recognizer, responsible for rapidly detecting potential aggregation needs in the system. This network employs a biologically sound Leaky Integrate-and-Fire (LIF) neuron model, utilizing the membrane potential dynamic equation... This simulates the working mechanism of a biological neuron. When a neuron receives an input stimulus, the membrane potential v gradually accumulates; when it reaches the threshold voltage Vth... th At this time, the neuron generates a pulse and resets the membrane potential to the resting potential V. rest This mechanism enables the network to efficiently process time-varying signals and extract temporal information features. The input layer contains 64 LIF neurons that receive system state data; the hidden layer contains 128 fully connected LIF neurons, with connection weights initialized using the Xavier method; the output layer has 4 neurons, corresponding to four basic aggregation needs: load peak-valley regulation, frequency regulation, voltage support, and emergency response. The network converts continuous state values ​​into pulse sequences through a time-coding mechanism and uses a winner-take-all mechanism to determine the final recognition result.

[0090] Once the static spiking neural network identifies a potential aggregation requirement, an adaptive dynamic spiking neural network classifier (the second-layer classifier) ​​is activated for accurate problem classification. This classifier employs a Growing When Required (GWR) mechanism, which dynamically adjusts the network structure based on the novelty of the input data. Specifically, it calculates the Euclidean distance between the input feature vector and the existing neuron prototype vectors, and when the minimum distance exceeds a preset novelty threshold θ... sim (Usually set to 0.8) and network activity a t Below the activity threshold θ act When the value is set to 0.1, the current input is determined to be a new pattern, and new neurons are automatically added to the network to expand its capacity. The initial weights of the new neurons are determined by linear interpolation, and their prototype vector is the average of the weights of the best-matching neuron and the second-best-matching neuron.

[0091] Network connection weights are updated using an Adaptive Spike-Timing-Dependent Plasticity (Ad-STDP) rule. Inspired by biological synaptic plasticity mechanisms, this rule adjusts synaptic strength based on the time difference between preceding and following neuronal pulses. The current neuronal pulse occurs before the following neuronal pulse. When the value is greater than 0, the connection weight is increased; conversely, it is decreased. The update formula is as follows: (Activity), among which Let A be the learning rate. The function f is time-dependent, and f(activity) is an adaptive function that dynamically adjusts the learning intensity based on the historical activity level of neurons, effectively preventing overlearning and catastrophic forgetting. The classifier output layer contains 12 neurons, corresponding to more specific aggregate sub-problem types, including short-term peak shaving, medium-term peak shaving, long-term energy storage, primary frequency adjustment, and secondary frequency adjustment, providing precise guidance for subsequent strategy formulation.

[0092] This dual-layer spiking neural network architecture fully combines the rapid response capability of static networks with the adaptive learning capability of dynamic networks. It can efficiently identify complex aggregated demands in energy systems and continuously improve classification performance as system operation experience is accumulated, forming a deep understanding of the operation mode of energy complexes.

[0093] S3: Obtain the different aggregation demand types in the accurate classification results of the specific energy aggregation problem, use the multilateral stochastic flow matching method to perform alignment analysis on the energy data measured at non-equidistant time points, and use the measurement value spline enhancement technology and fraction matching to quantitatively evaluate the response characteristics of each flexibility resource. The response characteristics of each flexibility resource include the randomness of state transitions and the sequentiality of decision-making in the energy aggregation process.

[0094] In this embodiment, step S3 achieves resource response characteristic assessment based on multilateral random flow matching. This step aims to solve the problem of high-dimensional data processing with non-equidistant time points commonly found in energy systems, and to accurately quantify the dynamic response characteristics of various flexible resources through advanced statistical learning methods.

[0095] Specifically, based on the specific aggregation problem type identified in step S2, the corresponding data processing strategy and time window parameters are determined. For rapid response needs (such as frequency regulation), a short time window (seconds to minutes) and a high sampling rate are used; for medium- to long-term regulation needs (such as load peak-valley regulation), a longer time window (hours) and a moderate sampling rate are used.

[0096] Then, multiple marginal distributions are constructed, each corresponding to data characteristics of a specific resource type or time period. For electrical energy storage resources, the marginal distribution includes key dimensions such as power output, response time, charge / discharge efficiency, and cycle life degradation coefficient. For thermal energy storage resources, the marginal distribution covers parameters such as temperature change gradient, thermal storage capacity, heat release rate, and heat loss coefficient. For gas energy storage resources, the marginal distribution includes indicators such as pressure change rate, volumetric flow rate, and compression efficiency. The system uses the kernel density estimation method for nonparametric modeling, makes no prior assumptions about the data distribution form, and uses a Gaussian kernel function. The original data points are smoothed. The bandwidth parameter h is determined through k-fold cross-validation to optimize the fit while avoiding overfitting.

[0097] To address the challenges of processing non-uniformly spaced time series data, a measured value spline technique is employed to significantly enhance the processing capability for irregular snapshot times. In practical energy systems, due to factors such as communication delays, equipment failures, or network congestion, data acquisition time points are often irregularly distributed, leading to poor performance of traditional interpolation methods. First, an appropriate spline node sequence is determined, and the node density is set according to the time interval distribution of the observed data, increasing the number of nodes in data-dense regions. Then, a cubic B-spline basis function is constructed. The basis function is a cubic polynomial within four consecutive node intervals and zero in other intervals, exhibiting C² continuity to ensure the smoothness of the interpolation results. The observations at irregular time points are represented as a linear combination of spline functions: The coefficient c is determined by minimizing the reconstruction error. i At the same time, a smoothing term is introduced to control the smoothness of the function and prevent overfitting of noise.

[0098] In high-dimensional data spaces, a fractional matching algorithm is employed to avoid the curse of dimensionality and numerical instability. This algorithm effectively circumvents the difficulty of calculating the normalization constant in traditional maximum likelihood estimation by minimizing the difference between the true data distribution and the fractional function (gradient of the log probability density function) of the distribution model. The system first calculates the log density gradient of the true data distribution as the true fractional function, and simultaneously calculates the log density gradient of the distribution model as the fractional function of the distribution model. Then, the fractional matching objective function is constructed:

[0099]

[0100] The least squares method is used to calculate the squared error between the two fractional functions, and the result is obtained by integration and summation. Finally, the Adam optimizer is used to optimize the parameters θ using stochastic gradient descent, achieving precise matching of each marginal distribution through multiple iterations.

[0101] The system ultimately outputs quantified resource response characteristics, including response time constant (reflecting resource adjustment speed), adjustment accuracy (reflecting resource control precision), and stability indicators (reflecting resource output stability). These characteristic indicators form a weight matrix. Where T is the number of time steps, N is the number of resource types, and M is the number of feature dimensions, it provides an accurate description of resource capabilities for subsequent reinforcement learning strategy optimization, enabling aggregation decisions to fully consider the actual characteristics and coordination potential of each resource.

[0102] S4: The energy aggregation process is modeled as a Markov decision process, and the twin-delay deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model.

[0103] Step S4 constructs the reinforcement learning-based aggregation policy optimization model. This step formalizes the dynamic aggregation problem of the energy complex into a Markov Decision Process (MDP) framework, and learns the optimal decision policy through advanced reinforcement learning algorithms. The system first defines the quintuple of the MDP, where S is the state space, A is the action space, P is the state transition probability function, R is the reward function, and γ is the discount factor (usually set to 0.95).

[0104] The state space S is constructed to fully consider the multidimensional characteristics and time-varying nature of the energy system, comprising five main dimensions: load demand, renewable energy, energy storage status, power quality, and economic signals. The load demand dimension includes the current total load demand (MW), load forecast uncertainty (standard deviation), and load change rate (MW / min). The renewable energy dimension includes wind power output forecast (MW), photovoltaic power output forecast (MW), forecast confidence interval, and weather condition coding. The energy storage status dimension includes the state of charge (%) of electrical storage, temperature distribution of thermal storage (°C), pressure status of gas storage (bar), and available capacity of each energy storage device. The power quality dimension includes system frequency (Hz), voltage amplitude at key nodes (pu), power factor, and harmonic distortion rate. The economic signals dimension includes real-time electricity price (yuan / MWh), ancillary service prices, and carbon emission factors. The total dimensions of the state vector are 17, comprehensively capturing system operating status information.

[0105] The action space A adopts a hybrid continuous-discrete design. The continuous action component is defined as the participation allocation vector of various resources. ,in This represents the contribution ratio of the i-th type of resource in the aggregation, satisfying the normalization constraint. Discrete action part This indicates the aggregation mode selection: 0 - Economic Optimization Mode, 1 - Rapid Response Mode, 2 - Stable Operation Mode, 3 - Emergency Support Mode. The complete action vector is... It can flexibly express complex resource scheduling strategies.

[0106] The reward function employs a weighted multi-objective approach to comprehensively evaluate the performance of the aggregation strategy across various aspects. The response speed reward assesses the system's timeliness in responding to aggregation demands by comparing the deviation between the actual response time and the target response time, and penalizing overdue responses. The adjustment accuracy reward assesses the degree of matching between the aggregation output and the target value, considering both overall accuracy and individual deviations of each resource. The operating cost reward calculates the economic cost of executing the strategy, including energy cost, equipment wear and tear cost, and start-up and shutdown cost. The system stability reward assesses the strategy's impact on system stability, including frequency stability, voltage stability, and power oscillation suppression. The environmental benefit reward encourages the consumption of renewable energy, penalizes the use of fossil fuels, and promotes low-carbon operation. The total reward function combines these five indicators using a weighted combination. The weight parameters are determined using the analytic hierarchy process: w1=0.25, w2=0.30, w3=0.20, w4=0.15, w5=0.10, reflecting the relative importance of different performance objectives.

[0107] The system employs an improved Twin Delayed Deep Deterministic Policy Gradient (ETD3) algorithm for training and optimization. This algorithm incorporates several optimizations tailored to the characteristics of the energy aggregation problem. The Actor network utilizes a multi-layer structure, with the input layer receiving a 17-dimensional state vector, which is then processed by the Layer... norm The algorithm performs normalization; the hidden layer uses the Swish activation function and introduces residual connections to alleviate the gradient vanishing problem; the output layer is divided into a continuous action part (using Softmax activation to ensure constraints) and a discrete action part (using Gumbel-Softmax to achieve differentiable sampling). The Critic network adopts a dual-path design, where the state path and action path are processed separately and then merged in the fusion layer. The two Critic networks have the same structure but their parameters are initialized independently, improving the robustness of Q-value estimation.

[0108] During training, the system employs a priority-based experience replay mechanism, allocating sampling probabilities based on TD error to improve the utilization of effective samples. The target network update uses an adaptive soft update mechanism to achieve a balance between rapid learning in the early stages of training and stable convergence in the later stages. Policy updates introduce a delay mechanism and noise regularization, updating the Actor network every 3 steps, and adding truncated Gaussian noise to the target action to improve the policy's exploration ability and robustness. The learning rate is dynamically adjusted using a cosine annealing strategy, changing periodically during training to help the model escape local optima. When the average reward change is less than 0.1% over 100 consecutive episodes and the policy loss stabilizes at 10... -4 The following is considered as training convergence, ultimately yielding a high-performance reinforcement learning aggregation policy optimization model.

[0109] S5: The reinforcement learning aggregation strategy optimization model receives state data in real time and outputs the optimal aggregation strategy. Using the optimal aggregation strategy, the distributed control system sends adjustment instructions to each flexible resource and collects actual response data for online learning and updating, thereby achieving adaptive dynamic optimization.

[0110] This step deploys the trained reinforcement learning model to the actual operating environment, enabling real-time optimized scheduling of flexible resources within the energy complex. Initially, the reinforcement learning aggregation strategy optimization model is deployed to the edge computing nodes of the energy management system, establishing stable, low-latency communication connections with each resource controller via high-speed Ethernet. The edge computing architecture reduces data transmission latency and improves system response speed, making it particularly suitable for rapid response scenarios requiring millisecond-level decision-making.

[0111] During operation, real-time status data is collected and processed at 1-second intervals, including equipment operating parameters obtained from the SCADA system, real-time meteorological data from weather stations, short-term load forecast results from the load forecasting system, and real-time price signals from the energy market. This multi-source heterogeneous data, after preprocessing and feature extraction, forms a 17-dimensional state vector which is input into the Actor network. The network quickly calculates and outputs continuous action vectors, representing the optimal participation ratio of various resources and the discrete aggregation mode selection results.

[0112] Before execution, the system verifies the feasibility of the strategy through a constraint checking module. Constraint checking comprises two levels: physical device constraints and system operational constraints. Physical device constraints include the power limitations of the energy storage device (…). ), capacity limit ( ), gradient limit ( And minimum operating / downtime of equipment, etc. System operating constraints include power balance constraints ( ), line power flow constraints ( ), node voltage constraints Constraints include system reserve capacity, etc. If the strategy violates any constraints, the system uses a quadratic programming projection algorithm to project the action vector into the feasible region. The projection formula is... This ensures that the final executed strategy satisfies both physical constraints and is as close as possible to the original optimization strategy.

[0113] The adjusted strategy is sent to each resource controller via standard communication protocols (such as Modbus, IEC 61850, or a custom API). The control command structure contains three key parts: the target output value (such as the power setpoint), the response time requirement, and the priority flag. After receiving the command, the resource controller confirms the feasibility of execution based on its own status and capability boundaries, and returns confirmation information. If a resource cannot execute the command, the strategy is quickly adjusted, and the task is reallocated to other available resources. During execution, the system continuously monitors the actual output and response characteristics of each resource, recording key performance indicators such as response latency, adjustment deviation, energy conversion efficiency, and operating costs.

[0114] To ensure the model can continuously adapt to system changes, an online learning and update mechanism is adopted. At the end of each control cycle, the system calculates the actual reward value based on the actual execution performance and updates the state transition samples (s). t , a t , r t , s t+1 The data is stored in the experience replay buffer. When the buffer accumulates enough new samples (typically 1000), the system initiates the incremental learning process. Online learning uses a smaller learning rate (10). -5The batch size (32) is set to avoid excessive perturbation of the already trained effective knowledge. To prevent catastrophic forgetting, the system employs Elastic Weight Consolidation (EWC) to add a regularization term to the loss function. ,in Importance weights are used to measure parameters. The degree of impact on model performance These are the key parameter values. Importance weights are estimated using the Fisher information matrix to ensure the stability of important parameters.

[0115] The system updates its model parameters hourly, adapting to new operating modes and equipment characteristics while maintaining historically valid knowledge. As the system's operating time increases, model performance continuously improves, significantly enhancing resource coordination efficiency and accuracy. Practical applications demonstrate that, compared to traditional static optimization methods, this adaptive dynamic optimization system can increase renewable energy absorption by 20%-35%, reduce operating costs by 15%-25%, and shorten regulation response time by 40%-60%, providing strong technical support for the efficient, economical, and environmentally friendly operation of energy complexes.

[0116] The energy complex in an industrial park under the jurisdiction of a municipal power supply company is a typical multi-energy coupled system. The park covers an area of ​​approximately 15 square kilometers and includes 120 manufacturing enterprises, 3 large commercial complexes, and 20,000 residential users. The park has various energy facilities, including a 50MW distributed photovoltaic power station, a 20MW / 40MWh lithium battery energy storage station, a 15MW gas turbine combined heat and power unit, thermal energy storage systems, and electric vehicle charging stations. With the increasing penetration of renewable energy and the diversification of energy demand, the park faces challenges such as large peak-valley load differences, difficulties in renewable energy absorption, and fluctuating power quality, urgently requiring the establishment of an intelligent energy aggregation and management system.

[0117] In 2024, the power company launched a reinforcement learning-based flexible dynamic aggregation project for energy complexes, aiming to achieve coordinated and optimized operation of various flexible resources within the park through advanced artificial intelligence technology. The project team deployed a comprehensive data acquisition network within the park, established a centralized energy management platform, and developed an intelligent dispatching system with self-learning capabilities. After six months of system construction and three months of trial operation, the project has achieved significant results in improving renewable energy absorption rates, reducing operating costs, and improving power quality.

[0118] The first phase of the project focused on establishing a comprehensive data acquisition system and a dynamic system model. For data acquisition, the power company installed over 500 smart sensors and monitoring devices within the park, and built a high-speed, reliable communication network via industrial Ethernet. For the 20MW lithium battery system of the energy storage station, a high-frequency sampling rate of 10Hz was used to monitor key parameters such as voltage, current, temperature, and state of charge of each battery cluster in real time, while also recording health status indicators such as charge / discharge cycle count, internal resistance changes, and capacity decay. The monitoring system for the distributed photovoltaic power station collected environmental parameters such as output power, voltage, current, module temperature, and irradiance of each inverter at a frequency of 1Hz.

[0119] Data acquisition for gas turbine combined heat and power (CHP) units covers multi-dimensional parameters such as unit operating status, fuel consumption, power output, heat output, and flue gas temperature. The thermal energy storage system primarily monitors thermodynamic parameters such as temperature distribution in the storage tank, hot water flow rate, and changes in stored heat capacity. The electric vehicle charging station network collects real-time data on the usage status of each charging pile, charging power, and vehicle information. This raw data undergoes outlier detection and standardization via a preprocessing module. Approximately 2% of outlier data points are eliminated using the three Sigma method, and Z-score standardization eliminates dimensional differences between different parameters.

[0120] During the dynamic modeling phase, the project team used neural network differential equations (NRCs) to establish accurate dynamic models for each type of major equipment. Taking a lithium battery energy storage system as an example, the model uniformly represents the complex dynamic processes such as battery state of charge changes, power response characteristics, temperature effects, and aging effects as continuous-time differential equations. By collecting three months of operating data from the energy storage station, including response characteristics under different charge and discharge modes, performance under various ambient temperature conditions, and voltage response characteristics within different state of charge ranges, the trained model achieved a root mean square error (RMS) of less than 1.5% in state of charge prediction, a voltage prediction error of less than 0.03V, and a power response time prediction error of less than 30ms.

[0121] For gas turbine units, the project team paid particular attention to modeling the dynamic characteristics of their start-up and shutdown processes and load regulation processes. By recording the unit's operating data under different seasons and load levels, including detailed parameter changes during cold starts, hot starts, and different load ramp-up processes, the established dynamic model can accurately predict key indicators such as the unit's start-up time, load response speed, and fuel consumption rate. Model validation results show that the power prediction accuracy reaches 2.1% of the rated power, and the thermal efficiency prediction error is controlled within 0.8 percentage points.

[0122] Based on dynamic modeling, the project team developed an intelligent aggregation problem identification system based on a two-layer spiking neural network. The first layer of the system, a static spiking neural network, is responsible for quickly identifying potential aggregation demands within the system. In actual operation, when the park experiences a sharp increase in load during the summer peak electricity consumption period, the system can identify load peak-valley regulation needs within 50 milliseconds. When a fault occurs at a nearby 110kV substation, causing the system frequency to deviate from the rated value by 0.3Hz, the network immediately identifies frequency regulation needs. When the voltage at a key load center within the park drops to 0.92 per unit, the system quickly identifies voltage support needs.

[0123] The static spiking neural network achieved a recognition accuracy of 91.5% in actual operation, basically meeting the needs of rapid preliminary screening. For identified potential aggregation needs, the system automatically activates a second-layer adaptive dynamic spiking neural network for precise classification. This dynamic classifier employs a growth mechanism based on demand, automatically expanding the network structure according to newly emerging operating modes. In the first two months of project operation, the network gradually expanded from an initial 36 neurons to 68 neurons. The newly added neurons were mainly used to represent new aggregation scenarios such as centralized charging of electric vehicles and coordinated shutdown of industrial loads within the park.

[0124] The dynamic classifier further subdivides the four basic aggregated demand categories into 12 specific subcategories. For example, load peak-valley regulation is subdivided into three subcategories: short-term peak shaving, medium-term peak shaving, and long-term energy storage; frequency regulation is subdivided into primary frequency regulation and secondary frequency regulation; voltage support is subdivided into voltage reactive power support and harmonic mitigation, etc. This refined classification provides more accurate guidance for subsequent resource allocation and strategy formulation. In practical applications, when the system identifies short-term peak shaving demand, it will prioritize the use of energy storage resources with fast response times; while for long-term energy storage demand, it will coordinate resources such as thermal energy storage systems and interruptible loads.

[0125] The third phase of the project focused on addressing the issues of time alignment and quantitative assessment of resource characteristics for multi-source heterogeneous data. Due to inconsistent data sampling frequencies among different devices within the park—the power monitoring system uses 50Hz high-frequency sampling, the thermal system uses 1Hz medium-frequency sampling, and the gas system uses only 0.2Hz low-frequency sampling—this time inconsistency posed a challenge to unified analysis. The project team adopted a multi-marginal stochastic flow matching method to construct marginal distribution models for electrical energy storage, thermal energy storage, and gas energy storage resources.

[0126] Taking the 20MW lithium battery energy storage system in the industrial park as an example, the project team constructed the marginal distribution of key parameters such as power output, response time, charge / discharge efficiency, and cycle life degradation coefficient. Analysis of three months of continuous monitoring data revealed that the power output distribution of the energy storage system exhibits a clear bimodal characteristic, corresponding to the two main operating modes of charging and discharging. The response time distribution is mainly concentrated in the range of 200-800 milliseconds, with an average response time of 450 milliseconds. The charge / discharge efficiency exhibits nonlinear characteristics at different power levels, reaching a maximum efficiency of 94.5% at 50% of rated power and approximately 91.2% at rated power.

[0127] For the thermal energy storage system, the project team focused on analyzing the distribution characteristics of parameters such as the temperature gradient, storage capacity, and heat release rate of the thermal storage tank. Monitoring data showed that the temperature gradient of the thermal storage tank was typically in the range of 0.8-3.2℃ / min, with the heat release rate reaching its peak near the phase change material storage region. Analysis of the gas turbine system's response characteristics indicated that its cold start time was approximately 25 minutes, its hot start time was approximately 8 minutes, and its load ramp-up rate was 5% of the rated power per minute.

[0128] Using measured spline techniques, the project team successfully converted these non-uniformly spaced observation data into a continuous time series representation. A fractional matching algorithm was then employed to precisely quantify the response characteristics of various resources, yielding a complete set of characteristic parameters including response time constants, regulation accuracy, and stability indices. These parameters provide an accurate description of resource capabilities for subsequent reinforcement learning model training.

[0129] Based on the aforementioned analysis, the project team modeled the energy aggregation problem in the industrial park as a Markov decision process, constructing a 17-dimensional state space and a hybrid continuous-discrete action space. The state space includes key information such as the park's current load demand, renewable energy output forecasts, the status of each energy storage device, power quality indicators, and real-time electricity prices. In actual operation, these state variables are updated once per second, providing real-time system state awareness for intelligent decision-making.

[0130] The action space design consists of a participation allocation vector for various flexible resources plus an aggregation mode selection. The continuous action portion includes the output allocation ratio of resources such as energy storage systems, combined heat and power units, and interruptible loads. The discrete action portion includes four aggregation modes: economic optimization, rapid response, stable operation, and emergency support. The project team designed a multi-objective reward function that comprehensively considers response speed, regulation accuracy, operating costs, system stability, and environmental benefits, and determined the weight coefficients of each objective using the analytic hierarchy process (AHP).

[0131] The reinforcement learning model was trained using an improved Siamese-delayed deep deterministic policy gradient algorithm. The actor network received 17-dimensional state input and output a resource allocation policy through a three-layer hidden network. The critic network employed a dual-path design to evaluate action value. Model training utilized historical operational data from the park over the past six months, including typical operational scenarios across spring, summer, autumn, and winter, as well as handling cases of various anomalies and emergencies. The training process employed priority-based experience replay and adaptive soft-update mechanisms. After 1500 iterations, the model converged, achieving an average reward of -42.3 on the validation set and an aggregate accuracy of 96.8%.

[0132] The trained reinforcement learning model was officially deployed to the park's energy management system in May 2025, beginning to guide the coordinated operation of various flexible resources in real time. The system receives real-time status data at a 1-second interval, verifies the feasibility of the strategy through the constraint checking module, and then sends adjustment commands to each resource controller. In actual operation, when the system detected a sudden drop of 8MW in photovoltaic output due to cloud cover, the intelligent dispatching system formulated a response strategy within 1.2 seconds: the energy storage system immediately output 5MW of power, the gas turbine increased its output by 2MW, and simultaneously initiated a 1MW interruptible load reduction, successfully maintaining the system's power balance.

[0133] In a typical summer peak electricity demand regulation scenario, when the park's load reached a peak of 52MW at 19:30, exceeding the power supply capacity by 3MW, the system automatically activated a multi-resource coordinated regulation strategy. The energy storage system output its rated power of 20MW for 30 minutes, the combined heat and power (CHP) unit increased its output from 13MW to 15MW, and simultaneously coordinated with some industrial users to stagger their electricity consumption by 2MW. The entire regulation process was smooth and orderly, successfully alleviating the power shortage and avoiding power outages.

[0134] The system's adaptive learning capability was fully demonstrated in actual operation. When a batch of new electric vehicle charging piles were put into operation in the park, the system automatically adapted to the new load mode through an online learning mechanism. In the first two weeks of operation, the system continuously optimized model parameters by collecting actual response data of electric vehicle charging, enabling the new resources to effectively participate in the park's coordinated scheduling. Similarly, when the response characteristics of a gas turbine changed due to equipment aging, the system could also automatically adjust the corresponding control strategy through continuous learning.

[0135] After three months of trial operation, the project has achieved significant results in several aspects. Regarding renewable energy consumption, the utilization rate of photovoltaic power generation in the park has increased from 82.4% to 94.7%, with an average monthly increase in clean energy consumption of approximately 180 MWh. In terms of operating cost control, through intelligent resource scheduling and peak-valley arbitrage, the park saves approximately 150,000 yuan in electricity expenses per month, while also reducing the number of ineffective equipment start-ups and shutdowns and extending equipment lifespan.

[0136] Regarding power quality improvement, after system regulation, the voltage qualification rate in the park increased from 96.8% to 99.2%, the standard deviation of frequency deviation decreased from 0.08Hz to 0.03Hz, and the power factor remained above 0.95. In terms of emergency response capabilities, the system's response time to various disturbances was reduced by an average of 45%, and regulation accuracy improved by more than 30%. Regarding environmental benefits, by optimizing energy allocation and increasing the utilization rate of renewable energy, the park reduced carbon dioxide emissions by approximately 850 tons per month.

[0137] The successful implementation of the project has provided a more stable and reliable power supply for enterprises in the park, improved the business environment, and also accumulated valuable experience in smart grid construction for the power supply company. With the extension of system operation time and the increase in data accumulation, the system performance is expected to be further improved, providing strong technical support for the construction of a new type of power system.

[0138] In step S1, operational data of flexible resources within the energy complex is acquired. Real-time operational parameters of electrical, gas, thermal, and multi-energy coupling equipment are collected via the industrial Ethernet protocol. A continuous-time dynamic system model is then constructed using neural network differential equation process modeling techniques. Specifically, this includes:

[0139] S1.1: Obtain the operational data of each flexibility resource, use the three sigma rule to remove data points that exceed the normal range, and eliminate dimensional differences through Z-score standardization to generate preprocessed operational data;

[0140] In step S1.1, real-time monitoring of various resources within the energy complex is achieved through the industrial Ethernet protocol. This includes collecting real-time parameters such as power, capacity, and charge / discharge efficiency of electrical resources; pressure, flow rate, and storage status of gaseous resources; temperature, heat capacity, and heat transfer coefficient of thermal resources; and the operating mode and conversion efficiency of multi-energy coupling devices. The data sampling frequency is determined based on the dynamic response characteristics of the resources; 10Hz high-frequency sampling is used for electrical energy storage devices, and 1Hz standard sampling is used for thermal energy storage devices. Outlier detection is performed on the collected data, using the 3σ rule to remove data points exceeding the normal range, and Z-score standardization is used to eliminate dimensional differences.

[0141] Specifically, step S1.1 mainly involves acquiring and preprocessing operational data from various flexible resources within the energy complex. A high-performance, low-latency data acquisition network is established using the industrial Ethernet protocol to achieve comprehensive monitoring of electrical, gas, thermal, and multi-energy coupling devices. Different sampling strategies are employed for different types of energy equipment during data acquisition. For electrical resources with fast response times and significant dynamic characteristics, such as battery energy storage systems, supercapacitors, and power electronic equipment, a high-frequency sampling rate of 10Hz is used to ensure the capture of millisecond-level transient response characteristics. For thermal resources with relatively slow responses, such as thermal storage devices and phase change material thermal storage systems, a standard sampling rate of 1Hz is used. For gas resources and multi-energy coupling devices, sampling rates of 0.5-5Hz are set according to their dynamic characteristics. This differentiated sampling strategy ensures the accuracy of key data acquisition while avoiding the storage and processing pressure caused by data redundancy.

[0142] The collected raw data is diverse, including hundreds of parameters such as voltage, current, power, frequency, temperature, pressure, flow rate, energy storage status, and conversion efficiency. These raw data inevitably contain outliers, noise, and missing values, requiring systematic preprocessing. First, the Three Sigma (3σ) rule is applied for outlier detection and removal. Specifically, the mean μ and standard deviation σ are calculated for each data type. Data points deviating from the mean by more than three standard deviations (i.e., falling outside the range [μ-3σ, μ+3σ]) are marked as outliers and removed. This statistical distribution-based outlier detection method assumes the data roughly conforms to a normal distribution and is applicable to most engineering measurement data. For specific parameters known to deviate from a normal distribution, the system employs an improved interquartile range (IQR) method or a density-based local anomaly factor (LOF) algorithm for anomaly detection, improving accuracy and applicability.

[0143] After outlier removal, the remaining data is standardized to eliminate dimensional differences between parameters. Standardization uses the Z-score method to transform the original data x into standardized data z: This process ensures that the transformed data exhibits a standard normal distribution with a mean of 0 and a standard deviation of 1. This allows parameters of different dimensions and orders of magnitude to be compared and calculated on the same numerical scale, providing a unified data foundation for subsequent modeling. For parameters that need to retain their physical meaning, the system employs Min-Max normalization to linearly map the data to the [0,1] interval. Maintain the relative relationship between parameters.

[0144] In addition to outlier handling and standardization, the system also needs to address missing values ​​in the data. For short-term data gaps (less than three times the sampling period), linear interpolation is used to fill in the gaps. For medium-term gaps (3-10 sampling periods), spline interpolation is used to ensure data smoothness. For long-term gaps (more than 10 sampling periods), the system constructs a temporary predictive model based on historical data from the same period or similar operating conditions to estimate and fill the gaps, and labels these filled data, assigning them lower confidence weights in subsequent analyses. Through this series of processes, the original heterogeneous, incomplete, and inconsistent data is transformed into a high-quality, standardized preprocessed dataset, laying a solid foundation for modeling neural ordinary differential equations.

[0145] S1.2: Based on the preprocessed operating data, the energy system state evolution process is represented as a differential equation. A neural network with a residual network architecture is used to approximate the differential equation function, and the gradient is calculated by the adjoint sensitivity method to establish a battery energy storage model, a gas turbine model, and a thermal energy storage model.

[0146] In step S1.2, a continuous-time dynamic system model is constructed based on the theory of constant differential equations, and the state evolution process of the energy system is represented as follows: The differential equation is given, where x(t) is the system state vector, u(t) is the control input vector, and θ is the neural network parameter. A ResNet-based neural network approximates the function f, and the gradient is calculated using the adjoint sensitivity method to achieve efficient parameter optimization. Independent sub-models are established for each type of resource: the battery energy storage model considers dynamic changes in state of charge and capacity decay effects; the gas turbine model includes start-stop characteristics and thermal inertia effects; and the thermal energy storage model integrates heat transfer processes and phase change characteristics.

[0147] Specifically, step S1.2 uses preprocessed operational data to represent the energy system's state evolution process as differential equations, and then approximates these differential equations using a neural network. The core idea of ​​this step is to integrate traditional physical modeling methods with modern deep learning technology, maintaining the physical interpretability of the model while improving the modeling accuracy of complex nonlinear systems. The system first represents the dynamic process of the energy equipment as ordinary differential equations: ,in Represents the system state vector (such as energy storage state of charge, equipment temperature, pressure, etc.), u(t)∈R m Let θ represent the control input vector (such as charging / discharging power, valve opening, etc.), θ represent the set of parameters to be learned, and f represent the state derivative function, which describes how the system state evolves over time.

[0148] Traditional methods typically construct the function f based on simplified physical models or linearized approximations, which struggle to accurately characterize the nonlinear and time-varying properties of complex energy systems. This invention innovatively employs a deep neural network with a ResNet architecture to approximate the function f, fully leveraging the advantages of deep learning in fitting nonlinear functions. A key feature of ResNet is the introduction of skip connections, enabling the network to learn state changes rather than absolute state values, which aligns closely with the approach of describing state derivatives in differential equations. The neural network structure uses a multilayer perceptron design, typically configured as follows: input layer (n+m neurons, receiving state and control inputs) → first hidden layer (256 neurons) → second hidden layer (128 neurons) → third hidden layer (64 neurons) → output layer (n neurons, outputting state derivatives). The hidden layers use the Swish activation function f(x) = x·sigmoid(βx), where the β parameter is learnable. This function combines the sparse activation properties of ReLU with the smoothness of the hyperbolic tangent function, making it suitable for differential equation approximation tasks.

[0149] The training of neural network parameters θ employs the adjoint sensitivity method to calculate gradients. This method, derived from optimal control theory, efficiently calculates the sensitivity of differential equation systems to parameters. Specifically, the loss function is constructed as follows:

[0150] This measures the deviation between the model's predicted trajectory and the measured data. Then, the adjoint equation is solved.

[0151]

[0152] Calculate the gradient of the loss function with respect to the parameters. , where a is the accompanying variable. This method is better suited for handling differential equation systems than traditional backpropagation, with high computational efficiency and low memory footprint, and can handle long-term sequences and high-dimensional state spaces.

[0153] Based on this framework, dynamic models for three typical energy devices are established. The battery energy storage model considers the dynamic changes in state of charge (SOC), the variation of internal resistance with temperature and cycle number, the nonlinear characteristics of charge-discharge efficiency, and the capacity decay effect. The core differential equation of the model is: Where η is the efficiency factor (related to SOC, temperature T, and power P), For rated capacity, a constant of 3600 is used for unit conversion. The gas turbine model includes start-stop characteristics, operating inertia, output regulation rate limits, and the characteristics of thermoelectric conversion efficiency varying with load. The core equations include the power dynamic equation. and temperature dynamic equation ,in The power time constant, The heat capacity is given. The thermal energy storage model integrates the dynamic characteristics of the heat transfer process, the heat loss mechanism, and the nonlinear characteristics of phase change materials. The core equation is: For phase change materials, special attention needs to be paid to their nonlinear behavior near the phase change temperature. Becomes a function of temperature It exhibits a sharp peak near the phase transition point.

[0154] Each device model is trained using a dedicated neural network, while incorporating physics knowledge to guide the network's learning. For example, physical constraints are added to the loss function to ensure that fundamental physical laws such as energy conservation and the second law of thermodynamics are met. This "physics-informed neural network" method significantly improves the model's generalization ability and physical plausibility, enabling it to maintain reasonable predictions even outside the scope of the training data.

[0155] S1.3: Using a variable step size ordinary differential equation solver, the battery energy storage model, gas turbine model, and thermal energy storage model are trained respectively to obtain the continuous-time dynamic system model.

[0156] In step S1.3, a variable-step-size ordinary differential equation solver is used for model training to ensure numerical stability and computational efficiency. Model performance is evaluated by comparing the mean square error between model predictions and actual measurements, and early stopping is employed to prevent overfitting. The resulting continuous-time dynamic system model accurately describes the dynamic characteristics of various energy resources, providing a foundation for subsequent aggregation problem identification and strategy optimization.

[0157] In step S1.3, a variable-step-size ordinary differential equation solver is used to train and validate the battery energy storage model, gas turbine model, and thermal energy storage model, ultimately integrating them into a complete continuous-time dynamic system model. The training of the neural ordinary differential equation model differs significantly from traditional deep learning models, requiring specialized numerical solution techniques.

[0158] The God Ordinary Differential Equation (NDE) model is a novel modeling paradigm that combines traditional ordinary differential equation theory with modern deep learning techniques. It is specifically designed to handle the modeling and prediction of continuous-time dynamic systems. The core idea of ​​this model is to treat the solution of the differential equation as a continuous limit of a neural network, using the neural network to parameterize the vector field function of the differential equation, thereby achieving accurate modeling of complex dynamic systems. Unlike traditional discrete-time neural networks, the God NEE model can handle predictions at arbitrary time points, without being limited by a fixed time step. This characteristic makes it particularly suitable for handling practical problems in energy systems with irregular time sampling and diverse time scales.

[0159] The fundamental principle of the divine ordinary differential equations (DCEs) is based on the numerical solution theory of initial value problems for ordinary differential equations. Mathematically, any continuous-time dynamic system can be described by a system of first-order ordinary differential equations. The divine DCE approximates the vector field of this system by introducing trainable neural networks. This approximation is not a simple function fitting, but rather learning the intrinsic laws governing the evolution of the system's state while maintaining the continuity of the system's dynamics. The neural network here plays a role similar to empirical formulas or physical laws in traditional modeling, but with greater expressive power, capable of capturing complex nonlinear relationships and higher-order coupling effects.

[0160] In the context of energy system modeling, neural ordinary differential equations (NADEs) are particularly well-suited for describing the dynamic response characteristics of energy devices. Energy devices such as batteries, capacitors, and heat exchangers exhibit significant dynamic characteristics; their state changes follow specific physical laws but are also influenced by a variety of complex factors. Traditional modeling methods often require numerous simplifying assumptions, while neural NDEs can learn these complex relationships through a data-driven approach while preserving physical meaning. For example, the state of charge (SOC) change of a battery depends not only on the charging and discharging current but also on factors such as temperature, aging degree, and cycle count. These complex relationships are difficult to accurately describe using analytical formulas, but neural NDEs can build accurate dynamic models by learning patterns from historical data.

[0161] The training process for neural network ordinary differential equations (NDEs) differs significantly from that of traditional neural networks. Traditional neural network training is based on discrete input-output pairs, while training NDEs requires considering the continuity of the entire time trajectory. Training data is typically in time series form, containing state observations of the system at different times. The training objective is to make the model's predicted state trajectory as close as possible to the observed trajectory. This requires solving the differential equations to obtain the predicted trajectory and then calculating the difference between the prediction and the observation. Because solving differential equations is involved, gradient calculation becomes complex, necessitating the use of the previously mentioned adjoint sensitivity method for efficient gradient computation.

[0162] Numerical solution of the model is a key technical step in the application of neural network differential equations. Since the vector fields parameterized by neural networks are usually highly nonlinear, direct analytical solutions are nearly impossible, necessitating numerical integration methods. Commonly used numerical solvers include the Euler method, the Runge-Kutta method, and multistep methods, each with its own applicable scope and accuracy characteristics. In practical applications, the choice of solver requires a trade-off between computational accuracy and computational efficiency. For real-time applications such as energy systems, computational efficiency is particularly important; therefore, adaptive step-size solvers are often used, dynamically adjusting the integration step size based on the smoothness of the solution.

[0163] Another important characteristic of the neural network's frequent differential equation (NDE) model is its memory efficiency. Traditional recurrent neural networks need to store all intermediate states when processing long sequences, and the memory requirement increases linearly with the sequence length. However, when the neural network calculates gradients using the adjoint sensitivity method, it only needs to store the final state and the adjoint state, making the memory requirement independent of the sequence length. This allows the model to handle very long time series data. This characteristic has significant practical value for long-term series such as energy system monitoring data.

[0164] In applications involving energy complexes, the God constant differential equation model also exhibits excellent generalization ability. Because the model learns the system's intrinsic dynamics rather than simple input-output mappings, a well-trained model can be extended beyond the training timeframe to predict system evolution over longer periods. This time extrapolation capability is of great significance for long-term planning and predictive control of energy systems. Furthermore, the model's continuous-time characteristics enable it to naturally handle data with different sampling frequencies, providing an effective approach for the fusion and processing of multi-source heterogeneous data.

[0165] An advanced variable-step solver is employed to automatically adjust the integration step size based on the severity of state changes, ensuring numerical stability while improving computational efficiency. For rigid differential equation systems (such as thermodynamic systems with multi-timescale characteristics), implicit solvers (such as the backward Euler method or the backward differential formula method) are used; for non-rigid systems, more computationally efficient explicit solvers (such as the Runge-Kutta method or the Adams method) are employed.

[0166] The training process employs a batch training strategy, with each batch containing 10-20 time-series samples, each sample consisting of 100-500 time steps. The optimization algorithm uses the Adam optimizer, with an initial learning rate of 0.001 and a learning rate decay mechanism, reducing the learning rate to 0.8 times its original value every 50 epochs. To prevent overfitting, early stopping is employed, stopping training when the loss function on the validation set shows no improvement for 10 consecutive epochs. An L2 regularization term (with a weight decay coefficient of 0.0001) is also introduced to suppress excessive growth of model parameters. In addition to the main state prediction error, the loss function during training also includes a derivative matching term. This enables the model to learn the correct dynamic derivative relationship.

[0167] The training of the battery energy storage model focuses particularly on the dynamic response characteristics under different operating conditions. The training data includes various charge and discharge modes: constant power mode, impulse response mode, random fluctuation mode, and high-frequency fluctuation mode under actual grid frequency regulation conditions. The model needs to accurately capture the nonlinear efficiency characteristics under different charge and discharge powers, performance changes under different temperature conditions, and voltage response characteristics within different SOC ranges. Validation criteria include a root mean square error (RMSE) of less than 2% for SOC prediction, a voltage prediction RMSE of less than 0.05V, and a power response time error of less than 50ms.

[0168] The training of the gas turbine model focuses on the dynamic characteristics during start-up and shutdown processes and load changes. Training data covers start-up and shutdown conditions such as cold start, hot start, normal shutdown, and emergency shutdown, as well as steady-state operation and load ramp-up processes at different load levels (25%, 50%, 75%, and 100%). The model needs to accurately predict start-up time, time to reach rated load, fuel consumption rate, exhaust temperature, and thermal efficiency changes under different loads. Validation criteria include a power prediction RMSE of less than 3% of rated power, a thermal efficiency prediction RMSE of less than 1 percentage point, and a temperature prediction RMSE of less than 5°C.

[0169] The training of thermal energy storage models pays special attention to the dynamic changes in temperature distribution and heat loss characteristics during the heat charging and discharging process. For phase change material thermal energy storage systems, the focus is on capturing nonlinear behavior near the phase change temperature; for stratified hot water energy storage systems, the focus is on modeling the temperature stratification effect and mixing process. Validation criteria include a temperature prediction RMSE of less than 2°C, a heat metering error of less than 5%, and a thermal response time error of less than 2 minutes.

[0170] After each sub-model is trained, a unified continuous-time dynamic system model is constructed using an ensemble learning approach. The ensemble employs a multi-model switching strategy, automatically selecting the most suitable sub-model based on the current system state and operating mode. Simultaneously, a soft switching mechanism ensures a smooth model switching process, avoiding abrupt changes in prediction results. The final dynamic system model undergoes comprehensive testing on a validation dataset to ensure accurate state predictions and dynamic response characteristics under various operating conditions. This model can predict the system's state evolution trajectory over future time periods, providing a reliable simulation environment and predictive basis for subsequent aggregation problem identification and optimized control.

[0171] In step S2, based on the continuous-time dynamic system model, a static spiking neural network is used to initially identify the energy aggregation problem. After identifying potential aggregation demands, an adaptive dynamic spiking neural network classifier is activated to accurately classify specific energy aggregation problems. Specifically, this includes:

[0172] S2.1: Obtain the normalized state data, and use the leak-integral discharge neuron model of the static pulse neural network to perform pulse coding, and identify four types of aggregated needs: load peak-valley regulation, frequency regulation, voltage support and emergency response.

[0173] In step S2.1, the static spiking neural network serves as the first layer recognizer, responsible for rapidly detecting potential aggregation needs in the system. The network employs a LIF (Leaky Integrate-and-Fire) neuron model, which simulates the integral firing characteristics of biological neurons; the membrane potential dynamic equation is... Where τ is the membrane time constant (usually set to 20ms), R is the membrane resistance (valued at 1MΩ), and I is the input current. When the membrane potential v exceeds the threshold voltage V... th When set to -50mV, the neuron generates a pulse and resets the membrane potential to the resting potential V. rest (-70mV). The input layer contains 64 neurons, receiving normalized system state data, specifically including: load forecast deviation (relative error range ±20%), renewable energy output fluctuation (15-minute sliding variance), energy storage state of charge (SOC percentage), grid frequency deviation (±0.5Hz range), node voltage amplitude (per unit value 0.95-1.05), and other key indicators. The hidden layer adopts a fully connected structure, containing 128 LIF neurons. The connection weights are initialized using the Xavier method, and the synaptic delay is set to a random distribution of 1-5ms. The output layer has 4 neurons corresponding to four types of aggregation needs: load peak-valley regulation, frequency regulation, voltage support, and emergency response. A winner-take-all mechanism is used to determine the final recognition result.

[0174] The core of step S2.1 is to replace the traditional artificial neural network with a spiking neural network (SNN), which is closer to the working mechanism of a biological nervous system, and to process the dynamic information of the energy system using a time-encoding mechanism. First, normalized state data from the continuous-time dynamic system model constructed in step S1 is obtained. After preprocessing, this data forms a unified normalized representation within the [0,1] interval, including key indicators such as load forecast deviation (relative error range ±20%, normalized to [0,1]), renewable energy output fluctuation (15-minute sliding variance, normalized to [0,1]), energy storage state of charge (SOC percentage, originally in the [0,1] interval), grid frequency deviation (±0.5Hz range, normalized to [0,1]), and node voltage amplitude (per unit value 0.95-1.05, normalized to [0,1]).

[0175] Static spiking neural networks employ a biologically plausible Leaky Integrate-and-Fire (LIF) neuron model, which processes information by simulating the dynamic membrane potential of biological neurons. The dynamic equation for the membrane potential of an LIF neuron is as follows: Where τ is the membrane time constant (usually set to 20ms), representing the time characteristic of the natural decay of the neuron's membrane potential; R is the membrane resistance (valued at 1MΩ), representing the proportionality coefficient of input current to membrane potential; I is the input current, representing the input signal from the preceding layer of neurons; and v is the membrane potential, representing the neuron's activation state. When the membrane potential v exceeds the threshold voltage V... th When the voltage is set to -50mV, the neuron generates a spike and resets the membrane potential to the resting potential V. rest (-70mV), then enters a refractory period (approximately 2ms), during which it does not respond to any input. This mechanism enables SNNs to naturally process temporal information, making them particularly suitable for identifying dynamic aggregation problems in energy systems.

[0176] The static spiking neural network (SSN) employs a three-layer architecture: an input layer, hidden layers, and an output layer. The input layer contains 64 LIF neurons, each receiving a normalized system state parameter. The input data is converted into a pulse sequence using rate coding; a higher rate code generates a higher frequency of pulses. A typical coding formula is... Where x is the normalized input value, f min and f max These represent the minimum and maximum pulse frequencies (typically set to 10Hz and 100Hz). The hidden layer employs a fully connected structure, containing 128 LIF neurons. The connection weights are initialized using the Xavier method to ensure stable signal propagation within the network. Synaptic transmission characteristics include synaptic delay and synaptic weights. The synaptic delay is set to a random distribution of 1-5ms to simulate signal transmission delays in biological neural systems, enhancing the network's time sensitivity.

[0177] The output layer has four LIF neurons, each corresponding to one of four basic aggregation needs: load peak-valley regulation (for demand-side management), frequency regulation (for grid stability), voltage support (for local grid quality), and emergency response (for sudden events). The network uses a Winner-Take-All (WTA) mechanism to determine the final identification result. Within a simulation time window (typically 100ms), the output neuron with the most pulses represents the category of the primary aggregation need for the current system. In practical applications, if the pulse counts of two or more output neurons are close (difference less than 10%), the system will label multiple possible aggregation needs for further evaluation by a subsequent precise classifier.

[0178] The network training employs a supervised learning method, combining temporal backpropagation and approximate gradient descent. Since the impulse function is inherently nondifferentiable, the system uses a surrogate gradient method with a smoothed sigmoid function. The gradient of an approximate impulse function is used to achieve end-to-end error backpropagation. The training dataset contains 10,000 labeled samples, covering various typical aggregation scenarios, such as peak load periods, periods of significant renewable energy fluctuations, and grid fault recovery periods. The training objective is to maximize the impulse frequency of neurons in the target class while minimizing the impulse frequency of neurons in the non-target class. The final static spiking neural network can quickly identify potential aggregation needs in the system with an accuracy of over 90%, providing preliminary screening results for subsequent accurate classification.

[0179] S2.2: Based on the four types of aggregation requirements, activate the adaptive dynamic spiking neural network classifier, and through the adaptive dynamic spiking neural network classifier, dynamically adjust the network structure according to the novelty of the input data, and automatically add neurons to expand the network capacity;

[0180] Step S2.2: Based on the four basic aggregation requirements identified by the static spiking neural network, an adaptive dynamic spiking neural network classifier is activated for accurate classification. Unlike the static SNN, the dynamic SNN employs a Growing When Required (GWR) mechanism, which dynamically adjusts the network structure according to the novelty of the input data. This characteristic makes it particularly suitable for handling the constantly evolving new aggregation requirements in energy systems. The dynamic SNN receives the output of the static SNN and combines it with more detailed system state data to form a high-dimensional feature vector for accurate classification.

[0181] The core of adaptive dynamic SNNs is the time-dependent growth mechanism, based on competitive learning theory. This mechanism dynamically adds neurons and adjusts connection weights to gradually form a topological mapping of the input space. The network initially contains only a small number of neurons (typically 2-3 per class), gradually expanding as the learning process progresses. Each neuron has an associated prototype vector w, representing its position in the input feature space. When a new feature vector x is input into the network, the system calculates the Euclidean distance d(x,w) between x and all existing neuron prototype vectors, finding the minimum distance. The corresponding neuron is called the Best Matching Unit (BMU).

[0182] When growth is required, it depends on two key indicators: minimum distance d. bmu and network activity at Network activity is defined as a. t = exp(-d bmu This reflects the degree of matching between the current input and the existing knowledge in the network. When d bmu Greater than the preset novelty threshold θ sim (Usually set to 0.8) and network activity a t Less than the activity threshold θ act When the threshold is set to 0.1, the system determines that the current input represents a new pattern and new neurons need to be added to expand the network capacity. This dual-threshold mechanism can both identify truly new patterns and avoid over-responding to noise and outliers.

[0183] The addition of new neurons follows specific rules. First, the system needs to determine the location of the new neuron in the network, i.e., its prototype vector. An effective strategy is to use the prototype vector w of the new neuron... new Set as the current input x and the BMU prototype vector w bmu Linear combination: w new = 0.5x + 0.5w bmu This allows new neurons to better represent input regions that are not currently fully covered. Furthermore, the system needs to identify the Second Best Matching Unit (SBMU) and establish lateral connections between the BMU and SBMU to form the topology of the input space. Each neuron also maintains an age counter that records the number of training samples passed since the last activation, used for subsequent cleaning of invalid neurons.

[0184] In addition to structural adjustments, the network also updates the prototype vectors of existing neurons using a competitive Hebbian learning rule. For BMU, its prototype vectors are updated according to the following rule: ,in The learning rate is typically set initially to 0.1 and gradually reduced to 0.01 during training. For BMU's direct topological neighbors, a similar update rule with a smaller learning rate is applied: ,in This ensures smooth changes in the topology.

[0185] As training progresses, neurons may become inactive for extended periods, consuming computational resources without providing useful information. The system periodically checks the age counters of all neurons. When a neuron's age exceeds a preset threshold (typically 100 training samples) and its topological connections are less than two, it is removed from the network, freeing up resources for new, effective neurons. This dynamic structural adjustment mechanism enables the network to efficiently represent the input space while adapting to changes in data distribution.

[0186] The final adaptive dynamic SNN can not only identify known aggregated demand types, but also detect and learn newly emerging patterns, forming a continuous learning and adaptation to the operating characteristics of the energy complex. Once a new aggregated demand pattern is identified and stably learned, the system assigns it a specific semantic label and provides an explanation to the administrator through a human-computer interaction interface, achieving interpretability and transparency of the artificial intelligence system.

[0187] In step S2.2, based on the four types of aggregation requirements, an adaptive dynamic spiking neural network classifier is activated. This classifier dynamically adjusts the network structure according to the novelty of the input data, automatically adding neurons to expand the network capacity. Specifically, this includes:

[0188] S2.2.1: Obtain the feature vectors of the four types of aggregation requirements, calculate the Euclidean distance between the input data and the existing neuron prototype vectors, and determine the new mode when the minimum distance exceeds the preset novelty threshold, triggering the network structure adjustment mechanism;

[0189] In step S2.2.1, the dynamic classifier employs a Growing When Required (GWR) mechanism, which is based on self-organizing map theory and dynamically adjusts the network structure according to the novelty of the input data. Novelty is evaluated by calculating the Euclidean distance d between the input vector and the best-matching neuron. bmu Proceed, when d bmu >θ sim (Threshold set to 0.8) and network activity a t <θ act When the threshold is set to 0.1, new neurons are automatically added to expand the network capacity.

[0190] In this embodiment, at this stage, feature vectors for four basic aggregated demands output from a static spiking neural network are first obtained. These feature vectors contain rich multidimensional information. For load peak-valley regulation demand, features include the shape of the load forecast curve, the ratio of peak-valley difference, duration estimation, and the distribution of load types eligible for regulation; for frequency regulation demand, features include the magnitude of frequency deviation, rate of change, oscillation characteristics, and system inertia state; for voltage support demand, features include the magnitude of voltage deviation, reactive power distribution, location of sensitive nodes, and fluctuation period; for emergency response demand, features include event type identification, severity assessment, impact range prediction, and necessary response time. After standardization, these features form high-dimensional feature vectors in a unified format, typically with dimensions between 20 and 30.

[0191] The core of novelty detection is to calculate the distance between the input feature vector and the prototype vectors of existing neurons in the network, thus assessing the similarity between the current input and the learned knowledge. The system uses Euclidean distance as the similarity metric, calculated using the following formula: Here, x is the input feature vector, and w is the neuron prototype vector. To improve computational efficiency, especially in high-dimensional (greater than three dimensions) feature spaces, a KD-tree-based nearest neighbor search algorithm is implemented, reducing the search complexity from O(n) to O(log n), where n is the number of neurons in the network. In practical applications, the system also introduces an early stopping strategy: when the distance to the found nearest neighbor is less than the novelty threshold, the search is immediately stopped, further optimizing computational performance.

[0192] Determining whether an input represents a new pattern requires meeting two conditions simultaneously: the minimum distance exceeds a preset novelty threshold, and the network activity is below an activity threshold. The novelty threshold, θsim, is a key parameter that directly affects the network's growth rate and generalization ability. A threshold that is too low will lead to excessive network growth and oversensitivity to noise and mutations; a threshold that is too high may ignore important new patterns. In this system, θsim employs an adaptive setting strategy, with an initial value of 0.8, gradually increasing as the network size grows. The calculation formula is as follows:

[0193]

[0194] Where n is the current number of neurons, and n0 is the baseline value (set to 50). This adaptive strategy allows the network to grow rapidly in the early stages to cover the input space, while adding new neurons more cautiously in the later stages to avoid redundancy and overfitting.

[0195] Network activity (at) is a secondary metric evaluating the degree of match between the current input and the network, defined as at = exp(-dbmu), with a range of (0,1]. The smaller the optimal matching distance (dbmu), the closer the activity is to 1, indicating a stronger expressive ability of the network for the current input. The activity threshold (θact) is typically set to 0.1, meaning that when the network's expressive activity for the input falls below this value, adding new neurons is considered. In practice, the activity threshold is also dynamically adjusted, initially set to a lower value (0.05) to promote rapid network coverage of the input space, and gradually increased to a stable value (0.15) as training progresses, allowing the network to pay more attention to subtle differences.

[0196] When the current input is determined to represent a new pattern, a network structure adjustment mechanism is triggered, preparing to add new neurons at the corresponding positions. This data-driven dynamic structure adjustment enables the network to adaptively represent the distribution characteristics of the input space, providing a flexible learning framework for capturing the constantly evolving new aggregation demands in energy systems. Compared with traditional fixed-structure networks, this dynamic growth mechanism significantly improves the model's expressive power and adaptability, making it particularly suitable for handling energy system data with non-stationary characteristics. Simultaneously, through a dual-threshold decision mechanism and adaptive parameter strategy, the system effectively controls network complexity while maintaining learning capabilities, achieving a good balance between expressive power and computational efficiency.

[0197] S2.2.2: Based on the network structure adjustment mechanism, create new hidden layer neurons in the corresponding output category, use the feature vector of the current input data as the prototype vector of the new neuron, and initialize the connection weights between the new neuron and the input and output layers.

[0198] In step S2.2.2, the initial weights of the new neuron are determined using a linear interpolation method: w new = 0.5(w bmu +w smu ), where w bmu and w smu These are the weight vectors for the best and second-best matched neurons, respectively.

[0199] Once it is determined that a new neuron needs to be added, the first step is to determine which output category this new neuron should be associated with. In the adaptive dynamic spiking neural network of this invention, each hidden layer neuron is associated with a specific output category, forming a category-specific representation structure. There are two strategies for determining the category association: supervised and semi-supervised methods. During supervised training, the system directly uses the label information of the training samples to associate the new neuron with the corresponding output category; during deployment or in the case of unlabeled samples, the system adopts a semi-supervised strategy, inferring the most likely category based on the category association of the best-matching neuron and the activation mode of the current network.

[0200] After determining the category association, the system creates new hidden layer neurons in the corresponding output category. This process includes three key steps: prototype vector initialization, connection weight setting, and topology update. Prototype vector initialization employs a hybrid strategy, combining the feature vector of the current input data with the prototype vector of the best-matching neuron to form the initial representation of the new neuron. The specific formula is wnew = λx + (1-λ)wbmu, where λ is the hybrid factor, initially set to 0.5 and gradually increased as the network develops, ensuring that later-added neurons focus more on the features of new samples. This hybrid strategy preserves the general features of the current category while also expressing the uniqueness of new samples, ensuring smooth network expansion.

[0201] Setting connection weights is a crucial step in establishing the functional role of new neurons in the network. New neurons need to establish two types of connections: forward connections with the input layer and lateral connections with other hidden layer neurons. Forward connection weights are initialized through supervised learning, employing a hybrid strategy similar to the prototype vector, combining the optimal weights for the current sample (calculated via temporary backpropagation) and the weights of the best-matching neuron to ensure the new neuron can generate appropriate activation for similar inputs. Lateral connections are established through a competitive learning mechanism: new neurons establish excitatory connections (positive weights) with spatially neighboring neurons of the same class and inhibitory connections (negative weights) with neighboring neurons of different classes, forming a local competition-global cooperation activation pattern.

[0202] Topology updating is a core characteristic of self-organizing networks. By dynamically adjusting the connections between neurons, a topology-preserving mapping of the input space is formed. After adding a new neuron, the system needs to update the network's topology, mainly in two aspects: first, establishing a direct connection between the new neuron and the best-matching neuron, setting the initial age to 0; second, disconnecting the best-matching neuron from its farthest neighbors and connecting these neurons to the new neuron. This connection reorganization mechanism ensures that the network structure accurately reflects the topological relationships of the input space, contributing to the formation of a semantically coherent representation space.

[0203] In addition to the basic mechanisms mentioned above, this invention introduces several innovative designs to enhance the stability and adaptability of the network. First, there is a trial period mechanism for new neurons. Newly added neurons enter a trial period of 30 training samples, during which their prototype vectors can be rapidly adjusted (learning rate increased by 50%), but they do not participate in further reorganization of the network structure. This mechanism allows new neurons to fully adapt to their representation regions while avoiding network structure fluctuations caused by immature neurons. Second, there is a neuron integration mechanism. The system periodically checks the similarity of adjacent neurons. When the prototype vectors of two neurons are highly similar (Euclidean distance less than 0.2) and belong to the same category, they are merged into one neuron, reducing redundant representations. Finally, there is an inactive neuron elimination mechanism. Neurons that have not been activated for a long time (more than 200 training samples) are removed, freeing up computational resources.

[0204] Through this dynamic equilibrium structural adjustment mechanism, adaptive dynamic spiking neural networks can efficiently represent and adapt to the complex and ever-changing aggregated demand patterns in energy systems, continuously improving representational power and classification accuracy while maintaining computational efficiency. This self-organizing learning method is particularly suitable for handling emerging patterns and concept drift problems in energy systems, enabling the system to continuously adapt to the ever-changing energy environment and user needs.

[0205] S2.2.3: Update the activation intensity of the new neurons in the network through a competitive learning mechanism, determine the optimal response neuron, and finally determine the capacity of the expanded network.

[0206] Step S2.2.3 updates the activation intensity of newly added neurons in the network through a competitive learning mechanism, determines the optimal responding neuron, and thus effectively expands the network capacity. Competitive learning is an unsupervised learning method inspired by biological neural systems. Its core idea is that neurons compete with each other to respond to specific input patterns. Only the "winning" neuron and its neighboring neurons can update their weights, thereby forming a self-organizing mapping of the input space. In adaptive dynamic spiking neural networks, the competitive learning mechanism plays a crucial role in ensuring that the network can efficiently represent the input space and optimize the allocation of computational resources.

[0207] The competitive learning process begins with the calculation of neuron activation intensity. When the input feature vector x enters the network, the system calculates the activation intensity of each hidden layer neuron using the following basic formula: ,in Let be the prototype vector of neuron i, and σ be the width parameter of the Gaussian function (affecting the size of the neuron's receptive field). To enhance the network's dynamic representation capability, this invention employs a context-sensitive activation mechanism, modifying the activation intensity as follows:

[0208]

[0209] in These are context modulation coefficients (dynamically adjusted based on the neuron's historical activation patterns). This is a temporal modulation factor (considering the temporal structure of the input sequence). This enhanced activation computation enables the network to better capture temporal dependencies and contextual information in energy systems, improving the accuracy and stability of classification.

[0210] Identifying the optimal responding neuron is the core of competitive learning. In the traditional winner-take-all strategy, only the neuron with the highest activation intensity is allowed to respond and update its weights. However, this extreme competitive mechanism easily leads to uneven resource utilization and over-specialization. This invention employs an improved k-winners-take-all strategy, allowing the top k most activated neurons to participate in response and learning, where k is dynamically adjusted according to the network size, typically set to 5%-10% of the total number of hidden layer neurons. The learning intensity of each winning neuron is weighted according to its activation intensity ranking, ensuring that the dominant neuron receives the maximum learning opportunity while allowing the next best neurons to learn moderately. This soft competitive mechanism significantly improves the network's expressive efficiency and generalization ability.

[0211] After the winning neuron is determined, the system updates its prototype vector and connection weights using a competitive Hebbian learning rule. The basic formula is used for updating the prototype vector. ,in A learning rate specific to the neuron. Unlike traditional fixed learning rates, this invention implements a multi-level adaptive learning rate mechanism. At the global level, the base learning rate... The learning rate is gradually decreased as training progresses, starting with an initial value of 0.1 and decreasing to 90% of its original value every 1000 training samples processed, until it reaches a minimum of 0.01. At the neuron level, each neuron adjusts its individual learning rate based on its age and historical update frequency. ,in This is an age-modulated function (new neurons have a higher learning rate). The frequency modulation function is used (the learning rate of frequently updated neurons is appropriately reduced to prevent over-specialization). This multi-level adaptive learning mechanism ensures that the network can balance the need to quickly learn new patterns with the need to retain learned knowledge, significantly improving learning efficiency and stability.

[0212] Connection weight updates employ a similar competitive rule, but with a greater emphasis on topology optimization. For connections between the winning neuron and its direct topological neighbors, the system resets their age to 0, indicating that these connections have been "refreshed" by recent activity. For connections that are not active, their age is increased by 1. When a connection's age exceeds a preset threshold (typically 50-100), the connection is removed, indicating that the corresponding neuron no longer represents an adjacent input region. Simultaneously, the system dynamically creates new connections; when two non-adjacent neurons simultaneously produce strong responses to an input, a new connection is established between them, reflecting the topology of the input space. This dynamic connection management mechanism enables the network to continuously optimize its topology, accurately reflecting the intrinsic relationships within the input space.

[0213] Through the aforementioned competitive learning mechanism, the Adaptive Dynamic Spiking Neural Network (DSN) can efficiently utilize computational resources to form an optimized representation of the input space. The network capacity dynamically expands during the learning process, capable of representing both frequently occurring common patterns and capturing rare but important special cases. As training progresses, the network gradually forms a stable representation structure, and the aggregation requirements of different categories form clear boundaries in the feature space, providing a solid foundation for subsequent accurate classification. This self-organizing mapping capability makes the system particularly suitable for handling complex pattern recognition problems in energy complexes, automatically discovering and adapting to latent structures in the data without requiring extensive manual feature engineering or prior knowledge.

[0214] S2.3: Update the connection weights of the dynamic spiking neural network classifier by adaptive peak timing-dependent plasticity rules, and accurately classify specific energy aggregation problems by the dynamic spiking neural network classifier.

[0215] In step S2.3, the network connection weights are updated using the Adaptive Spike-Timing-Dependent Plasticity (Ad-STDP) rule, which adjusts the synaptic strength based on the time difference Δt between preceding and following pulses. The update formula is as follows: (Activity), among which The learning rate (initial value 0.01, decaying with training) It is a time-dependent function, expressed in double exponential form. or time constant

[0216] f(activity) is an adaptive function, defined as follows: α is the adaptive coefficient (value 0.1). The Ad-STDP rule can adaptively adjust the learning intensity based on the historical activity level of neurons, avoiding overlearning and catastrophic forgetting by maintaining historical activity records (sliding window length 1000 time steps). The classifier output layer contains 12 neurons, corresponding to subdivided aggregate sub-problem types: short-term peak shaving (<15 minutes), medium-term peak shaving (15 minutes - 4 hours), long-term energy storage (>4 hours), primary frequency regulation, secondary frequency regulation, voltage and reactive power support, harmonic mitigation, islanded operation support, emergency load reduction, multi-energy coordination optimization, demand response management, and ancillary service provision, providing precise guidance for subsequent strategy formulation.

[0217] Specifically, step S2.3 updates the connection weights of the dynamic spiking neural network using the Adaptive Spike-Timing-Dependent Plasticity (Ad-STDP) rule to achieve accurate classification of specific aggregation problems. STDP is a learning rule inspired by biological neural systems that adjusts synaptic strength based on the relative timing of impulses from preceding and following neurons, and is an important learning mechanism discovered in neuroscience. In traditional STDP, if the impulse from the preceding neuron causes the impulse from the following neuron to occur (i.e., the preceding impulse occurs before the following impulse), the synaptic connection is strengthened; conversely, if the impulse from the following neuron occurs before the impulse from the preceding neuron, the connection is weakened.

[0218] This invention employs an improved adaptive STDP rule, which not only considers the temporal relationship of impulses but also incorporates historical information about neuronal activity, achieving a more flexible learning process. Specifically, the synaptic weight update formula is as follows:

[0219] (Activity), among which The learning rate is set initially to 0.01 and gradually decreased during training. It is a time-dependent function describing the pulse time difference. Regarding the impact of weight changes, f(activity) is an adaptive function that adjusts the learning intensity based on the historical activity level of neurons.

[0220] Time-dependent functions Using the classic double-exponential form: when When >0 (the preceding pulse precedes the following pulse), ;when When <0 (the later pulse precedes the earlier pulse), Among them, A + and A - The strengths of long-term reinforcement (LTP) and long-term inhibition (LTD) are controlled separately, typically set to 0.5 and 0.55, respectively, to make the learning slightly biased towards inhibition and enhance the selectivity of the network; This is a time constant that controls the width of the time window. It is usually set to 20ms, which is consistent with the typical value of biological nervous systems.

[0221] The adaptive function f (activity degree) is a key innovation of this invention, defined as follows: Where α is the adaptive coefficient (with a value of 0.1), The historical activity level of neurons is calculated using an exponential moving average: β is a smoothing factor (value 0.95). This mechanism reduces the learning rate of high-frequency active neurons and increases the learning rate of low-frequency active neurons, preventing a few neurons from dominating network behavior, promoting more balanced feature expression, and preventing catastrophic forgetting problems.

[0222] To further improve classification accuracy, a supervised learning layer is added to the dynamic SNN. This layer contains 12 output neurons, corresponding to more specific aggregate sub-problem types: short-term peak shaving (<15 minutes), medium-term peak shaving (15 minutes - 4 hours), long-term energy storage (>4 hours), primary frequency regulation, secondary frequency regulation, voltage and reactive power support, harmonic mitigation, islanded operation support, emergency load reduction, multi-energy coordination optimization, demand response management, and ancillary service provision. Supervised learning employs a teacher forcing strategy, providing additional stimulation to the target category neurons during the training phase to accelerate the formation of correct mappings.

[0223] Training data is processed using an online learning approach, with the system continuously collecting new samples from the real-world operating environment and periodically updating network parameters. To evaluate classification performance, the system employs confusion matrix analysis to calculate precision, recall, and F1 score for each aggregation requirement. A target precision of at least 95% is set to ensure that subsequent strategy optimization is based on accurate problem classification. For boundary cases that are difficult to classify, the system uses a soft classification method, outputting multiple possible categories and their probability distributions for the decision-making module to reference.

[0224] The final accurate classification results not only include type labels for aggregation needs but also key attribute information, such as response urgency (categorized as immediate, short-term, and routine response), expected duration (from minutes to days), and resource preference index (indicating the applicability of different resource types). This rich semantic information provides precise guidance for subsequent resource response characteristic evaluation and strategy optimization, ensuring the selection of the most suitable resource combination and control strategy for specific aggregation needs.

[0225] Through this precise classification, the system can transform raw system state data into aggregated demand descriptions with clear semantics, establishing a bridge between the physical world of energy and the decision-making and control system, laying the foundation for subsequent optimization decisions. As system operating experience accumulates, the classifier's performance will continue to improve, gradually forming a deep understanding of the operating modes of the energy complex, realizing the transformation from data to knowledge.

[0226] In step S3, the multilateral stochastic flow matching method is used to perform alignment analysis on energy data measured at non-equidistant time points. Through measured value spline augmentation and fractional matching, the response characteristics of each flexibility resource are quantitatively evaluated, specifically including:

[0227] S3.1: Obtain the data distribution characteristics of different types of energy equipment, construct multiple marginal distributions of electric energy storage resources, thermal energy storage resources and gas energy storage resources, perform nonparametric modeling through kernel density estimation method, and generate a smoothed distribution model;

[0228] In step S3.1, the construction and modeling of multiple marginal distributions first involves constructing multiple marginal distributions, each corresponding to the data distribution characteristics of a specific resource type or time period. Specifically, for electrical energy storage resources (such as lithium batteries and supercapacitors), the marginal distribution includes key dimensions such as power output (kW), response time (ms level), charge / discharge efficiency (85%-95%), and cycle life decay coefficient; for thermal energy storage resources (such as molten salt thermal storage and phase change materials), the marginal distribution covers parameters such as temperature change gradient (°C / min), thermal storage capacity (MWh), heat release rate (MW), and heat loss coefficient; for gas energy storage resources (such as compressed air energy storage), the marginal distribution includes pressure change rate (bar / min), volumetric flow rate (… Indicators such as compression efficiency, etc., are used. Each marginal distribution is modeled nonparametrically using kernel density estimation, avoiding prior assumptions about the data distribution. A Gaussian kernel function is employed.

[0229] The original data points are smoothed, where h is the bandwidth parameter, which directly affects the smoothness and accuracy of the estimation. The bandwidth parameter is determined through k-fold cross-validation (k=5), with the goal of minimizing the mean squared error. This avoids both overfitting and underfitting.

[0230] Specifically, this step first acquires the data distribution characteristics of different types of energy equipment, constructs multiple marginal distributions, and performs nonparametric modeling using the kernel density estimation method to generate a smoothed distribution model. The multi-marginal distribution framework is a key technology for processing high-dimensional heterogeneous energy data. It allows the system to model the characteristics of different energy dimensions separately, and then combine them through a complex dependency structure, effectively avoiding the curse of dimensionality problem faced by directly modeling high-dimensional joint distributions.

[0231] For energy storage resources, the system constructed marginal distributions of several key characteristics. The power output distribution describes the output capability of energy storage devices under different load conditions, typically exhibiting a bimodal distribution reflecting the two main operating modes: charging and discharging. The response time distribution characterizes the time characteristics of the device from receiving a command to reaching the target output; energy storage devices from different technologies show significant differences: supercapacitors and flywheel energy storage have response times concentrated in the millisecond range (5-50 ms), lithium-ion batteries in the second range (0.5-5 s), and lead-acid and sodium-sulfur batteries in the tens of seconds range (10-60 s). The charge / discharge efficiency distribution reflects the energy conversion process losses, which are not only related to the technology type but also closely related to the operating state: efficiency is usually lower under high-power conditions, and higher under shallow charge / discharge cycles. The cycle life degradation coefficient distribution describes the capacity degradation of energy storage devices over the usage period, showing obvious technology correlation and usage mode dependence.

[0232] For thermal energy storage resources, the system constructs the marginal distributions of key parameters such as temperature gradient, thermal storage capacity, heat release rate, and heat loss coefficient. The temperature gradient distribution typically exhibits skewed characteristics, with higher density in the lower gradient region (0.5-2℃ / min), reflecting the thermal inertia of the thermal system. The thermal storage capacity distribution often displays a multimodal structure, corresponding to equipment clusters of different scales and technology types. The heat release rate distribution is highly correlated with equipment type and operating mode: hydroelectric thermal storage systems typically exhibit a uniform distribution, while phase change material thermal storage systems show a temperature-dependent non-uniform distribution, with the heat release rate reaching its peak near the phase change temperature. The heat loss coefficient distribution reflects the energy retention capacity under different insulation technologies and environmental conditions, typically exhibiting a log-normal distribution.

[0233] The marginal distribution of compressed air energy storage resources includes key parameters such as pressure change rate, volumetric flow rate, and compression efficiency. The pressure change rate distribution reflects the dynamic characteristics of the gas storage system's charging and discharging, typically exhibiting a conditional distribution, meaning its shape depends on the current pressure level and the amount of gas stored. The volumetric flow rate distribution is directly related to the system's power output capacity; large compressed air energy storage systems typically... Operating within a certain range. The compression efficiency distribution reflects the energy loss during the energy conversion process and is closely related to the compression ratio, temperature, and equipment characteristics, with a typical distribution range of 65%-85%.

[0234] After acquiring these raw distribution data, the system employs kernel density estimation (KDE) for nonparametric modeling. Compared to parametric distributions (such as the normal distribution and Weibull distribution), KDE makes no prior assumptions about the data distribution form, allowing it to more flexibly adapt to the complex characteristics of various energy devices. The basic idea of ​​KDE is to treat each data point as the center of a kernel function, and to form a smooth probability density estimate by superimposing these kernel functions.

[0235] This system uses a Gaussian kernel function. This function possesses favorable mathematical properties and computational efficiency. The bandwidth parameter h directly affects the smoothness of the estimation and is the most critical hyperparameter in KDE. Too small a bandwidth leads to an overly "rugged" estimation, overfitting noise; too large a bandwidth results in excessive smoothing, masking the true structure of the data. This invention employs an improved Silverman rule to adaptively determine the initial bandwidth. σ is the sample standard deviation, IQR is the interquartile range, and n is the sample size. Based on this, the bandwidth parameter is further optimized using k-fold cross-validation (k=5), with the goal of minimizing the log-likelihood difference between the estimated distribution and the actual data.

[0236] To handle sparse regions and extreme values ​​in marginal distributions, the system introduces an adaptive kernel bandwidth technique. Smaller bandwidth is used in data-dense regions to preserve detailed structure, while larger bandwidth is used in sparse regions to improve estimation stability. This adaptive bandwidth strategy significantly improves the estimation quality in marginal regions of the distribution, which is particularly important for capturing the extreme performance and anomalous behavior of energy devices.

[0237] Using the methods described above, the system ultimately obtains a set of high-quality smooth marginal distribution models that accurately characterize the distribution of various energy devices across different dimensions. These distribution models not only provide a statistical description of device performance but also reflect the randomness and variability of device response characteristics, providing a solid statistical foundation for subsequent precise matching and response characteristic evaluation.

[0238] S3.2: Based on the smoothed distribution model, a cubic B-spline basis function is constructed using the measured value spline technique to process the observations at irregular time points and generate a continuous time series;

[0239] Step S3.2, based on the smoothed distribution model, employs measurement splines to construct cubic B-spline basis functions, processing observations at irregular time points to generate continuous time series. This step addresses the common problem of non-equidistant sampling in energy system monitoring, achieving a reliable conversion from discrete observation points to continuous time functions and providing a unified time reference framework for subsequent analysis. Measurement splines are an advanced interpolation technique specifically designed for irregularly sampled data, balancing data fitting accuracy and function smoothness, and are particularly suitable for processing multi-source heterogeneous time series data in energy systems.

[0240] Initially, the system needs to obtain observations at irregular time points from the smoothed distribution model and determine appropriate spline node sequences. Monitoring data from energy complexes typically come from different systems and equipment from different manufacturers, resulting in inconsistent sampling frequencies and timestamps: power monitoring systems may use high-frequency sampling of 50Hz or 100Hz, thermal systems typically use medium-frequency sampling of 1Hz or 0.2Hz, while gas systems may use only 0.1Hz or lower sampling rates. This temporal inconsistency makes traditional uniform grid interpolation methods difficult to apply, necessitating specialized irregular node processing techniques.

[0241] The determination of the node sequence directly affects the quality and computational efficiency of spline interpolation. This invention employs an adaptive node distribution strategy, setting the node density based on the time interval distribution of the observed data, increasing the number of nodes in dense data regions and decreasing the number of nodes in sparse data regions. Specifically, the system first calculates the time interval distribution of adjacent observation points, obtaining their percentiles (typically 10%, 25%, 50%, 75%, and 90%). Then, based on these statistical characteristics, a node distribution strategy is designed: in regions where the time interval is less than the 25th percentile, the node density is set to one node per original observation point; in regions where the time interval is between the 25th and 75th percentiles, the node density is one node for every two observation points; in regions where the time interval is greater than the 75th percentile, the density is dynamically adjusted according to the interval size, typically not exceeding three times the original observation point interval. This adaptive node strategy ensures the representation accuracy in dense data regions while avoiding overfitting and wasted computational resources in sparse regions.

[0242] After the node sequence is determined, the system constructs cubic B-spline basis functions, which are the core component of spline interpolation. A B-spline is a set of compactly supported polynomial functions; each B-spline basis function is non-zero in a finite interval and zero in other intervals. This local support property gives B-splines significant advantages in computational efficiency and numerical stability. Cubic B-spline basis functions are represented as cubic polynomials over four consecutive node intervals. These basis functions have… Continuity, meaning that the function values ​​and their first and second derivatives are continuous at the nodes, ensures the high smoothness of the interpolation results, making it particularly suitable for representing the continuous change of physical quantities in energy systems.

[0243] For a given sequence of nodes The number of cubic B-spline basis functions is Each basis function spans four adjacent nodes and is always zero outside this interval. The calculation of the basis functions adopts a recursive definition. First, a 0th-order B-spline (i.e., a piecewise constant function) is defined. Then, higher-order B-splines are calculated using recursive formulas. This recursive definition gives the B-spline basis functions good numerical properties and avoids Runge's pherawenon phenomenon commonly found in higher-order polynomial interpolation.

[0244] After the basis functions are constructed, the system determines the coefficients of the spline function using the least squares fitting method, so that the spline curve passes through or approximates all observation points. Let the observed data be (tj, yj), j=1,2,...,n, where tj is the time point and yj is the corresponding observation value.

[0245] To avoid overfitting and control the smoothness of the curve, a regularization term is introduced, forming a penalized least squares problem. This system uses Generalized Cross-Validation (GCV) to automatically select the optimal λ value. , where Sλ is the smoothing matrix, and tr(Sλ) is its trace, representing the "equivalent degrees of freedom".

[0246] By employing the aforementioned measurement spline technique, the system successfully transforms discrete observations at irregular time points into continuous-time function representations, resolving the key challenge of temporal inconsistencies in energy data. The generated continuous time series exhibits excellent interpolation accuracy and smoothness, accurately reflecting the dynamic response processes of various energy devices. This provides a unified time reference framework and a high-quality data foundation for subsequent score matching and characteristic evaluation. This continuous representation is particularly advantageous for capturing dynamic characteristics at different time scales, accurately describing everything from millisecond-level transient electrical responses to hourly-level thermodynamic system changes.

[0247] In step S3.2, based on the smoothed distribution model, a cubic B-spline basis function is constructed using measured value spline technology to process the observations at irregular time points, generating a continuous time series. Specifically, this includes:

[0248] S3.2.1: Obtain the irregular time point observations in the smoothed distribution model, determine the spline node sequence, set the node density according to the time interval distribution of the observation data, and increase the number of nodes in the data-dense region;

[0249] Step S3.2.1 details the feature extraction and analysis process for non-uniformly spaced time series data. In the actual operating environment of energy complexes, data acquisition equipment is often affected by factors such as communication delays, equipment failures, or network congestion, resulting in observation sequences with irregular time intervals. This non-uniform interval characteristic poses a challenge to traditional time series analysis methods, as most standard methods assume that data points are uniformly distributed over time. To address this issue, the system first performs comprehensive feature extraction on the collected raw time series, capturing the basic statistical characteristics, temporal structure features, and frequency domain characteristics of the data.

[0250] Extracting basic statistical features is the first step in understanding data distribution. The system calculates central tendency indicators (mean, median, mode), dispersion indicators (standard deviation, interquartile range, range), distribution shape indicators (skewness, kurtosis), and extreme value characteristics (maximum, minimum, and their occurrence times) for each time series. Unlike equally spaced series, the calculation of these statistics needs to consider the non-uniformity of the time intervals. For example, the formula for calculating the weighted mean is... The weight The time weighting is inversely proportional to the time interval between adjacent observations, ensuring that long time intervals do not excessively affect the overall statistical properties. This time-weighted method allows the statistics to more accurately reflect the actual distribution of physical quantities, unaffected by sampling patterns.

[0251] Temporal structure feature extraction focuses on the dynamic evolution patterns of the data. The system employs Lomb-Scargle periodogram analysis to identify periodic components in non-uniformly spaced data. This method does not require equally spaced data sampling and is particularly suitable for processing irregular observation data in energy systems. The Lomb-Scargle method is used to calculate the normalized power spectrum. Using this method, the system can identify the main periodic components in the data and obtain reliable results even when the sampling intervals are irregular.

[0252] In addition to periodic analysis, the system also detects trend components and change point characteristics of time series. Trend detection employs the Mann-Kendall nonparametric test, which does not require equally spaced data distribution but only focuses on the relative magnitudes of data points, calculating the statistic S. Frequency domain feature extraction transforms the time series into the frequency domain for analysis, revealing the contributions of different frequency components. Since traditional Fast Fourier Transform (FFT) requires equally spaced sampling, the system uses Non-Uniform FFT (NUFFT) technology to process equally spaced data. NUFFT maps equally spaced data onto a uniform grid using interpolation and convolution methods, then applies the standard FFT algorithm, and finally performs appropriate compensation corrections. This method retains the computational efficiency of FFT while adapting to the characteristics of equally spaced sampling. From the spectrum, the system extracts features such as main frequency components, frequency band energy distribution, and power spectral density shape. These features are valuable for identifying the oscillation characteristics, control response characteristics, and noise characteristics of energy equipment.

[0253] For different types of energy equipment, the system also extracts domain-specific features. For electrical energy storage equipment, features such as charge / discharge state transition frequency, maximum power change rate, and response delay distribution are extracted; for thermal energy storage equipment, features such as thermal response curve shape, temperature gradient distribution, and thermal inertia characteristics are extracted; and for gas energy storage equipment, features such as pressure fluctuation characteristics, flow rate change patterns, and compression response curves are extracted. These domain-specific features fully consider the physical characteristics and technical constraints of different energy forms, providing a rich feature base for subsequent resource response characteristic assessment.

[0254] Through the comprehensive feature extraction process described above, the system transforms non-equidistant time series data into structured feature vectors. These feature vectors capture various aspects of the energy equipment's response characteristics, providing rich and accurate information input for subsequent measurement spline processing and fraction matching algorithms, effectively overcoming the analytical challenges brought about by the temporal irregularity of the original data.

[0255] S3.2.2: Based on the node sequence, construct cubic B-spline basis functions, where each B-spline basis function is a cubic polynomial in four consecutive node intervals and zero in other intervals;

[0256] The detailed explanation of step S3.2.2 is as follows:

[0257] The quality of spline interpolation is highly dependent on the choice of node sequence. Inappropriate node distribution can lead to overfitting (too few nodes) or overfitting (too many nodes). This step employs a multi-level adaptive strategy to intelligently determine the optimal node sequence based on the temporal distribution characteristics, value range variation characteristics, and physical model constraints of the data, ensuring an optimal balance between computational efficiency and fitting accuracy.

[0258] The system first performs a preliminary node distribution analysis based on the characteristics of the original non-uniformly spaced time series. The core idea is to increase node density in regions of rapid data change and decrease node density in regions of relatively stable data, thereby achieving the optimal fitting effect with the minimum number of nodes. Specifically, the system calculates the time interval between adjacent data points. Change in sum Then calculate the rate of change. Based on these fundamental quantities, the system constructs an initial node distribution strategy: for regions with a rate of change exceeding the 90th percentile (rapidly changing areas), one node is set for each original data point; for regions with a rate of change between the 50th and 90th percentiles (moderately changing areas), one node is set for every 2-3 data points; for regions with a rate of change below the 50th percentile (stable areas), one node is set for every 5-10 data points. This data-driven adaptive strategy ensures a high degree of matching between the node distribution and the characteristics of data change.

[0259] In addition to the basic strategy based on the rate of change, the system also considers the distribution characteristics of time intervals. For regions with particularly large time intervals (greater than the 95th percentile), which may represent data acquisition interruptions or equipment offline periods, the system adds special processing logic: if the data changes significantly on both sides of the large interval (exceeding 20% ​​of the data range), transition nodes are added on both sides of the interval; if the changes are not significant, a small number of nodes are evenly inserted into the large interval region to avoid "information vacuum zones." This special handling for abnormal time intervals significantly improves the system's adaptability to data interruption situations.

[0260] For energy devices with clear physical model constraints, the system introduces a model-guided node optimization strategy. For example, for battery energy storage systems, the state of charge (SOC) change typically follows a specific charge-discharge curve, with the slope changing significantly as the SOC approaches 0% or 100%. Utilizing this prior knowledge, the system automatically increases node density in the critical SOC regions (0-10% and 90-100%), ensuring the physical plausibility of the fitted curve even if the original data is insufficiently sampled in these regions. Similarly, for thermal energy storage systems, node density is increased near the phase transition temperature (within ±2°C) to accurately capture the nonlinear thermal characteristics during the phase transition process; for compressed air energy storage systems, node density is increased in regions where the pressure approaches the maximum rated value to reflect the nonlinear compression characteristics under high pressure.

[0261] After the initial node sequence is determined, the system further applies an adaptive node optimization algorithm to optimize the node distribution through an iterative process. The core idea is to identify areas with insufficient node distribution through local fitting error analysis and then add or adjust node positions accordingly. Specifically, the system first performs spline fitting using the initial node sequence, and then calculates the fitting error for each original data point. , where s(t) is a spline function. Based on the error distribution, error hotspots (errors greater than twice the average error) are identified, and nodes are added in these regions. The number of added nodes is proportional to the magnitude of the local error, but an upper limit is set (typically 5 new nodes) to prevent overfitting. The precise location of the new nodes is determined by minimizing the local error function, ensuring that the fitting error is minimized to the greatest extent possible.

[0262] To prevent the number of nodes from growing indefinitely, the system also introduces a node redundancy detection and merging mechanism. When the spline function between two adjacent nodes is almost linear (the second derivative is close to zero) and the fitting error is small, removing one of the nodes can be considered without significantly affecting the fitting quality. The system calculates a node importance index. ,in The contribution of each node is evaluated using a distance-weighted function. Nodes with importance below a threshold are marked as redundant and can be removed while maintaining overall fitting accuracy. This dynamic balancing mechanism ensures that the number of nodes remains within a reasonable range, avoiding both underfitting and overfitting, as well as the waste of computational resources.

[0263] Through the aforementioned multi-level adaptive strategy, the system ultimately obtained a highly optimized sequence of spline nodes. These nodes are distributed in a manner highly consistent with the temporal structure, value range variations, and physical characteristics of the data, laying a solid foundation for the subsequent construction of cubic B-spline basis functions. Compared to fixed-interval or simple heuristic methods, this intelligent node determination strategy significantly improves the accuracy and computational efficiency of spline interpolation, making it particularly suitable for processing complex and variable non-uniformly spaced time series data in energy systems.

[0264] S3.2.3: The coefficients of the cubic B-spline basis function are determined by the least squares fitting method, so that the spline curve passes through or approximates all observation points, thus obtaining the continuous time series with time continuity and smoothness.

[0265] Step S3.2.3 completes the construction of cubic B-spline basis functions based on optimized node sequences, and applies the regularized least squares method to determine the spline function coefficients, ultimately generating a high-quality continuous time series representation. B-splines are a class of piecewise polynomial functions with local support. Compared with traditional interpolation polynomials, they have advantages such as high computational efficiency, good numerical stability, and strong flexibility, making them particularly suitable for processing large-scale and complex energy system data.

[0266] The system first uses the optimized node sequence determined in step S3.2.2. Construct a set of cubic B-spline basis functions. A cubic B-spline is a B-spline of order k=3, where each segment is a cubic polynomial with C at the nodes. 2 Continuity (i.e., the function values ​​and their first and second derivatives are continuous). This smoothness makes cubic B-splines particularly suitable for representing continuous state changes in physical systems, accurately capturing the dynamic response characteristics of energy devices while avoiding unnatural oscillations or jumps.

[0267] S3.3: The fractional matching algorithm is used to calculate the difference between the minimum true data distribution of the continuous time series and the fractional function of the distribution model to obtain the response characteristics of each flexibility resource.

[0268] Step S3.3 employs a score matching algorithm to calculate the difference between the true data distribution of the continuous time series and the score function of the distribution model, quantitatively evaluating the response characteristics of each flexibility resource. Score matching is an advanced distribution estimation technique, particularly suitable for density estimation problems in high-dimensional data spaces. By comparing the score functions (gradients of the log probability density function) of the distributions rather than directly comparing probability densities, it cleverly avoids the difficulty of calculating normalization constants in traditional methods, greatly improving computational efficiency and numerical stability.

[0269] First, the system acquires real data samples of the continuous time series generated in the previous step. These samples have been processed by measurement splines and transformed into a continuous representation under a unified time reference. Based on these samples, the system calculates the logarithmic density gradient of the real data distribution pdata(x). , as the true fractional function. Specifically, for each data point x, its fractional function is calculated using the gradient of the kernel density estimate: Where K is the kernel function, Its derivative. This nonparametric estimation method avoids prior assumptions about the form of data distribution and can flexibly adapt to the complex characteristic distribution of various energy devices.

[0270] At the same time, the system establishes a parameterized model distribution. And calculate its logarithmic density gradient. The distribution model's fractional function is used as the basis for its selection. The choice of model distribution must consider the characteristics of the problem and computational complexity. For simpler unimodal distributions, the system uses exponential family distributions (such as normal, gamma, and beta distributions) and their mixtures; for complex multimodal distributions or distributions with special constraints, the system uses more flexible representations such as normalizing flows or energy-based models. Regardless of the model chosen, its fractional function... Probability density and its normalization constant can usually be calculated directly using analytical derivatives or automatic differentiation techniques, avoiding the difficulty of directly calculating them.

[0271] Based on the true fractional function and the fractional function derived from the distribution model, the system constructs a fractional matching objective function, calculates the squared error between the two fractional functions using the least squares method, and integrates and sums the results. Mathematically, this is represented as follows: , where Epdata represents the expectation of the true data distribution. This objective function can be further simplified by integrating parts, avoiding direct estimation. ,in This represents the trace of the Hessian matrix of the logarithmic density function, which is the sum of the second derivatives of each dimension. This transformation allows score matching to only require calculating the derivative of the model distribution, without needing to estimate the derivative of the true distribution, further improving computational efficiency and stability.

[0272] The optimization of the score matching objective function employs a gradient descent-type algorithm, with the Adam optimizer widely used due to its ability to adaptively adjust the learning rate for different parameters. The optimization parameters are set as follows: initial learning rate of 0.001, momentum parameters β1=0.9, β2=0.999, and a gradient clipping threshold of 10.0 to prevent gradient explosion. To prevent overfitting, the system uses early stopping, stopping training when the loss function on the validation set shows no improvement for 30 consecutive epochs. The optimization process typically requires 500-1000 iterations, with the convergence criterion being a relative change in the loss function of less than 10 for 10 consecutive epochs. -6 .

[0273] Through score matching optimization, the system obtains the optimal model parameters θ*, making the score function of the distribution model approximate the true score function to the greatest extent. Based on these optimized parameters, the system quantitatively evaluates the response characteristics of each flexibility resource, mainly including three types of key indicators:

[0274] 1. Response Time Constant: Characterizes the time characteristic of a resource from receiving an instruction to reaching the target output. For electrical energy storage devices, the response time constant is typically in the range of milliseconds to seconds, reflecting its rapid adjustment capability; for thermal energy storage devices, the response time constant is in the range of minutes to hours, reflecting the thermal inertia characteristics of the thermal system; for gas energy storage devices, the response time constant is in the range of seconds to minutes, reflecting the dynamic characteristics of the compression and release process.

[0275] 2. Regulation accuracy: Characterizes the degree of matching between the actual resource output and the target command. Regulation accuracy is usually expressed as relative error or root mean square error, reflecting the accuracy and stability of the control system. High-precision regulation (error <1%) is suitable for fine-grained power grid frequency regulation services, medium-precision regulation (error 1%-5%) is suitable for conventional load following, and low-precision regulation (error >5%) is only suitable for extensive peak shaving and valley filling.

[0276] 3. Stability Indicators: These indicators characterize the volatility and sustainability of resource output. Stability indicators include multiple dimensions such as output variance, maximum duration, and decay rate, comprehensively reflecting the reliability and durability of resources across different time scales. High-stability resources are suitable for providing long-term basic support services, medium-stability resources are suitable for short- to medium-term adjustment services, and low-stability resources are only suitable for short-term impulse response services.

[0277] These response characteristic indicators constitute a capability profile of each flexible resource, forming a response characteristic matrix that accurately describes the randomness of state transitions and the sequentiality of decisions for each resource during energy aggregation. These indicators not only reflect the physical characteristics and technical limitations of the resources but also encompass the combined impact of factors such as the control system's reaction speed, communication latency, and measurement accuracy. This provides a comprehensive and accurate description of resource characteristics for subsequent reinforcement learning strategy optimization, ensuring that optimization decisions fully consider the actual capabilities and coordination potential of each resource, thereby achieving optimal overall system performance.

[0278] S3.3.1: Obtain the true data distribution of the continuous time series, calculate the log density gradient of the true data distribution and use the log density gradient of the true data distribution as the true fractional function, and at the same time calculate the log density gradient of the distribution model and use the log density gradient of the distribution model as the fractional function of the distribution model;

[0279] In step S3.3.1, the fractional matching algorithm effectively solves the curse of dimensionality and numerical instability problems in high-dimensional data spaces by cleverly avoiding the calculation of normalization constants. The core idea of ​​this algorithm is to minimize the difference between the true data distribution and the fractional function (gradient of the log probability density function) of the distribution model.

[0280] In the first crucial step of the score matching algorithm, the system needs to acquire real data samples of continuous time series and calculate the log density gradient of the real data distribution based on these samples. This process begins by extracting representative sample points from the high-quality continuous time series generated in the previous steps. These sample points not only include the original observation data but also estimates of intermediate time points generated through spline interpolation. The sample selection strategy considers both the uniformity of the time distribution and the representativeness of data variations, ensuring coverage of both typical system operating states and various anomalies and transitional states.

[0281] The calculation of the true fractional function employs a nonparametric method based on kernel density estimation. The system first constructs a Gaussian kernel function for each data point, and then estimates the log density gradient of the overall distribution by calculating the weighted gradients of all kernel functions. The advantage of this method is that it does not require any prior assumptions about the form of the data distribution and can automatically adapt to various complex distribution patterns. In actual calculations, the system employs an adaptive bandwidth selection strategy, dynamically adjusting the width parameter of the kernel function according to the local density characteristics of the data to ensure accurate gradient estimation in densely populated regions and maintain the stability of the estimation in sparsely populated regions.

[0282] Meanwhile, the system establishes a parameterized model distribution to approximate the real data distribution. The choice of model distribution is based on the physical characteristics of the energy system and the statistical features of the data. For simple distributions exhibiting unimodal characteristics, the system uses appropriate members from the exponential family of distributions, such as the normal distribution for symmetric distributions, the gamma distribution for right-skewed distributions, and the beta distribution for bounded distributions. For distributions exhibiting multimodal characteristics or complex shapes, the system uses a hybrid model or a more flexible energy model. The logarithmic density gradient of the model distribution can usually be directly calculated using analytical methods or automatic differentiation techniques, which greatly simplifies the subsequent optimization process.

[0283] When processing high-dimensional data, the computation of fractional functions faces the challenge of the curse of dimensionality. The system employs various dimensionality reduction and simplification strategies to address this problem. First, principal component analysis or other dimensionality reduction techniques are used to identify the main directions of data variation, projecting the high-dimensional problem into a low-dimensional subspace for processing. Second, a decomposition strategy is employed, decomposing the high-dimensional joint distribution into a combination of multiple low-dimensional marginal and conditional distributions. Fractional function estimations are performed on each of these combinations before they are synthesized. This decomposition not only reduces computational complexity but also improves the stability and interpretability of the estimation.

[0284] S3.3.2: Based on the true fractional function and the fractional function of the distribution model, construct a fractional matching objective function, use the least squares method to calculate the squared error between the true fractional function and the fractional function of the distribution model, and integrate and sum the calculation results;

[0285] In step S3.3.2, the objective function is defined as follows: ,in For the true data distribution, For parameterized model distribution, This represents a fractional function. In actual calculations, the true fractional function is obtained through empirical estimation: ,in These are sample points.

[0286] Based on the true score function and the score function of the distribution model calculated in the previous step, the system enters the stage of constructing the score matching objective function. The core idea of ​​this objective function is to minimize the squared error between the two score functions, thereby making the gradient characteristics of the model distribution as close as possible to the gradient characteristics of the true data distribution. Compared with the traditional maximum likelihood estimation method, the score matching method cleverly avoids the computational challenge of the normalization constant of the probability density function, which has a significant computational advantage when dealing with complex high-dimensional distributions.

[0287] The objective function is constructed using a squared loss in the form of expectation, which calculates the expected square of the difference between two fractional functions under the true data distribution. In actual calculations, this expectation is approximated by averaging a finite number of samples. To improve the accuracy of the estimation, the system employs advanced sampling techniques such as importance sampling and stratified sampling. Importance sampling reduces the estimation variance by adjusting sample weights, and is particularly effective when dealing with tail distributions. Stratified sampling ensures uniform coverage of the samples throughout the data space, avoiding insufficient samples in certain important regions.

[0288] In the specific implementation of the objective function, the system also considers the balance between computational efficiency and numerical stability. For large-scale datasets, directly calculating the score differences of all sample pairs is computationally infeasible. The system employs a mini-batch stochastic gradient method, calculating the objective function value of only a subset of samples at a time, and approximating the global optimum through multiple iterations. The mini-batch size needs to be carefully chosen; too small a batch size will lead to excessive noise in gradient estimation, while too large a batch size will reduce computational and memory usage efficiency.

[0289] To enhance the stability of the optimization process, the system introduces regularization terms into the objective function. These regularization terms include L2 norm penalties for the parameters to prevent overfitting; gradient smoothness constraints to ensure the learned fractional function has good mathematical properties; and penalties based on physical constraints to ensure the model distribution is physically reasonable. The regularization parameters are automatically determined through methods such as cross-validation, achieving an optimal balance between fitting accuracy and generalization ability.

[0290] The objective function also needs to address the scale differences between different dimensions and physical quantities. Different parameters in energy systems often have different orders of magnitude and units, and directly calculating errors may lead to certain dimensions dominating the optimization process. The system employs an adaptive weight adjustment strategy, dynamically adjusting the weight coefficients based on the variance and importance of each dimension. Simultaneously, a multi-scale objective function is introduced, defining sub-objectives at different time and spatial scales, and finding the global optimum through multi-objective optimization techniques.

[0291] S3.3.3: Optimize the model parameters of the score matching objective function using the gradient descent algorithm, so that the score function of the distribution model approximates the true score function, and obtain the response characteristics of each flexibility resource including response time constant, adjustment accuracy and stability index.

[0292] In step S3.3.3, the Adam optimizer is used to optimize the parameters θ using stochastic gradient descent, with a learning rate of 0.001 and momentum parameters β1=0.9 and β2=0.999. Through multiple iterations (typically 500-1000 rounds), precise matching of each marginal distribution is achieved. The convergence criterion is that the change in the loss function is less than 10 for 10 consecutive rounds. -6 .

[0293] Specifically, after the objective function is constructed, the system enters the iterative optimization phase of the model parameters. This phase employs gradient-based optimization algorithms, with the Adam optimizer being the preferred choice due to its ability to adaptively adjust the learning rate for different parameters. The Adam optimizer combines the advantages of momentum and RMSprop algorithms, automatically adjusting the learning rate for each parameter during training, making it particularly effective for problems with sparse or noisy gradients. In complex optimization problems such as energy systems, the gradient information for different parameters often varies significantly; Adam's adaptive nature ensures that all parameters are updated appropriately.

[0294] The hyperparameter settings for the optimization process need to be tuned according to the specific characteristics of the problem. The initial learning rate is set to 0.001, a widely validated default value that neither causes the optimization process to diverge nor slows down the convergence speed. Momentum parameter... Setting it to 0.9 controls the influence of historical gradient information; higher values ​​help accelerate convergence and reduce oscillations. The decay rate of the second-order moment estimation. The value is set to 0.999 to estimate the variance of the gradient; a higher value helps maintain stable parameter updates later in training. To prevent gradient explosion, the system sets a gradient clipping threshold of 10.0; gradients exceeding this threshold are scaled proportionally.

[0295] The convergence criteria during training employ multiple standards. First, the trend of the loss function is considered; when the relative change in the loss function over ten consecutive iterations is less than one ten-thousandth, it is considered close to convergence. Second, the magnitude of parameter changes is assessed; when the update magnitude of all parameters is less than one-thousandth of their current value, the parameters are considered stable. Finally, the performance on the validation set is monitored to prevent overfitting. An early stopping mechanism is implemented, halting training when the validation set loss shows no improvement for thirty consecutive epochs; this helps select the model with the best generalization performance.

[0296] The optimization process typically requires hundreds to thousands of iterations to reach a satisfactory convergence. During training, the system monitors the changing trends of various metrics in real time, including training loss, validation loss, gradient norm, and parameter norm. By visualizing the curves of these metrics, problems during training, such as vanishing gradients, exploding gradients, and overfitting, can be identified promptly, and corresponding adjustments can be made. The system also records key intermediate results for subsequent analysis and debugging.

[0297] After model training, the system extracts response characteristic indicators for various flexibility resources from the optimized parameters. The extraction of the response time constant is based on the time decay characteristics of the model distribution, determined by analyzing the rate of change of the fractional function over time. For electrical energy storage devices, the response time constant is typically in the range of milliseconds to seconds, reflecting its rapid adjustment capability; for thermal energy storage devices, the response time constant is in the range of minutes to hours, reflecting the thermal inertia characteristics of the thermal system; for gas energy storage devices, the response time constant is in the range of seconds to minutes, reflecting the dynamic characteristics of the compression and release processes.

[0298] The regulation accuracy index is calculated by analyzing the deviation distribution between the model's predicted values ​​and the actual observed values. The system calculates various accuracy indices such as mean absolute error, root mean square error, and maximum error, and determines the corresponding accuracy level according to the requirements of different application scenarios. High-precision regulation capability makes the resources suitable for participating in fine-grained grid ancillary services, medium-precision regulation capability is suitable for routine load following tasks, while low-precision regulation capability is mainly used for extensive peak shaving and valley filling applications.

[0299] The calculation of stability metrics encompasses multiple time scales and performance dimensions. Short-term stability is assessed by analyzing the high-frequency fluctuation characteristics of resource output, including parameters such as output variance, fluctuation frequency, and decay rate. Medium-term stability focuses on the resource's ability to maintain performance during continuous operation, including metrics such as maximum continuous uptime and performance decay rate. Long-term stability evaluates the reliability and lifespan characteristics of resources during long-term use. These stability metrics provide crucial information for formulating resource scheduling strategies and scheduling maintenance plans.

[0300] In step S4, the Siamese-Delay Deep Deterministic Policy Gradient Algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model, specifically including:

[0301] S4.1: Model the energy aggregation process as a Markov decision process, construct a state space including load demand, renewable energy output and real-time resource status, and an action space including resource combination and output allocation ratio. Design a reward function that integrates response speed, adjustment accuracy and operating cost. Based on the Markov decision process framework, state space, action space and reward function, use the twin-delay deep deterministic policy gradient algorithm for training and optimization to generate a reinforcement learning aggregation policy optimization model.

[0302] In step S4.1, based on the Markov decision process framework, state space, action space, and reward function, the Siamese-delay deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregated policy optimization model, specifically including:

[0303] S4.1.1: Obtain multidimensional characteristic data, construct a state space including load demand dimension, renewable energy dimension, energy storage status dimension, power quality dimension and economic signal dimension, and an action space including resource participation allocation vector and aggregation mode selection, and establish a Markov decision process framework.

[0304] In step S4.1.1, the dynamic aggregation problem of the energy complex is formalized as a Markov Decision Process (MDP), defined as a quintuple, where S is the state space, A is the action space, P is the state transition probability, R is the reward function, and γ is the discount factor. The construction of the state space S fully considers the multidimensional characteristics and time-varying nature of the energy system. Specifically, it includes: load demand dimension (total load demand at the current moment, load forecast uncertainty, load change rate), renewable energy dimension (wind power output forecast, photovoltaic power output forecast, forecast confidence interval, weather state coding), energy storage state dimension (electric energy storage state of charge, thermal energy storage temperature distribution, gas energy storage pressure state, available capacity of each energy storage device), power quality dimension (system frequency, voltage amplitude of key nodes, power factor, harmonic distortion rate), and economic signal dimension (real-time electricity price, ancillary service price, carbon emission factor). The action space A adopts a hybrid continuous-discrete design, where the continuous action part is defined as the participation degree allocation vector of various resources, and the discrete action part represents the aggregation mode selection.

[0305] S4.1.2: In the aforementioned Markov decision process framework, design a multi-objective reward function that includes a response speed reward, a regulation accuracy reward, an operating cost reward, a system stability reward, and an environmental benefit reward, and generate a comprehensive performance evaluation index;

[0306] In step S4.1.2, the reward function adopts a weighted multi-objective form to comprehensively evaluate the performance of the aggregation strategy:

[0307] Response speed bonus items:

[0308]

[0309] in Measuring response time deviation Severe penalties will be imposed for exceeding the time limit.

[0310] Adjustment accuracy bonus items:

[0311]

[0312] The first item assesses overall accuracy, while the second item considers individual biases of each resource.

[0313] Operating cost incentives:

[0314]

[0315] It includes three parts: energy cost, wear and tear cost, and start-up and shutdown cost.

[0316] System stability bonus items:

[0317]

[0318] The effects of frequency stability, voltage stability, and power oscillation suppression were evaluated.

[0319] Environmental benefit awards:

[0320]

[0321] Encourage the consumption of renewable energy and punish the use of fossil fuels.

[0322] Total reward function:

[0323]

[0324] The weight parameters are determined using the analytic hierarchy process (AHP): .

[0325] S4.1.3: Based on the comprehensive performance evaluation index, the actor network and the critic network are constructed using the improved twin-delay deep deterministic policy gradient algorithm to obtain the reinforcement learning aggregation policy optimization model.

[0326] In step S4.1.3, the improved twin-delay deep deterministic policy gradient (Empowered TD3, ETD3) algorithm is adopted and optimized for the characteristics of the energy aggregation problem. Actor network structure: The input layer receives a 17-dimensional state vector, which is then processed by Layer... normalization Normalization was performed. The first hidden layer has 256 neurons, and the Swish activation function was used. β=1.702. The second hidden layer has 128 neurons, with residual connections introduced to alleviate gradient vanishing. The third hidden layer has 64 neurons, with a Dropout layer (p=0.1) added to prevent overfitting. The output layer is divided into two parts: continuous action output uses Softmax activation to ensure... Constraints and discrete action outputs are implemented using Gumbel-Softmax for differentiable sampling. The Critic network structure employs a dual-path design, processing the state path and action path separately before merging. State path: 17-dimensional input → 128-dimensional hidden layer → 64-dimensional representation. Action path: n+1-dimensional input → 64-dimensional hidden layer → 64-dimensional representation. Fusion layer: 128-dimensional concatenation → 64-dimensional → 32-dimensional → 1-dimensional Q-value output. The two Critic networks, Q1 and Q2, have identical structures but their parameters are initialized independently.

[0327] The experience playback buffer uses a priority sampling mechanism based on TD error. Assign sampling probability Where α=0.6 controls the priority strength. The buffer capacity is set to 2×10. 6 This ensures sufficient experience diversity. The target network update employs an adaptive soft update mechanism: , where τ base =0.001, τ max =0.01, λ=0.001, achieving a balance between rapid learning in the early stages of training and stable convergence in the later stages. Policy updates introduce a delay mechanism and noise regularization: the Actor network is updated every d=3 steps, and truncated Gaussian noise is added to the target action. , -c, c), where σ=0.1, c=0.2, to improve the robustness of the policy. Loss function optimization: The Critic loss adopts Huber loss. or To reduce the impact of outliers, an entropy regularization term H(π) = -Σπ(a|s)logπ(a|s) is added to the Actor loss to encourage policy exploration. Learning rate scheduling: a cosine annealing strategy is used. ,in , T represents the total number of training steps, enabling periodic adjustment of the learning rate. The learning rate is adjusted when the average reward change over 100 consecutive episodes is less than 0.1% and the policy loss stabilizes at 10. -4The following conditions are considered met: training convergence. The final model's average reward on the validation set is no less than -50, and the aggregate accuracy reaches over 95%.

[0328] In step S5, the reinforcement learning aggregation strategy optimization model receives state data in real time and outputs the optimal aggregation strategy. Using the optimal aggregation strategy, adjustment commands are sent to each flexible resource through the distributed control system, and actual response data is collected for online learning and updating to achieve adaptive dynamic optimization. Specifically, this includes:

[0329] S5.1: Use the reinforcement learning aggregation strategy to optimize the model and obtain real-time state data. Verify the feasibility of the strategy through the constraint checking module. If the strategy violates the constraints, use the projection algorithm to project the action vector into the feasible region to generate the optimal aggregation strategy.

[0330] In step S5.1, the trained reinforcement learning model is deployed to the edge computing nodes of the energy management system and establishes communication connections with each resource controller via high-speed Ethernet. During system operation, the model receives real-time status data at 1-second intervals, including equipment operating parameters collected by the SCADA system, meteorological forecast data, and load forecast results. After preprocessing and feature extraction, the data is input into the Actor network, which outputs continuous action vectors representing the optimal participation ratio of various resources. The feasibility of the strategy is verified through a constraint checking module before action execution.

[0331] After the reinforcement learning model is trained and deployed to the actual operating environment, the system enters the real-time policy generation and execution phase. The first key step in this phase is to use the trained reinforcement learning aggregation policy optimization model to acquire real-time state data and generate preliminary control policies. The system collects the latest state information from various monitoring points at one-second intervals, covering the overall operating status of the energy complex. Real-time data acquisition is achieved through a high-speed Ethernet network, ensuring low latency and high reliability of data transmission.

[0332] Preprocessing of state data is a crucial step in ensuring the quality of model input. Raw monitoring data may contain outliers caused by measurement noise, communication interference, or equipment malfunctions; these issues, if not addressed promptly, can severely impact decision-making quality. The system employs a multi-layered data cleaning and validation mechanism. First, obvious outliers are filtered out through a rationality check. Then, potential data errors are detected and corrected through time series analysis. For short-term data gaps, the system uses interpolation methods to complete the data; for longer-term data interruptions, the system activates backup sensors or uses historical data for estimation.

[0333] Preprocessed state data is input into the Actor network for policy reasoning. Based on the current state, the Actor network quickly calculates the optimal participation ratio for various resources and the aggregation mode selection. This calculation process is typically completed within tens of milliseconds, meeting the latency requirements of real-time control. The network output policy includes a continuous resource allocation vector and discrete mode selection results, forming a complete control command framework.

[0334] Constraint checking of the strategy is a crucial mechanism to ensure the safe and reliable operation of the system. The constraint checking module verifies whether the generated strategy meets various physical and operational constraints. Physical constraints include technical boundary conditions such as power limits, ramp rate limits, minimum operating time, and minimum downtime for various devices. Operational constraints include system power balance requirements, line transmission capacity limits, node voltage range constraints, and system stability requirements. The constraint checking employs a fast numerical calculation method, enabling the verification of all constraints within milliseconds.

[0335] When a policy violation is detected, the system automatically initiates a projection algorithm to correct infeasible action vectors into the feasible region. The projection algorithm employs quadratic programming, making minimal adjustments to the constraint-violating parts while preserving the basic policy intent. This adjustment strategy ensures that the corrected control instructions satisfy all constraints and are as close as possible to the original optimized policy. The projection process also considers constraint priority, giving the highest priority to safety-related hard constraints and allowing moderate relaxation for economy-related soft constraints.

[0336] After constraint checks and projection corrections are completed, the system generates the final aggregation strategy. This strategy includes not only the target output values ​​for each resource but also detailed information such as the execution time window, priority indicators, and backup plans. The strategy is represented using a standardized data format for easy subsequent transmission and execution. Simultaneously, the system generates a confidence assessment of the strategy, evaluating the reliability of its execution based on the uncertainty of the current state and the model's prediction accuracy.

[0337] S5.2: Based on the optimal aggregation strategy, adjustment instructions are sent to each flexible resource through the communication protocol and actual execution effect data is collected to obtain actual response data;

[0338] In step S5.2, the adjusted strategy is sent to each resource controller via the Modbus protocol to achieve distributed coordinated control. The control command contains two parts: target output power and response priority. After receiving the command, the resource controller confirms the feasibility of execution based on its own state and capability boundaries, and returns feedback information to the central control system. During execution, the system continuously monitors the actual output and response characteristics of each resource, recording key indicators such as response delay, adjustment deviation, and energy conversion efficiency.

[0339] After the strategy is generated, the system enters the command transmission and execution monitoring phase. Based on the generated optimal aggregation strategy, the system needs to transmit specific control commands to the controllers of various flexible resources. Command transmission employs multiple standardized industrial communication protocols, including Modbus, IEC61850, and DNP3, to ensure reliable communication with equipment from different manufacturers and using different technologies. The choice of communication protocol is optimized based on the characteristics of the equipment type and application scenario. High-speed Ethernet protocols are used for energy storage devices requiring high real-time performance, while lower-cost serial communication protocols can be used for thermal equipment with relatively lower real-time requirements.

[0340] The control commands are designed using a structured information format, with each command containing multiple key components. The target output value is the core of the command, explicitly specifying the power output level or other control objectives the device needs to achieve. The response time requirement specifies the timeframe within which the device must complete the adjustment action; this requirement is dynamically adjusted based on current system demands and device capabilities. Priority indicators help the device controller determine the execution order when faced with multiple commands, ensuring that the most important adjustments are executed first. Backup commands provide alternatives when the primary command cannot be executed, enhancing the system's fault tolerance.

[0341] Upon receiving an adjustment command, the resource controller first performs a local feasibility check. This check considers factors such as the current status, health condition, maintenance status of the equipment, and external environmental conditions. If the controller determines that the command can be safely executed, it sends a confirmation message to the central system and begins the adjustment action. If the controller finds that the command poses a security risk or is technically infeasible, it reports the specific problem information to the central system and provides possible alternatives. This distributed security check mechanism establishes an effective balance between central decision-making and local security.

[0342] During command execution, the system continuously monitors the actual response of each resource. Monitoring includes multiple dimensions such as actual output power, response latency, regulation accuracy, energy conversion efficiency, and equipment status. Monitoring actual output power helps the system understand whether the regulation effect has met expectations; monitoring response latency helps evaluate the dynamic performance of the equipment; monitoring regulation accuracy reflects the accuracy of the control system; and monitoring conversion efficiency relates to the system's economic efficiency.

[0343] The system also establishes a mechanism for detecting and handling abnormal responses. When the actual response of a resource deviates significantly from the expected response, the system automatically initiates anomaly handling procedures. Minor deviations may only trigger fine-tuning of adjustment parameters, while severe deviations may lead to the resource being temporarily excluded from scheduling and other resources being activated to compensate. Anomaly detection employs a statistical learning-based method, which can distinguish between normal performance fluctuations and genuine abnormal conditions, avoiding false alarms and missed alarms.

[0344] To optimize communication efficiency and reduce network load, the system employs an intelligent data transmission strategy. Frequently changing critical parameters are transmitted at high frequencies, while relatively stable status information is transmitted at lower frequencies. The system also implements data compression and incremental transmission techniques, transmitting only the changed data portions, significantly reducing bandwidth consumption. Under poor network conditions, the system can automatically adjust its data transmission strategy to ensure the timely and reliable delivery of the most critical information.

[0345] S5.3: Calculate the actual reward value for the actual response data and store it in the experience replay buffer, start online learning and updating, and realize the adaptive dynamic optimization of the system.

[0346] In step S5.3, the online learning update mechanism ensures that the model continuously adapts to system changes. After each control cycle, actual execution performance data is collected, including actual output power, response time, and adjustment accuracy of each resource. The actual reward value is calculated and stored in the experience replay buffer. When the buffer accumulates sufficient new experience, online learning updates are initiated. Online learning uses a relatively small learning rate (10^6). -5 The batch size (32) is adjusted to avoid excessive perturbation of the trained parameters. To prevent catastrophic forgetting, Elastic Weights Combining (EWC) is used, adding a regularization term to the loss function. ,in As importance weight, These are the key parameter values. The model parameters are updated hourly to ensure that historical knowledge is retained while adapting to new operating modes.

[0347] Based on instruction execution and response monitoring, the system enters the continuous learning and adaptive optimization phase. The core objective of this phase is to enable the reinforcement learning model to continuously improve its decision-making performance based on actual operational experience, adapting to changes in system characteristics and new operating modes. First, the system needs to collect and organize actual response data, which includes not only the output performance of each resource but also comprehensive indicators such as the overall system performance, user satisfaction, and economic benefits.

[0348] The calculation of actual reward values ​​is a crucial step in online learning. Unlike the theoretical reward function used in the training phase, actual reward values ​​are calculated based on real-world performance results, more accurately reflecting the actual effectiveness of the strategy. Response speed rewards are calculated based on measured response time data; adjustment accuracy rewards are calculated based on the deviation between the actual adjustment effect and the target value; operating cost rewards are calculated based on actual expenses incurred; system stability rewards are calculated based on power quality monitoring data; and environmental benefit rewards are calculated based on actual energy consumption and emission data. This reward calculation based on actual performance results makes the learning process more realistic and improves the model's practicality.

[0349] The management of the experience replay buffer employs advanced storage and retrieval strategies. Due to the massive amount of data generated during online operation, the system needs to intelligently select and retain the most valuable experience samples. A priority sampling mechanism allocates storage priority based on the learning value of experiences; those experiences containing rare situations, important decisions, or significant learning effects are given higher retention priority. Simultaneously, the system also implements time-sensitivity management of experiences, gradually eliminating outdated experiences to free up storage space for new ones. This dynamic experience management ensures that the buffer always contains the most relevant and valuable learning samples.

[0350] Online learning parameter update strategies need to strike a balance between learning new knowledge and retaining historical knowledge. The learning rate is set much lower than the initial training phase, typically one-tenth to one-hundredth of the training rate. This setting avoids new experiences overwhelmingly impacting already learned knowledge. The batch size is also adjusted accordingly to a small value, usually between 16 and 64, ensuring that each update is a gradual improvement rather than a drastic change. The update frequency is set to once per hour, maintaining learning sensitivity while avoiding overly frequent parameter adjustments.

[0351] To prevent catastrophic forgetting, the system implements a resilient weight merging technique. This technique calculates the importance of each model parameter to historical task performance, imposing stronger constraints on important parameters when learning new tasks. Specifically, the system maintains an importance weight matrix, recording the sensitivity of each parameter to key performance indicators. During parameter updates, changes in highly important parameters are limited to a smaller range, while less important parameters are allowed larger adjustments. This selective constraint mechanism effectively protects critical knowledge from being destroyed during the new learning process.

[0352] Adaptive optimization is also reflected in the dynamic adjustment of the model structure. When the system detects a new operating mode or resource type, it may need to expand the model's expressive power. The system implements a progressive network structure expansion mechanism, which can add new neurons or connections without affecting existing functionality. The newly added network components quickly adapt to new task requirements through transfer learning, while the original network structure remains relatively stable. This progressive expansion ensures that the system can continuously adapt to the ever-changing application environment.

[0353] The system also establishes a performance monitoring and evaluation mechanism to regularly assess the effectiveness of online learning and the changing trends of model performance. Key performance indicators include multiple dimensions such as decision accuracy, response speed, economic efficiency, and system stability. By comparing performance changes before and after online learning, the system can objectively evaluate the learning effect and adjust the learning strategy or roll back to a previous model version when necessary. This continuous performance monitoring ensures that the system always remains in optimal operating condition.

[0354] like Figure 2 As shown, corresponding to the above method, the present invention also provides a flexible dynamic aggregation system for energy complexes based on reinforcement learning, including: an acquisition module, an identification module, an analysis module, a construction module, and an output module.

[0355] The acquisition module is used to acquire operational data of flexible resources within the energy complex. It collects real-time operational parameters of electrical, gas, thermal, and multi-energy coupling equipment via the industrial Ethernet protocol and constructs a continuous-time dynamic system model using the neural network differential equation process modeling technique.

[0356] The identification module is used to perform preliminary identification of energy aggregation problems based on the continuous-time dynamic system model using a static spiking neural network. After identifying potential aggregation needs, it activates an adaptive dynamic spiking neural network classifier to accurately classify specific energy aggregation problems.

[0357] The analysis module is used to determine the corresponding data processing strategy and time window parameters by utilizing the different aggregation demand types in the accurate classification results of the specific energy aggregation problem, and to perform alignment analysis on energy data measured at non-equidistant time points using the multilateral stochastic flow matching method. Through measurement value spline enhancement technology and fraction matching, the response characteristics of each flexibility resource are quantitatively evaluated. The response characteristics of each flexibility resource include the randomness of state transitions and the sequentiality of decision-making in the energy aggregation process.

[0358] The module is used to model the energy aggregation process as a Markov decision process, and to train and optimize it using the twin-delay deep deterministic policy gradient algorithm to generate a reinforcement learning aggregation policy optimization model.

[0359] The output module is used to receive state data in real time and output the optimal aggregation strategy using the reinforcement learning aggregation strategy optimization model. Using the optimal aggregation strategy, it sends adjustment instructions to each flexible resource through the distributed control system and collects actual response data for online learning and updating to achieve adaptive dynamic optimization.

[0360] This invention organically combines advanced methods such as neural network differential equation modeling, two-layer spiking neural network classification, multi-marginal stochastic flow matching, and reinforcement learning strategy optimization to construct a complete flexible dynamic aggregation method and system for energy complexes. This method can accurately model the dynamic characteristics of complex, multi-element, heterogeneous energy systems, effectively process high-dimensional, non-equidistant time series data, and realize the dynamic aggregation and optimized scheduling of flexible resources within energy complexes, demonstrating broad application prospects.

[0361] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various changes and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of protection of the present invention.

[0362] This disclosure also provides a computer-readable storage medium storing a computer program. When a processor executes the computer program, it performs the steps of the reinforcement learning-based flexible dynamic aggregation system method for energy complexes described in the above-described method embodiments. The storage medium can be volatile or non-volatile computer-readable storage.

[0363] Furthermore, this disclosure also provides a computer program product storing a computer program. When the computer program is run by a processor, it executes the steps of the reinforcement learning-based flexible dynamic aggregation system method for energy complexes provided in any of the above embodiments of this disclosure. For details, please refer to the above method embodiments, which will not be repeated here.

[0364] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0365] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices and apparatuses described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0366] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0367] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0368] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0369] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A reinforcement learning based energy complex dynamic optimization method, characterized in that, The method comprises the following steps: acquiring operation data of flexible resources in an energy complex, collecting real-time operation parameters of electric, gas, heat and multi-energy coupling equipment, and using neural differential equation process modeling technology to build a continuous time dynamic system model; based on the continuous time dynamic system model, using a static pulse neural network to preliminarily identify an energy aggregation problem, and activating an adaptive dynamic pulse neural network classifier after identifying a potential aggregation demand to accurately classify a specific energy aggregation problem, including: acquiring normalized state data, using a leaky integral discharge neuron model of the static pulse neural network to perform pulse coding, identifying four types of aggregation demands of load peak-valley adjustment, frequency adjustment, voltage support and emergency response; based on the four types of aggregation demands, activating an adaptive dynamic pulse neural network classifier, and dynamically adjusting the network structure according to the novelty of the input data, automatically increasing neurons to expand the network capacity through the adaptive spike timing-dependent plasticity rule to update the connection weights of the dynamic pulse neural network classifier, and classifying the specific energy aggregation problem through the dynamic pulse neural network classifier; acquiring different aggregation demand types in the accurate classification result of the specific energy aggregation problem, using a multi-marginal random flow matching method to align and analyze energy data measured at non-equidistant time points, quantitatively evaluating the response characteristics of each flexible resource through measurement value spline enhancement technology and fractional matching, including: acquiring data distribution characteristics of different types of energy equipment, constructing multiple marginal distributions of electric energy storage resources, thermal energy storage resources and gas energy storage resources, performing non-parametric modeling through kernel density estimation method to generate a smoothed distribution model; based on the smoothed distribution model, using measurement value spline technology to construct a cubic B-spline basis function to process observation values at irregular time points, generating a continuous time sequence; using a fractional matching algorithm to calculate the difference between the minimum real data distribution of the continuous time sequence and the fractional function of the distribution model, obtaining the response characteristics of each flexible resource; wherein the response characteristics of each flexible resource include the randomness of state transition and the sequence of decision-making in the energy aggregation process; modeling the energy aggregation process as a Markov decision process and using a twin-delayed deep deterministic policy gradient algorithm for training and optimization to generate a reinforcement learning aggregation strategy optimization model; using the reinforcement learning aggregation strategy optimization model to receive state data in real time and output an optimal aggregation strategy, and using the optimal aggregation strategy to send adjustment instructions to each flexible resource through a distributed control system and collect actual response data for online learning and updating to realize adaptive dynamic optimization.

2. The method of claim 1, wherein, acquiring operation data of flexible resources in an energy complex, collecting real-time operation parameters of electric, gas, heat and multi-energy coupling equipment, and using neural differential equation process modeling technology to build a continuous time dynamic system model, including: Obtaining operation data of the flexible resources, removing data points beyond the normal range by using the three-sigma rule and eliminating dimension differences by Z-score standardization to generate preprocessed operation data; Based on the preprocessed operation data, the state evolution process of the energy system is represented as a differential equation form, a neural network with a residual network architecture is used to approximate the differential equation function and the gradient is calculated by the adjoint sensitivity method to establish a battery energy storage model, a gas turbine model and a thermal energy storage model; A variable step size ordinary differential equation solver is used to train the battery energy storage model, the gas turbine model and the thermal energy storage model respectively to obtain the continuous-time dynamic system model.

3. The method of claim 1, wherein, Based on the four types of aggregated demand, an adaptive dynamic pulse neural network classifier is activated, and the network structure is dynamically adjusted according to the novelty of the input data by the adaptive dynamic pulse neural network classifier, and the network capacity is automatically expanded, including: Obtaining the feature vector of the four types of aggregated demand, calculating the Euclidean distance between the input data and the existing neuron prototype vector, and determining a new mode when the minimum distance exceeds the preset novelty threshold to trigger the network structure adjustment mechanism; Based on the network structure adjustment mechanism, a new hidden layer neuron is created in the corresponding output category, the feature vector of the current input data is taken as the prototype vector of the new neuron, and the connection weights between the new neuron and the input layer and the output layer are initialized; The activation strength of the new neuron in the network is updated through the competitive learning mechanism to determine the optimal response neuron and finally determine the expanded network capacity.

4. The method of claim 1, wherein, The fractional matching algorithm is used to calculate the difference between the minimum real data distribution of the continuous time sequence and the fractional function of the distribution model to obtain the response characteristics of the flexible resources, including: Obtaining the real data distribution of the continuous time sequence, calculating the logarithmic density gradient of the real data distribution and taking the logarithmic density gradient of the real data distribution as the real score function, and calculating the logarithmic density gradient of the distribution model and taking the logarithmic density gradient of the distribution model as the score function of the distribution model; Based on the real score function and the score function of the distribution model, a score matching objective function is constructed, the least squares method is used to calculate the square error between the real score function and the score function of the distribution model, and the calculation result is integrated and summed; The model parameters of the score matching objective function are optimized by the gradient descent algorithm to make the score function of the distribution model approach the real score function to obtain the response characteristics of the flexible resources including response time constant, regulation accuracy and stability index.

5. The method of claim 1, wherein, Based on the smoothed distribution model, the measurement value spline technology is used to construct a cubic B-spline basis function to process irregular time point observations to generate a continuous time sequence, including: Obtaining the irregular time point observations in the smoothed distribution model, determining the spline node sequence, setting the node density according to the time interval distribution of the observation data, and increasing the number of nodes in the data intensive area; Based on the node sequence, the cubic B-spline basis functions are constructed, each of which is a cubic polynomial within four consecutive node intervals and zero in other intervals; The coefficients of the cubic B-spline basis functions are determined by a least square fitting method, so that the spline curve passes through or approximates all observation points, and the continuous time sequence with time continuity and smoothness is obtained.

6. The method of claim 1, wherein, The energy aggregation process is modeled as a Markov decision process, and a twin delayed deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model, including: The energy aggregation process is modeled as a Markov decision process, and a state space and an action space are constructed, and a reward function is designed; Based on the Markov decision process framework, state space, action space and reward function, a twin delayed deep deterministic policy gradient algorithm is used for training and optimization to generate the reinforcement learning aggregation policy optimization model.

7. The method of claim 6, wherein, Based on the Markov decision process framework, state space, action space and reward function, a twin delayed deep deterministic policy gradient algorithm is used for training and optimization to generate a reinforcement learning aggregation policy optimization model, including: Multi-dimensional characteristic data is obtained, a state space including load demand dimension, renewable energy dimension, energy storage state dimension, power quality dimension and economic signal dimension, and an action space including resource participation degree allocation vector and aggregation mode selection are constructed, and a Markov decision process framework is established; In the Markov decision process framework, a multi-objective reward function including response speed reward item, adjustment accuracy reward item, operation cost reward item, system stability reward item and environmental benefit reward item is designed, and a comprehensive performance evaluation index is generated; Based on the comprehensive performance evaluation index, an improved twin delayed deep deterministic policy gradient algorithm is used to construct an actor network and a critic network to obtain the reinforcement learning aggregation policy optimization model.

8. The method of claim 1, wherein, The reinforcement learning aggregation policy optimization model is used to receive real-time state data and output an optimal aggregation policy in real time, and the optimal aggregation policy is used to send adjustment instructions to each flexible resource through a distributed control system, and actual response data is collected for online learning and updating to realize adaptive dynamic optimization, including: The reinforcement learning aggregation policy optimization model is used to obtain real-time state data, and a constraint checking module is used to verify the feasibility of the reinforcement learning aggregation policy. If the reinforcement learning aggregation policy violates the constraints, a projection algorithm is used to project the action vector into the feasible region to generate the optimal aggregation policy; Based on the optimal aggregation policy, adjustment instructions are sent to each flexible resource through a communication protocol, and actual execution effect data is collected to obtain actual response data; The actual reward value of the actual response data is calculated and stored in an experience replay buffer, and online learning and updating are started to realize adaptive dynamic optimization.

9. A reinforcement learning based energy complex dynamic optimization system, characterized in that, including: An acquisition module is used to acquire operation data of flexible resources in an energy complex, real-time operation parameters of electrical, gas, thermal and multi-energy coupling devices are collected through an industrial Ethernet protocol, and a continuous time dynamic system model is constructed using neural ordinary differential equation process modeling technology; The identification module is configured to preliminarily identify an energy aggregation problem by using a static spiking neural network based on the continuous-time dynamic system model, and to activate an adaptive dynamic spiking neural network classifier to accurately classify a specific energy aggregation problem after identifying a potential aggregation demand, including: obtaining normalized state data, performing pulse coding by using a leaky integrate-and-fire neuron model of the static spiking neural network, and identifying four types of aggregation demands, i.e., load peak-valley regulation, frequency regulation, voltage support, and emergency response; based on the four types of aggregation demands, activating the adaptive dynamic spiking neural network classifier, and dynamically adjusting the network structure according to the novelty of the input data, automatically increasing neurons to expand the network capacity by using the adaptive dynamic spiking neural network classifier; updating the connection weights of the dynamic spiking neural network classifier by using an adaptive spike-timing-dependent plasticity rule, and classifying the specific energy aggregation problem by using the dynamic spiking neural network classifier; The analysis module is configured to determine a corresponding data processing strategy and time window parameter by using different aggregation demand types in the accurate classification result of the specific energy aggregation problem, perform alignment analysis on energy data measured at non-equidistant time points by using a multi-marginal stochastic flow matching method, and quantitatively evaluate the response characteristics of each flexible resource by using a measured value spline enhancement technique and a fractional matching, including: obtaining data distribution characteristics of different types of energy equipment, constructing multiple marginal distributions of electrical energy storage resources, thermal energy storage resources, and gas energy storage resources, performing non-parametric modeling by using a kernel density estimation method, and generating a distribution model after smoothing processing; based on the distribution model after smoothing processing, constructing a cubic B-spline basis function to process observation values at irregular time points by using a measured value spline technique, and generating a continuous time sequence; calculating the difference between the minimum real data distribution of the continuous time sequence and a fractional function of the distribution model by using a fractional matching algorithm, and obtaining the response characteristics of each flexible resource; wherein the response characteristics of each flexible resource include the randomness of state transition and the sequence of decision-making in the energy aggregation process; The construction module is configured to model the energy aggregation process as a Markov decision process, and train and optimize the model by using a twin-delayed deep deterministic policy gradient algorithm to generate a reinforcement learning aggregation strategy optimization model; The output module is configured to receive state data in real time and output an optimal aggregation strategy by using the reinforcement learning aggregation strategy optimization model, and to send adjustment instructions to each flexible resource by using the optimal aggregation strategy through a distributed control system, and collect actual response data for online learning and updating, thereby realizing adaptive dynamic optimization.

Citation Information

Patent Citations

  • Scheduling decision model establishment method based on SumTree-TD3 algorithm

    CN117291390A