Group scheduling and group control digital twinning method for pavement distributed photovoltaic energy storage system
Through the group scheduling and group control digital twin method and self-supervised learning energy prediction and scheduling strategy for road distributed photovoltaic energy storage systems, the system's scheduling problem under complex environments and variable load conditions is solved, and efficient energy scheduling and optimization management is achieved.
Patent Information
- Application Number
- CN202510331507.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing distributed photovoltaic energy storage systems are difficult to achieve efficient energy scheduling and optimization management under complex environments and variable load conditions, and lack real-time perception and adaptive scheduling capabilities.
The group-scheduling and group-control digital twin method is adopted for pavement distributed photovoltaic energy storage system, combined with self-supervised learning distributed energy prediction and scheduling strategies, and perceive environmental changes in real time and generate environmental adaptive scheduling strategies through multi-objective optimization algorithms.
It realizes efficient scheduling and optimization management of distributed photovoltaic energy storage systems in dynamic environments, improves energy utilization efficiency, and enhances the real-time and adaptive capabilities of the system.
Smart Images

Figure CN120237691A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital twins, and particularly relates to a group regulation and control digital twin method for a distributed photovoltaic energy storage system for road surfaces. Background Art
[0002] With the rapid development of renewable energy technologies, photovoltaic power generation, as a clean, efficient, and environmentally friendly energy method, has been widely applied globally. Especially with the popularization of distributed photovoltaic energy storage systems, more and more photovoltaic systems are deployed in urban and industrial scenarios, including building rooftops, traffic road surfaces, and public facilities, etc. These systems can not only provide local power support but also provide additional power output through energy storage units during peak energy demand periods, thereby optimizing energy usage efficiency. However, in actual operation, distributed photovoltaic energy storage systems face many challenges. Especially under complex environments and variable load conditions, how to achieve efficient energy scheduling and optimization management remains a difficult point in technological development.
[0003] Currently, most existing scheduling methods for photovoltaic energy storage systems mainly rely on rule-based scheduling mechanisms or make scheduling decisions based on traditional optimization algorithms (such as linear programming, dynamic programming, etc.). These methods usually assume that the operating parameters of the system (such as light intensity, temperature, traffic flow, etc.) are known and unchanged, or use simple empirical formulas for prediction and scheduling, ignoring the adaptability and self-optimization ability of the system in a dynamic environment. This static scheduling method often fails to cope with complex environmental changes, resulting in problems such as the following in actual operation of the system:
[0004] 1) The photovoltaic power generation efficiency is greatly affected by environmental factors (such as weather, temperature, light, etc.), and traditional scheduling methods cannot be adjusted in real time according to environmental changes;
[0005] 2) The charge and discharge scheduling of energy storage units cannot flexibly cope with the fluctuations of the power grid load, resulting in energy waste or instability;
[0006] 3) Most current scheduling algorithms lack sufficient real-time performance and adaptability, cannot be dynamically adjusted according to real-time data and future demands, and are prone to uneven energy distribution or low system efficiency.
[0007] In addition, although some more advanced scheduling methods have introduced machine learning and artificial intelligence algorithms, these methods usually rely on a large amount of historical data or manual annotation, lacking sufficient self-learning and adaptive adjustment capabilities. For example, when a reinforcement learning-based scheduling algorithm faces a high-dimensional complex system, it may not achieve satisfactory results due to insufficient data volume or inefficient learning process. At the same time, traditional machine learning methods require a large amount of prior knowledge or precisely labeled data and cannot be efficiently trained without sufficient data.
[0008] Therefore, the main problem of the existing technology is the lack of an intelligent scheduling method that can perceive environmental changes in real time, flexibly schedule, and optimize energy distribution. In addition, most of the existing energy prediction models and scheduling algorithms rely on artificially designed rules or known external data, lacking adaptability and the ability to cope with complex dynamic environments, which leads to inefficiency and limitations in energy scheduling and management. Especially in complex distributed photovoltaic energy storage systems, how to maintain efficient energy utilization in a dynamic environment remains an unsolved challenge. Summary of the Invention
[0009] The object of the present invention is to design a group regulation and control digital twin method for a road surface distributed photovoltaic energy storage system, combining a distributed energy prediction and scheduling strategy based on self-supervised learning. Through real-time perception of environmental changes and intelligent learning, efficient scheduling and optimized management of the distributed photovoltaic energy storage system in a dynamic environment are achieved.
[0010] To achieve the above object, the present invention provides a group regulation and control digital twin method for a road surface distributed photovoltaic energy storage system, and the method includes the following steps:
[0011] Collect environmental data and system state data in real time, and generate environmental features through adaptive filtering and standardization;
[0012] Construct a digital twin model based on a multi-level correlation regression network model, and dynamically predict the operating state of the photovoltaic energy storage system by combining environmental features and historical states; the multi-level correlation regression network model is used to capture the spatio-temporal cross-effect of environmental features and historical states as the system state at the current moment, and dynamically predict the system state;
[0013] Based on a self-supervised learning framework, use the predicted system operating state and environmental features to predict the photovoltaic power generation and energy demand in future periods;
[0014] According to the prediction results and real-time environmental data, generate an environment-adaptive scheduling strategy through a multi-objective optimization algorithm, and dynamically adjust the operating modes of the power generation and energy storage units;
[0015] Based on a closed-loop feedback mechanism, the execution result of the scheduling is monitored in real time, and the scheduling strategy is dynamically corrected in combination with environmental changes to ensure the stable operation of the system.
[0016] Furthermore, the environmental data is collected based on multiple sensors; the sensors include a light sensor, a temperature sensor, a humidity sensor, and a system status sensor; the environmental data includes light intensity, temperature, humidity, and traffic flow data; the system status data includes the battery power of the energy storage, the photovoltaic power generation, and the grid load;
[0017] The environmental characteristics include light intensity, temperature, and humidity data.
[0018] Furthermore, the calculation formula of the digital twin model is expressed as:
[0019] S(t) = α1·F e (t) + α2·S(t - 1) + γ1·F e (t) ⊙ S(t - 1) + R(t)
[0020] Among them, S(t) is the system state at the current moment; F e (t) is the environmental characteristic at the current moment; α1 and α2 are the weighting coefficients of the environmental characteristic and the historical state; γ1 is the weighting coefficient of the cross effect, which is used to measure the non-linear interaction effect between the environment and the historical state; R(t) is the regularization term, which prevents the model from overfitting and enhances its robustness.
[0021] Furthermore, the digital twin model also includes a dynamic adjustment mechanism; the dynamic adjustment mechanism is specifically:
[0022] When the prediction error ∈ t exceeds the set threshold, the weight coefficients α1 and α2 and the cross effect coefficient γ1 are automatically updated; the regularization term R(t) is determined based on the mean of the historical system state and the mean of the historical environmental characteristics.
[0023] Furthermore, the tasks of the self-supervised learning framework include: power generation prediction and energy demand prediction, which are expressed as:
[0024] P pred (t) = W1·S(t) + W2·F e (t) + C1·S(t - 1) + C2·F e (t - 1) + b
[0025] Among them, P pred (t) is the predicted power generation or energy demand; S(t) is the system state at the current moment; F e(t) is the environmental feature at the current moment; W1 and W2 are weight matrices for system state and environmental features; C1 and C2 are weight coefficients for the system state and environmental features at the current moment, reflecting the temporal dependence of the system; b is the bias term of the model, representing a constant offset.
[0026] Furthermore, the loss function of the self-supervised learning framework is obtained by weighted summation of the prediction error term, self-consistency constraint, and time-varying regularization term;
[0027] Among them, the prediction error term is determined by the square of the second norm based on the error between the predicted power generation or energy demand and the actual power generation or energy demand;
[0028] The self-consistency constraint is determined by the square of the second norm based on the error between the predicted power generation or energy demand at the previous time step and the current predicted power generation or energy demand;
[0029] The time-varying regularization term is used to control the penalty degree of temporal variation;
[0030] The self-supervised learning framework is trained through the backpropagation algorithm to optimize the weight parameters; during the training process, the mini-batch gradient descent method is adopted to improve the training efficiency, and at the same time, the dynamic learning rate adjustment strategy is used to avoid falling into local optima.
[0031] Furthermore, the multi-objective optimization algorithm is expressed as:
[0032]
[0033] Among them, P pv (t + k) is the predicted value of the photovoltaic power generation at the (t + k)-th moment; P demand (t + k) is the predicted value of the power demand at the (t + k)-th moment; S(t + k) is the power system state at the (t + k)-th moment; S target is the target system state, that is, the desired system operating state; ω1 and ω2 are weight coefficients, controlling the weight ratio between energy balance and system stability; is the regularization term of the scheduling strategy, used to suppress excessive fluctuations; λ(t) is the environmental adaptation coefficient, indicating the degree to which the system adjusts the scheduling strategy according to external environmental changes; is the environmental adaptability term, indicating the influence degree of environmental changes on scheduling;
[0034] Among them, during the scheduling optimization process, the following constraints are also included:
[0035] The photovoltaic power generation cannot exceed the maximum value;
[0036] The charging and discharging power of the energy storage device has upper and lower limits, and the stored energy of the battery must be within a certain range;
[0037] Within a certain time window, the error between the power generation and the demand should not be too large to ensure load matching;
[0038] According to the volatility of the environmental data, the adaptability of the control system to external environmental changes is adjusted through the environmental adaptation coefficient λ(t).
[0039] Furthermore, a hybrid optimization method based on deep reinforcement learning and evolutionary algorithm is used to solve the multi-objective optimization algorithm, specifically including:
[0040] The state space at each time step consists of the following elements: the current state of the power system, the predicted photovoltaic power generation and power demand, and environmental variables;
[0041] The action space is the photovoltaic strategy;
[0042] The reward function provides feedback based on the difference between the current scheduling strategy and the target, promotes the learning of the optimal scheduling strategy by reducing the energy difference, and at the same time suppresses excessive fluctuations, expressed as:
[0043] R(t + k) = -(P pv (t + k) - P demand (t + k)) 2
[0044] Among them, R(t + k) is the reward function value, P pv (t + k) is the predicted photovoltaic power generation, P demand (t + k) is the predicted power demand;
[0045] The evolutionary algorithm is a genetic algorithm, and the operations include:
[0046] Crossover: Mix two different scheduling schemes to form a new candidate solution;
[0047] Mutation: Make a small adjustment to the current solution to explore new solutions;
[0048] Selection: Select the best scheduling scheme according to the fitness evaluation and continue to evolve.
[0049] Furthermore, the closed-loop feedback mechanism includes real-time collection of execution feedback data, calculation of the scheduling error, and dynamic adjustment of the learning rate through the environmental adaptability factor to correct the scheduling strategy; the environmental adaptability factor is dynamically adjusted based on the real-time environmental change rate to enhance the system's response ability to sudden environmental fluctuations.
[0050] Further, the method further includes continuously updating the digital twin model and the self-supervised learning model based on newly collected data to adapt to long-term environmental evolution and system operating state changes.
[0051] The beneficial technical effects of the present invention are at least as follows:
[0052] By obtaining environmental data (such as light, temperature and humidity, traffic flow, etc.) and system state data (such as energy storage battery power, photovoltaic power generation, etc.) in real time, and combining digital twin technology to update and simulate the state of each photovoltaic energy storage unit in real time, the scheduling decision can be dynamically adjusted according to different environmental conditions. Compared with the existing rule-based scheduling method, this innovative method can more accurately capture the impact of environmental changes on factors such as photovoltaic power generation efficiency and energy storage state, and flexibly adjust the power generation output and energy storage strategy of photovoltaic units, thereby improving the overall energy utilization efficiency of the system.
[0053] The present invention also proposes to use a self-supervised learning algorithm for energy demand prediction and scheduling optimization. Through deep learning of historical data, this algorithm can automatically identify the internal relationships between parameters such as photovoltaic power generation, energy storage battery charge and discharge efficiency, and load demand, and predict the energy demand and power generation in future periods. Compared with the existing supervised learning algorithms, self-supervised learning does not rely on a large amount of labeled data and can still achieve high prediction accuracy when the data volume is small or incomplete. This innovation effectively solves the limitations of relying on a large amount of historical data and manual annotation in the existing methods, and improves the prediction accuracy and adaptive ability of the system.
[0054] The present invention also establishes a closed-loop feedback mechanism, combines the execution situation of the scheduling decision with real-time data feedback, and realizes the real-time adjustment and optimization of the system. Specifically, when the system detects a deviation between the actual operation and the prediction, it will update and adjust the scheduling strategy in real time through the digital twin model to ensure the stable operation of the system in a complex environment. This mechanism overcomes the deficiencies of the lack of real-time feedback and dynamic optimization in the existing scheduling methods, and improves the real-time performance and flexibility of the system.
[0055] Through the above innovations, the present invention can perform real-time perception and scheduling adjustment for dynamic environmental changes in the photovoltaic energy storage system, not only improving the accuracy of energy management, but also being able to intelligently adjust the power generation and energy storage strategies according to the fluctuations of the system load, thus effectively solving problems such as uneven energy distribution, large prediction errors, and low system efficiency in the existing technology. The technical solution of the present invention is of great significance in realizing efficient energy utilization, reducing energy waste, and improving the flexibility and intelligent level of the system, and is particularly suitable for distributed photovoltaic energy storage systems in complex environments such as road surfaces. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The present invention will be further described with reference to the accompanying drawings. However, the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on the following drawings without creative efforts.
[0057] Figure 1 This is a flow chart of the group regulation and control digital twin method for the road surface distributed photovoltaic energy storage system of the present invention. Specific embodiments
[0058] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0059] In one or more embodiments, as Figure 1 shown, a group regulation and control digital twin method for a road surface distributed photovoltaic energy storage system is disclosed. The method includes the following steps:
[0060] S1. Collect environmental data and system status data in real time, and generate environmental features through adaptive filtering and standardization.
[0061] Specifically, first, the present invention collects relevant environmental variables by deploying multiple sensors in the system operation area. The sensor array should include the following categories:
[0062] Light sensors: Collect light intensity data at different locations.
[0063] Temperature sensors: Monitor environmental temperature changes.
[0064] Humidity sensors: Provide humidity data, which helps to predict the attenuation of the performance of photovoltaic panels.
[0065] Traffic flow sensors: In the case of urban areas, traffic flow data helps to predict local load changes.
[0066] System status sensors: Include important data such as the battery level of energy storage and the photovoltaic power generation.
[0067] The data collected by the present invention is represented in matrix form:
[0068] D = [D1 D2 D3 … D N (1)
[0069] where D i represents the i-th type of data (such as light, temperature, etc.), D is an N×T matrix, N represents the number of sensor types, and T is the time step of the data.
[0070] Furthermore, after data acquisition, the present invention cleans, denoises, and normalizes the original data to eliminate measurement errors and external interferences and ensure data quality. Here, an "Adaptive Filter Algorithm (AFA)" defined by the present invention is used to remove short-term random fluctuations and outliers. This method combines weighted mean filtering and local regression models, and reduces the impact of noise on data quality through weighted averaging and linear regression processing of data in adjacent time periods. The specific steps are as follows:
[0071] Let the acquired original data be D. First, data smoothing is performed through the following formula:
[0072]
[0073] where w i is the smoothing weight coefficient, and k is the smoothing window size. The present invention uses a weighted window function w i to weight the data according to the time interval, and the weight function decays as the time distance increases. In this way, the present invention can eliminate sudden short-term fluctuations.
[0074] Secondly, local regression analysis is used to identify and correct outliers. For the original data D(t) with a time step of t, if its deviation from the front and back data exceeds a certain set threshold τ, then this data is considered an outlier and is corrected through weighted regression:
[0075]
[0076] where is the predicted value calculated based on historical data, and λ is the correction coefficient, usually taking a relatively small value (such as 0.1) to ensure the smoothness of data correction.
[0077] Finally, to avoid the influence caused by different dimensions between different variables, the present invention normalizes all data. The normalization formula is:
[0078]
[0079] where μ D and σ D are the mean and standard deviation of all acquired data respectively. Through this step, the ranges of all environmental data are mapped to the standard normal distribution, improving the stability of model training and optimization in subsequent steps.
[0080] Furthermore, the preprocessed data needs to further extract the most valuable features for the photovoltaic energy storage system. For example, the changes in light intensity and temperature have a greater impact on photovoltaic power generation. Therefore, it is necessary to calculate their correlation with the energy storage battery power and generate the environmental feature F e (t):
[0081] F e (t)=[D light (t), D temp (t), D humidity (t)](5)
[0082] Among them, F e (t) is the processed environmental feature, including light intensity, temperature and humidity data.
[0083] Finally, the output data set D processed will be used as the input for the digital twin model and self-supervised learning algorithm in the subsequent steps. After operations such as preprocessing, denoising, standardization, and feature extraction, this data set can accurately reflect the environmental state and system operation conditions, and become a reliable input for digital twin and energy prediction.
[0084] S2. Build a digital twin model based on a multi-level correlation regression network model, and dynamically predict the operation state of the photovoltaic energy storage system by combining environmental features and historical states; the multi-level correlation regression network model is used to capture the spatio-temporal cross-effect of environmental features and historical states as the system state at the current moment, and dynamically predict the system state.
[0085] Specifically, this solution adopts an innovative multi-level correlation regression network (MARN) model, which comprehensively considers the relationship between the environmental feature F e (t) and the historical state S(t - 1), and can effectively capture the spatio-temporal cross-effect. The core formula of the model is:
[0086] S(t)=α1·F e (t)+α2·S(t - 1)+γ1·F e (t)⊙S(t - 1)+R(t)(6)
[0087] Among them, S(t) is the system state at the current moment (such as photovoltaic power generation, battery charge state, etc.). F e (t) is the environmental feature vector at the current moment (such as solar radiation, temperature, etc.). α1, α2 are the weighted coefficients of environmental features and historical states. γ1 is the weighted coefficient of the cross-effect, which is used to measure the non-linear interaction between the environment and historical states. R(t) is the regularization term, which prevents the model from overfitting and enhances its robustness.
[0088] Furthermore, to improve the model's adaptability to sudden environmental changes, a dynamic adjustment mechanism is introduced. Specifically, when the prediction error ∈1 of the model exceeds the set threshold, the weight coefficients α1, α2 and the cross-effect coefficient γ1 are automatically updated. The prediction error ∈ t is calculated as follows:
[0089] ∈ t =|S pred (t)-S actual (t)|(7)
[0090] In addition, by introducing a regularization term R(t) to suppress the over-complication of the model and maintain the stability of the prediction results. The calculation formula of the regularization term is:
[0091]
[0092] where S avg (t) is the historical system state mean value, used as a comparison benchmark. F avg (t) is the historical environmental feature mean value. λ is the regularization coefficient, which controls the model complexity.
[0093] Furthermore, the operating state of the photovoltaic energy storage system is greatly affected by time and space factors. Therefore, an adaptive adjustment mechanism is added to the MARN model. This mechanism dynamically adjusts the parameters α1, α2, γ1 in the regression model according to the change rates of environmental features and historical data. This adjustment ensures that the model can adapt to changes in different time periods and spatial regions, improving the prediction accuracy. For example, when the system is in an extreme weather environment, the model will automatically adjust to better predict the system state.
[0094] Furthermore, after the model is constructed, the model is verified through historical data and real-time data, the prediction error of the model is calculated, and optimized adjustment is performed according to the feedback. By continuously evaluating the error, the accuracy of the digital twin model is ensured, and the weights and parameters of the regression model are continuously improved so that it can still maintain high-efficiency and stable performance during long-term operation.
[0095] Furthermore, as the system operation progresses and the environmental conditions change continuously, the digital twin model needs to be continuously optimized and updated. Every time new data is collected, the model uses the error feedback mechanism for retraining to continuously optimize the prediction accuracy. Finally, through the continuously optimized digital twin model, the system can long-term adapt to different environmental changes and operating conditions and provide accurate decision-making support for subsequent energy scheduling.
[0096] Through this series of steps, the present invention constructs a digital twin model with high adaptability and high prediction accuracy for a photovoltaic energy storage system. This model can provide real-time state prediction during actual operation and offer reliable data support for subsequent energy scheduling and optimization.
[0097] S3. Based on the self-supervised learning framework, utilize the predicted system operating state and environmental characteristics to predict the photovoltaic power generation and energy demand in future time periods.
[0098] Specifically, to achieve accurate prediction of energy demand and power generation, this step constructs a self-supervised learning framework. Different from traditional supervised learning methods, this framework does not rely entirely on labeled data, but automatically constructs state transition relationships through self-supervised learning using the historical behavior of the system and environmental characteristics.
[0099] Among them, this framework includes two core tasks: power generation prediction P pv (t) and energy demand prediction P demand (t). The goal is to learn the historical system states S(t - 1), S(t - 2),... and environmental characteristics F e (t - 1), F e (t - 2),... through the model, and combine a certain time series delay to achieve the prediction of power generation and energy demand at future moments through self-supervised learning.
[0100] Furthermore, the core prediction function of the model is as follows:
[0101] P pred (t) = W1·S(t) + W2·F e (t) + C1·S(t - 1) + C2·F e (t - 1) + b(9)
[0102] Among them, P pred (t) is the predicted power generation or energy demand. S(t) is the system state at the current moment (from the digital twin model in the previous step). F e (t) is the environmental characteristic at the current moment (such as light, temperature, etc.). W1, W2 are weight matrices for system state and environmental characteristics. C1, C2 are weight coefficients for the system state and environmental characteristics at the previous moment, reflecting the time series dependence of the system. b is the bias term of the model, representing a constant offset.
[0103] Furthermore, this model framework introduces the characteristics of time series by combining the current state S(t) and the historical data S(t - 1), F e (t - 1) at the previous moment. This design not only enhances the system's memory ability for historical states but also effectively captures the system's response to environmental changes at different time points.
[0104] Furthermore, to ensure effective training of the model without labels, a reasonable loss function needs to be designed that can simultaneously optimize the accuracy of power generation prediction and energy demand prediction and ensure the temporal consistency of the model.
[0105] The present invention designs a dual loss function that combines prediction error and self-consistency constraint. The prediction error term is used to measure the difference between the model's predicted value and the actual value, while the self-consistency constraint ensures that the model can maintain consistency over consecutive time steps and reduce prediction fluctuations. The formula for the prediction error term is:
[0106]
[0107] where P true (t) is the actual power generation or energy demand. is the square of the L2 norm, used to calculate the prediction error.
[0108] The formula for the self-consistency constraint is:
[0109]
[0110] P pred (t - 1) is the predicted value at the previous time step.
[0111] Furthermore, to improve the model's adaptability to time series data, the present invention also designs an innovation term based on regularization to penalize the instability of the model in the time series. Especially when the system state undergoes a sudden change, the predicted output of the model should remain stable. The formula for the innovation regularization term is:
[0112]
[0113] where α is the regularization coefficient, used to control the degree of penalty for temporal changes. is the derivative of the predicted value with respect to time, reflecting the rate of change of the system state.
[0114] The final loss function is:
[0115]
[0116] where λ1, λ2, λ3 are the weighting coefficients of the loss function, adjusting the relative importance of different loss terms.
[0117] The design of this loss function enables the model to not only focus on prediction accuracy but also prevent overfitting and temporal fluctuations through self-consistency constraints and regularization terms. This is of great significance for energy systems because power demand and power generation are often affected by various complex factors and require the model to maintain stability and reliability.
[0118] Further, in this step, the present invention trains the above model through the backpropagation algorithm to optimize all parameters (i.e., W1, W2, C1, C2, b). During the training process, the mini-batch gradient descent method (Mini-batch SGD) is adopted to improve the training efficiency, and at the same time, the dynamic learning rate adjustment strategy is used to avoid falling into local optima.
[0119] During the training process, the present invention also uses the "online learning" technology to ensure that the model can continuously adapt to the changes of new data. Through periodic verification, the prediction accuracy and stability of the model are monitored.
[0120] Once the model is trained and verified, it can be used in actual applications to predict future photovoltaic power generation and energy demand. According to the actual environmental input F e (t), and the real-time system state S(t), the model will generate P pred (t), which serves as the basis for the next step of scheduling optimization and energy allocation. This predicted value will provide decision support for the energy management system, help the scheduling system predict future power demands and match them with photovoltaic power generation, thereby optimizing the operation efficiency of the system. The prediction accuracy of the model directly affects the energy allocation and resource utilization of the system, ensuring the reasonable scheduling and conservation of energy.
[0121] S4. According to the prediction results and real-time environmental data, generate an environment-adaptive scheduling strategy through a multi-objective optimization algorithm to dynamically adjust the operation modes of the power generation and energy storage units.
[0122] The input of this step comes from the results obtained in the previous step (self-supervised learning of energy demand and power generation prediction), mainly including:
[0123] Energy demand prediction: P demand (t), which predicts the future power demand through historical data and the self-supervised learning model.
[0124] Power generation prediction: P pv (t), which is the photovoltaic power generation obtained through weather prediction and system model prediction.
[0125] Environmental characteristics: F e (t), such as environmental data like weather, temperature, humidity, etc.
[0126] System state: S(t), including battery energy storage state, grid load, etc.
[0127] At this time, the input data is power demand, power generation, environmental characteristics, and system state, which will serve as the core input of the scheduling optimization model.
[0128] Furthermore, in this step, the present invention designs a scheduling optimization framework based on environmental adaptability. The core of this framework is to achieve the best match between power demand and generation capacity through scheduling strategies, while ensuring the stability and reliability of the power system.
[0129] Furthermore, to achieve this goal, the objective function of the scheduling optimization model includes the following items:
[0130] Energy balance term: Balance the power generation and demand, and minimize the difference between generation and demand.
[0131] System stability term: Ensure that the system can be maintained within the target range by controlling the change of the system state.
[0132] Environmental adaptability term: Adjust the scheduling strategy so that the system can adapt to the changes in the external environment and avoid the instability of the system performance caused by environmental factors.
[0133] The mathematical expression of the objective function is:
[0134]
[0135] where, P pv (t + k) is the predicted value of the photovoltaic power generation at the (t + k)-th moment. P demand (t + k) is the predicted value of the power demand at the (t + k)-th moment. S(t + k) is the state of the power system at the (t + k)-th moment (such as battery energy storage, load, etc.). S target is the target system state, that is, the desired system operating state. ω1, ω2 are weight coefficients, which control the weight ratio between energy balance and system stability. is the regularization term of the scheduling strategy, which is used to suppress excessive fluctuations. λ(t) is the environmental adaptation coefficient, which represents the degree to which the system adjusts the scheduling strategy according to changes in the external environment (such as temperature, humidity, etc.). is the environmental adaptability term, which represents the influence degree of environmental changes on scheduling, and the influence brought by environmental changes can be quantified through model calculation.
[0136] Among them, the innovation points of this objective function are:
[0137] Environmental adaptability term: By introducing the environmental adaptability term the system can dynamically adjust the scheduling strategy to cope with environmental fluctuations. λ(t) adjusts the sensitivity of environmental adaptability to ensure that the scheduling system can make reasonable responses to environmental changes at different times.
[0138] Regularization term: Introduce This regularization term can effectively suppress high-frequency fluctuations during the scheduling process, ensure the smoothness of the scheduling scheme, and prevent overly frequent scheduling changes.
[0139] Design of the regularization term: To ensure the smoothness of scheduling, the regularization term adopts a new design that not only considers the error between power generation and demand but also incorporates a frequency penalty term:
[0140]
[0141] where α and β are penalty coefficients that control the penalties for the variation ranges of power generation and demand.
[0142] Frequency penalty term: By calculating the differences in power generation between each moment and the previous moment, and the differences in demand between each moment and the previous moment, high-frequency fluctuations are suppressed to avoid overly frequent scheduling adjustments.
[0143] Furthermore, during the scheduling optimization process, to ensure the stability of the system and the feasibility of the constraint conditions, a series of constraints must be introduced:
[0144] Power generation limit: The photovoltaic power generation cannot exceed the maximum value P max , that is:
[0145] P pv (t + k) ≤ P max (t + k)(16)
[0146] Battery energy storage limit: The charging and discharging power of the energy storage device has upper and lower limits, and the battery energy storage must be within a certain range:
[0147] E(t + k) ∈ [E min , E max (17)
[0148] Demand matching error: Within a certain time window, the error between power generation and demand should not be too large to ensure load matching:
[0149] |P pv (t + k) - P demand (t + k)| ≤ ∈(18)
[0150] Environmental adaptability adjustment: According to the volatility of environmental data, the system's adaptability to external environmental changes is controlled through λ(t):
[0151] P adjust (t + k) = P pred (t + k)·λ(t)(19)
[0152] Furthermore, the core task of this step is to solve the scheduling optimization problem. Since the scheduling optimization problem usually has complex constraints and non-linear characteristics, traditional optimization methods (such as linear programming and integer programming) cannot effectively handle it. To overcome this challenge, the present invention adopts a hybrid optimization method based on deep reinforcement learning (DQN) and evolutionary algorithm.
[0153] State space: In scheduling optimization, the state space includes all possible environmental variables, system states, and predicted data within a time period. Specifically, the state s(t + k) consists of the following elements at each time step:
[0154] Current state of the power system: such as the state of battery energy storage E(t + k), load state, etc.
[0155] Predicted photovoltaic power generation P pv (t + k) and power demand P demand (t + k).
[0156] Environmental variables: such as temperature F e (t + k), humidity, etc.
[0157] Action space: At each time step, the system can select different scheduling strategies P adjust (t + k). This includes strategies such as adjusting the usage of photovoltaic power generation and controlling battery charging and discharging.
[0158] Reward function: The core of DQN is the reward function, which provides feedback based on the difference between the current scheduling strategy and the target. In this scenario, the present invention designs a reward function based on the objective function to maximize the energy balance and minimize the system fluctuations and differences in environmental adaptability.
[0159] Reward function:
[0160] R(t + k) = -(P pv (t + k) - P demand (t + k)) 2 (20)
[0161] The reward function promotes the learning of the optimal scheduling strategy by reducing the energy difference and at the same time suppresses excessive fluctuations.
[0162] Furthermore, although the deep Q-network performs well in local optimization, it may be prone to falling into local optimal solutions. To avoid this problem and accelerate the search for the global optimal solution, the present invention introduces an evolutionary algorithm, which effectively explores the scheduling strategy space through a combination of global search and local optimization.
[0163] Genetic Algorithm: The present invention uses an evolutionary strategy based on the genetic algorithm, where individuals represent different scheduling schemes. Through operations such as crossover, mutation, and selection, the quality of candidate solutions is continuously optimized. The fitness of each individual is evaluated by the objective function and the operations of the genetic algorithm include:
[0164] Crossover: Mix two different scheduling schemes to form a new candidate solution.
[0165] Mutation: Make small adjustments to the current solution to explore new solutions.
[0166] Selection: Select the best scheduling scheme according to the fitness evaluation to continue evolving.
[0167] By combining DQN and the evolutionary algorithm, the present invention realizes an efficient hybrid optimization strategy, which can find the global optimal solution in a complex and high-dimensional scheduling space and dynamically adjust the system according to real-time data.
[0168] Furthermore, after the scheduling optimization algorithm determines the preliminary scheduling strategy, the system must be able to perform dynamic adaptive updates according to changes in the external environment. This process mainly depends on the incremental learning and environmental adaptability module.
[0169] According to the change of environmental feature F e (t) (such as temperature, humidity, sunlight intensity, etc.), the system should dynamically adjust its scheduling strategy. The environmental adaptability update mechanism is adjusted based on the environmental adaptability term in the objective function to carry out regulation.
[0170] Incremental Learning: The system adjusts the scheduling strategy in real time according to the environmental data and historical scheduling results at each moment. Through continuous learning, the scheduling decision is gradually optimized. For example, if the change in temperature causes fluctuations in photovoltaic power generation, the system automatically adjusts the scheduling strategy of battery energy storage and grid power through the environmental adaptability adjustment coefficient λ(t).
[0171] Environmental Adaptability Adjustment: The present invention defines the environmental adaptability term to quantify the impact of environmental changes on the scheduling strategy. By dynamically adjusting λ(t), the system can adapt to fluctuations in the external environment.
[0172] The process of environmental adaptability update is as follows:
[0173] P adjust (t + k) = P pred (t + k)·λ(t)(21)
[0174] Dispatch History Update: Whenever there are changes in the external environment, the system will gradually update the dispatch strategy based on historical dispatch decisions and feedback. This mechanism can improve the system's adaptability and robustness, prevent over-reliance on single decisions, and ensure that the system can respond to external changes in real time.
[0175] Once an environmental change triggers a dispatch update, the system will immediately adopt a new dispatch strategy within the next time step. By continuously adapting to environmental changes, the system ensures that it can achieve optimal power dispatch in the long run and optimize the stability and economy of the power system.
[0176] Furthermore, after the dispatch strategy is optimized and adapted to the environment, the final obtained dispatch strategy P adjust (t + k) will be passed as input to the Energy Management System (EMS). The EMS is responsible for actually executing the dispatch tasks and collecting feedback data during the execution process for further optimizing dispatch decisions.
[0177] Dispatch Execution: The EMS executes the energy dispatch tasks according to the current dispatch plan, which includes controlling photovoltaic power generation, adjusting the battery energy storage status, dispatching the grid load, etc. During the real-time execution process, the EMS collects new environmental data (such as temperature and humidity changes, light intensity changes, etc.) and system status data (such as battery power, grid load, etc.) through sensors and monitoring devices.
[0178] Feedback Mechanism: The EMS will return the feedback information during the execution process (such as the difference between the dispatch result and the actual demand) to the dispatch optimization model in real time. Based on this feedback data, the optimization model will fine-tune the dispatch strategy to ensure more accurate future dispatch decisions.
[0179] Fine-tuning and Optimization: When there is a deviation between the actual dispatch result and the predicted demand and power generation, the EMS will adjust the dispatch strategy according to the feedback signal and update the dispatch decision through an optimization algorithm. This feedback mechanism enables the system to achieve self-optimization and adaptation during long-term operation, continuously improving the operation efficiency and stability of the power system.
[0180] In this step, the present invention innovatively designs a dispatch optimization framework based on environmental adaptability, and combines reinforcement learning and evolutionary algorithms to achieve optimal dispatch under different environmental conditions. The introduced regularization term, environmental adaptability adjustment, and the combination of reinforcement learning and evolutionary algorithms solve the problems that traditional dispatch methods cannot flexibly cope with environmental fluctuations and system instability caused by frequent fluctuations. This solution can not only ensure the balance between power generation and demand, but also gradually improve the stability and economy of the power system in long-term operation.
[0181] S5. Based on the closed-loop feedback mechanism, monitor the dispatch execution results in real time, and dynamically correct the dispatch strategy in combination with environmental changes to ensure the stable operation of the system.
[0182] Specifically, after the power dispatching is executed, the system will collect new environmental data and system status information through real-time monitoring to form a closed-loop feedback. This feedback includes:
[0183] Environmental changes: environmental characteristics F e such as temperature, humidity, light intensity, etc., and the real-time changes of F
[0184] System status: real-time monitoring of system status S
[0185] such as battery energy storage status, grid load changes, etc., at time (t + k). adjust Feedback after the execution of the dispatching policy P
[0186] (t + k), mainly the deviation between the actual power generation and the demand.
[0187] These feedback information will be transmitted to the next stage through the following process to achieve dynamic adjustment. adjust Further, the present invention first calculates the actually executed dispatching policy P target (t + k) and the system expected target P
[0188] The error ΔP(t + k) between them is used as the adjustment basis. The calculation formula of the error term is as follows: adjust ΔP(t + k) = P target (t + k) - P
[0189] where P adjust (t + k) is the actually executed dispatching policy of the system at the (t + k)-th moment. P target (t + k) is the target dispatching policy predicted according to the model at the (t + k)-th moment (ideally, it should be close to P demand (t + k)).
[0190] According to the error ΔP(t + k), the system will judge whether the dispatching policy needs to be adjusted. If the error is large, it indicates that the difference between the current dispatching policy and the actual demand / generation is too large, and parameters need to be adjusted to bridge this gap.
[0191] Further, the system will use an adaptive adjustment algorithm to fine-tune the dispatching policy according to the error ΔP(t + k), the current environmental characteristics F e (t + k) and the system status S(t + k). Here, the present invention designs an "Error-Weighted Adjustment (EWA)" algorithm to adjust the policy. This algorithm takes into account the combined effects of dispatching errors and environmental factors:
[0192]
[0193] Among them, P adjust (t + k) new is the new adjusted scheduling strategy. γ(t + k) is the learning rate (or adjustment speed), which controls the amplitude of the adjustment and can be dynamically adjusted according to real-time feedback. ΔP(t + k) is the error term, indicating the gap between the actual and the target scheduling strategy. is the environmental adaptability factor, indicating the impact of environmental changes on the scheduling strategy. Specifically, is a regulation function based on environmental characteristics, and its calculation can refer to:
[0194]
[0195] Among them, α is the environmental change sensitivity parameter, reflecting the sensitivity of environmental changes to scheduling adjustments.
[0196] Through the above correction strategy, the system will dynamically adjust P adjust (t + k), and optimize the scheduling strategy at each moment to gradually approach the desired power demand and generation target.
[0197] Furthermore, the adjusted scheduling strategy will be fed back to the Energy Management System (EMS) to execute the new scheduling command, and continue to collect new feedback data during the execution. In this way, the system will continuously self-optimize through closed-loop feedback during long-term operation.
[0198] Among them, the iterative process of the feedback correction mechanism:
[0199] The initial scheduling strategy P adjust (t + k) is obtained and executed through an optimization algorithm.
[0200] Real-time collect the feedback data of the system, including the environmental characteristics F e (t + k) and the system state S(t + k).
[0201] Calculate the error ΔP(t + k) between the actually executed scheduling strategy and the target.
[0202] Based on the error and environmental adaptability, use the error weighted correction algorithm to adjust the scheduling strategy P adjust (t + k).
[0203] Transmit the adjusted scheduling strategy to the EMS for execution again, and continue to collect feedback.
[0204] This closed-loop adjustment process will be carried out at each time step to ensure that the system gradually optimizes the scheduling decision under changing environmental conditions, and finally realizes the stable and economic operation of the system.
[0205] Furthermore, to ensure that the power system can more effectively adapt to external environmental changes, the adjusted scheduling strategy P adjust (t + k) new will be readjusted according to real-time environmental data. For example, when the sunlight intensity changes sharply, the photovoltaic power generation may fluctuate significantly. The system will adjust to strengthen the response to environmental fluctuations and ensure that the scheduling strategy is more in line with the actual power generation and demand.
[0206] Through this dynamic feedback mechanism, the system can quickly respond to environmental changes and make corresponding adjustments, further optimizing the stability and economy of power dispatching. Through the closed-loop feedback and dynamic adjustment strategy, the system can real-time adjust the power dispatching plan to adapt to environmental changes and fluctuations in system status. Specifically, the Error-Weighted Adjustment algorithm (EWA) can adjust the scheduling strategy according to real-time feedback and environmental characteristics at each moment, making the dispatching process smoother and achieving system optimization in the long-term operation. This solution can effectively cope with the uncertainties and environmental changes in the actual power system, ensuring the reliability and efficiency of power dispatching.
[0207] In summary, the present invention can perform real-time perception and scheduling adjustment for dynamic environmental changes in the photovoltaic energy storage system, not only improving the accuracy of energy management, but also being able to intelligently adjust the power generation and energy storage strategies according to the fluctuations of the system load, thus effectively solving the problems of uneven energy distribution, large prediction errors, and low system efficiency in the prior art. The technical solution of the present invention is of great significance in realizing efficient energy utilization, reducing energy waste, and improving the flexibility and intelligent level of the system, and is particularly applicable to distributed photovoltaic energy storage systems in complex environments such as roads.
[0208] The above-disclosed are only some preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A digital twin method for group regulation and control of road-distributed photovoltaic energy storage systems, characterized in that: The method comprises the following steps: Collect environmental data and system status data in real time, and generate environmental features through adaptive filtering and standardization; A digital twin model based on a multi-level correlation regression network model is constructed to dynamically predict the operating status of the photovoltaic energy storage system by combining environmental characteristics and historical status; the multi-level correlation regression network model is used to capture the spatiotemporal cross-effects of environmental characteristics and historical status as the system status at the current moment, and dynamically predict the system status; Based on the self-supervised learning framework, the predicted system operation status and environmental characteristics are used to predict the photovoltaic power generation and energy demand in the future period; Based on the prediction results and real-time environmental data, an environmental adaptive scheduling strategy is generated through a multi-objective optimization algorithm to dynamically adjust the operating modes of power generation and energy storage units; Based on the closed-loop feedback mechanism, the scheduling execution results are monitored in real time, and the scheduling strategy is dynamically modified according to environmental changes to ensure stable operation of the system.
2. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 1 is characterized in that: The environmental data is collected based on multiple sensors; the sensors include light sensors, temperature sensors, humidity sensors and system status sensors; the environmental data include light intensity, temperature, humidity and traffic flow data; the system status data includes energy storage battery power, photovoltaic power generation power and grid load; The environmental characteristics include light intensity, temperature and humidity data.
3. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 1 is characterized in that: The calculation formula of the digital twin model is expressed as: S(t)=α1·F e (t)+α2·S(t-1)+γ1·F e (t)⊙S(t-1)+R(t) Among them, S(t) is the system state at the current moment; F e (t) is the environmental feature at the current moment; α1, α2 are the weighting coefficients of environmental features and historical states; γ1 is the weighting coefficient of the cross effect, which is used to measure the nonlinear interaction between the environment and historical states; R(t) is the regularization term, which prevents the model from overfitting and enhances its robustness.
4. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 3 is characterized in that: The digital twin model also includes a dynamic adjustment mechanism; the dynamic adjustment mechanism is specifically: When the model prediction error ∈ t When the set threshold is exceeded, the weight coefficients α1, α2 and the cross-effect coefficient γ1 are automatically updated; the regularization term R(t) is determined based on the historical system state mean and the historical environmental feature mean.
5. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 3 is characterized in that: The tasks of the self-supervised learning framework include: power generation prediction and energy demand prediction, which are expressed as: P pred (t)=W1·S(t)+W2·F e (t)+C1·S(t-1)+C2·F e (t-1)+b Among them, P pred (t) is the predicted power generation or energy demand; S(t) is the system state at the current moment; F e (t) is the environmental characteristic at the current moment; W1, W2 are the weight matrices for the system state and environmental characteristics; C1, C2 are the weight coefficients for the system state and environmental characteristics at the current moment, reflecting the timing dependence of the system; b is the bias term of the model, representing the constant offset.
6. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 5 is characterized in that: The loss function of the self-supervised learning framework is obtained by weighted summing of the prediction error term, the self-consistency constraint, and the time-varying regularization term; The prediction error term is determined by the square of the second norm based on the error between the predicted power generation or energy demand and the actual power generation or energy demand; The self-consistency constraint is based on the error between the predicted power generation or energy demand at the previous time step and the current predicted power generation or energy demand, determined by the square of the two norm; The time variation regularization term is used to control the degree of penalty for time series variation; The self-supervised learning framework is trained through a back-propagation algorithm to optimize weight parameters. During the training process, a small batch gradient descent method is used to improve training efficiency, and a dynamic learning rate adjustment strategy is used to avoid falling into a local optimum.
7. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 5 is characterized in that: The multi-objective optimization algorithm is expressed as: Among them, P pv (t+k) is the predicted value of photovoltaic power generation at the t+kth moment; P demand (t+k) is the power demand forecast value at the t+kth moment; S(t+k) is the power system state at the t+kth moment; S target is the target system state, i.e. the desired system operation state; ω1 and ω2 are weight coefficients, which control the weight ratio between energy balance and system stability; is the regularization term of the scheduling strategy, which is used to suppress excessive fluctuations; λ(t) is the environmental adaptation coefficient, which indicates the degree to which the system adjusts the scheduling strategy according to changes in the external environment; is the environmental adaptability term, which indicates the impact of environmental changes on scheduling; Among them, the following constraints are also included in the scheduling optimization process: The photovoltaic power generation power cannot exceed the maximum value; The charging and discharging power of energy storage equipment has upper and lower limits, and the battery storage capacity must be within a certain range; Within a certain time window, the error between power generation and demand should not be too large to ensure load matching; According to the volatility of environmental data, the adaptability of the system to external environmental changes is controlled by the environmental adaptability coefficient λ(t).
8. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 7 is characterized in that: A hybrid optimization method based on deep reinforcement learning and evolutionary algorithms is used to solve multi-objective optimization algorithms, including: The state space consists of the following elements at each time step: the current state of the power system, the predicted PV power generation and power demand, and environmental variables; The action space is the photovoltaic strategy; The reward function provides feedback based on the difference between the current scheduling strategy and the target, promoting the learning of the optimal scheduling strategy by reducing the energy difference while suppressing excessive fluctuations, which can be expressed as: R(t+k)=-(P pv (t+k)-P demand (t+k)) 2 Among them, R(t+k) is the reward function value, P pv (t+k) is the predicted photovoltaic power generation, P demand (t+k) is the predicted electricity demand; The evolutionary algorithm is a genetic algorithm, and the operations include: Crossover: Mix two different scheduling schemes to form a new candidate solution; Mutation: Make small adjustments to the current solution and explore new solutions; Selection: Select the best scheduling solution based on fitness evaluation to continue evolution.
9. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 1 is characterized in that: The closed-loop feedback mechanism includes real-time collection of execution feedback data, calculation of scheduling errors, and dynamic adjustment of learning rates through environmental adaptability factors to correct scheduling strategies; the environmental adaptability factors are dynamically adjusted based on the real-time environmental change rate to enhance the system's responsiveness to sudden environmental fluctuations.
10. The group adjustment and group control digital twin method for road distributed photovoltaic energy storage system according to claim 1, characterized in that: The method also includes continuously updating the digital twin model and the self-supervised learning model based on the newly collected data to adapt to the long-term environmental evolution and changes in the system operating status.
Citation Information
Patent Citations
Comprehensive energy control method and system based on digital twinning
CN116014715A
Distributed photovoltaic efficiency conversion optimization method and system based on digital twinning
CN116937661A
Group scheduling and group control digital twinning method for distributed photovoltaic system
CN118899914A