Energy management method and system based on big data
By constructing a big data-based energy management system, and utilizing digital twin modeling, multi-source data fusion, reinforcement learning, and closed-loop feedback mechanisms, the system addresses the shortcomings of existing technologies in real-time modeling, multi-source data fusion, and adaptive optimization in energy management. This achieves higher modeling accuracy, prediction accuracy, and optimization results, thereby reducing energy costs and carbon emissions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-27
Smart Images

Figure CN121414075B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of energy management, in particular to an energy management method and system based on big data, and more particularly to intelligent management and optimal configuration of enterprise energy systems using digital twin technology, multi-source data fusion technology and deep reinforcement learning technology. BACKGROUND
[0002] With the continuous growth of global energy demand and the increasing strictness of environmental protection requirements, enterprise energy management is facing dual pressures of cost control and carbon emission reduction. Traditional energy management methods mainly rely on manual experience and simple statistical analysis, which are difficult to cope with complex and variable energy consumption scenarios and dynamic market environment.
[0003] Chinese patent CN115994628B discloses an energy management method and device based on big data. The method collects historical carbon emission data of the target organization and its departments, calculates principal component vectors through principal component analysis, allocates carbon emission shares to each department using attention mechanism, and predicts active and passive carbon emissions through neural network. The method has the following shortcomings: first, the method only performs statistical analysis and prediction based on historical data, lacks dynamic modeling capability of real-time running state of energy system, and cannot accurately reflect the actual changes of physical system; second, the method only considers single dimension of carbon emission data, does not fuse multi-source data such as business operation, environmental parameters and price fluctuations, resulting in limited accuracy of prediction and decision-making; third, the method uses offline trained neural network for prediction, lacks online learning and adaptive optimization capability, and is difficult to cope with dynamic changes and uncertainties of energy system; fourth, the method lacks closed-loop feedback mechanism, and cannot dynamically adjust model and strategy according to actual execution effect, resulting in decline of long-term operation effect.
[0004] Therefore, there is an urgent need for an energy management method that can realize real-time modeling, multi-source data fusion, adaptive optimization and closed-loop feedback of energy system, to improve the intelligent level and economic benefits of enterprise energy management. SUMMARY
[0005] The present application aims to provide an energy management method and system based on big data, which solves the technical problems of lack of real-time modeling capability, insufficient multi-source data fusion, weak adaptive optimization capability and lack of closed-loop feedback mechanism in the prior art.
[0006] To achieve the above object, the application provides an energy management method based on big data, which constructs a deep coupling closed-loop collaborative system composed of a digital twin modeling module, a multi-source data fusion module, a reinforcement learning optimization module and a closed-loop feedback adjustment module. The digital twin modeling module realizes dynamic tracking and accurate description of the running state of the energy system by establishing a real-time mapping relationship between the physical energy system and the virtual energy model. The multi-source data fusion module integrates energy consumption data, business operation data, environmental parameter data and real-time electricity price data, and generates a unified energy state feature vector through an adaptive fusion weight algorithm, providing a comprehensive and accurate data basis for subsequent optimization decisions. The reinforcement learning optimization module uses a deep Q network algorithm to learn the optimal energy configuration strategy and can realize adaptive optimization in a complex dynamic environment. The closed-loop feedback adjustment module evaluates the energy configuration effect in real time, generates feedback adjustment parameters according to the actual energy efficiency deviation, and adjusts the digital twin model parameters, data fusion weights and learning strategies in reverse, forming a complete closed-loop feedback mechanism.
[0007] The application also provides an energy management system based on big data, which includes a memory and a processor, and realizes the above method by executing computer executable instructions.
[0008] The technical scheme of the application has the following beneficial effects:
[0009] Firstly, the digital twin modeling module establishes a real-time mapping relationship between the physical energy system and the virtual energy model, which can accurately reflect the dynamic running state of the energy system. Compared with the statistical analysis method based on historical data in the prior art, the application can track system changes in real time, and the modeling accuracy is improved by more than 40%.
[0010] Secondly, the multi-source data fusion module integrates multi-dimensional data such as energy consumption, business operation, environmental parameters and real-time electricity price. Compared with the prior art which only considers single-dimensional carbon emission data, the application can comprehensively consider various factors affecting energy management, and the prediction accuracy is improved by more than 25%.
[0011] Thirdly, the reinforcement learning optimization module uses a deep Q network algorithm to realize adaptive optimization. Compared with the method of offline training neural network in the prior art, the application can learn and continuously optimize the strategy online, adapt to the dynamic changes of the energy system, and the optimization effect is improved by more than 30%.
[0012] Fourthly, the closed-loop feedback adjustment module establishes a complete feedback adjustment mechanism. Compared with the prior art which lacks closed-loop feedback, the application can dynamically adjust the model parameters and optimization strategies according to the actual execution effect, ensure long-term stable operation, and the system reliability is improved by more than 35%.
[0013] Fifth, by constructing a deep coupling closed-loop collaborative system, mutual promotion and superposition of modules are realized, the overall performance of the system presents a nonlinear growth characteristic of 1+1>2, the comprehensive energy efficiency is improved by 20%, the energy cost is reduced by 15%-20%, and the carbon emission is reduced by more than 18%. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 It is a whole flow schematic diagram of the energy management method based on big data of the application.
[0015] Figure 2 It is a structural schematic diagram of the digital twin modeling module of the application.
[0016] Figure 3 It is a working flow schematic diagram of the multi-source data fusion module of the application.
[0017] Figure 4 It is a feedback mechanism schematic diagram of the closed-loop feedback adjustment module of the application.
[0018] Figure 5 It is a hardware structure schematic diagram of the energy management system based on big data of the application. DETAILED DESCRIPTION
[0019] Please refer to the accompanying Figures 1-5 , the purposes, technical solutions and advantages of the application will be clearer, the embodiments of the application will be further described in detail below with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the application, and are not used to limit the protection scope of the application.
[0020] Referring to Figure 1 , the application provides an energy management method based on big data, which is applied to an enterprise energy management system, and a deep coupling closed-loop collaborative system composed of a digital twin modeling module 1, a multi-source data fusion module 2, a reinforcement learning optimization module 3 and a closed-loop feedback adjustment module 4 is constructed. The system realizes intelligent management and optimal configuration of the energy system through deep coupling and closed-loop feedback between modules.
[0021] Referring to Figure 2 , the digital twin modeling module 1 is used for constructing a digital twin model of the enterprise energy system, and establishing a real-time mapping relationship between the physical energy system and the virtual energy model. The module includes a physical device parameter acquisition unit, a network topology construction unit, an operation state monitoring unit and a digital twin mapping unit.
[0022] The physical device parameter acquisition unit is responsible for acquiring the basic parameters of all energy devices in the enterprise energy system. The energy devices include but are not limited to transformers, power distribution cabinets, air conditioning systems, lighting systems, production equipment, gas boilers, heat pipe networks, etc. The acquired physical device parameters include device type, device number, rated power, operating efficiency, service life, maintenance record and device status, etc. In an embodiment of the present application, the Internet of Things sensors and industrial Ethernet technology are used to realize the automatic acquisition of device parameters, and the sensor sampling frequency is set to 1 Hz to 10 Hz, ensuring that the dynamic changes of device operation can be captured.
[0023] The network topology construction unit is used to construct the network topology structure of the energy system. The topology structure of the energy system describes the connection relationship and energy flow path between the energy devices. In an embodiment of the present application, the graph theory method is used to represent the energy network topology, the energy devices are abstracted as nodes, and the energy transmission paths between the devices are abstracted as edges, to construct a directed graph model of the energy system. The topology structure includes power network topology, gas pipe network topology, heat pipe network topology and cooling water pipe network topology, etc. The topology information includes parameters such as distance between nodes, pipe diameter, transmission loss and network impedance.
[0024] The running state monitoring unit is responsible for real-time acquisition of the running state parameters of each energy device. The running state parameters include real-time power, instantaneous energy consumption, cumulative energy consumption, running temperature, running pressure, flow, voltage, current, power factor and fault signal, etc. In an embodiment of the present application, by deploying intelligent electric meters, flow meters, temperature sensors, pressure sensors and other monitoring devices, comprehensive monitoring of the running state of the energy system is realized. The monitoring data is transmitted in real time to the data center through the industrial gateway, and the data transmission delay is controlled within 100 ms, ensuring that the digital twin model can reflect the state changes of the physical system in real time.
[0025] The digital twin mapping unit establishes the digital twin mapping relationship based on the acquired physical device parameters, network topology structure and running state parameters, and generates a virtual energy model. The core of digital twin mapping is to establish a two-way data flow channel between the physical entity and the virtual model, to realize the mapping from the physical space to the digital space and the control feedback from the digital space to the physical space.
[0026] In an embodiment of the present application, the digital twin mapping adopts an innovative state space mapping algorithm. The algorithm defines the physical system state vector and the digital twin model state vector , and establishes the relationship between the two through a mapping function :
[0027] ,
[0028] wherein, is the digital twin model state vector at time is the physical system state vector at time is the physical device parameter matrix, is the network topology matrix, is the mapping function. The physical system state vector contains real-time operating data of all monitoring points, the digital twin model state vector represents the state of the virtual model at the corresponding time, the physical device parameter matrix describes the inherent properties of each device, and the network topology matrix characterizes the structural relationship of the system.
[0029] To ensure that the digital twin model and the physical system remain highly synchronized, the present application adopts an adaptive parameter adjustment mechanism. This mechanism calculates the deviation between the virtual model output and the actual measured value of the physical system in real time, and dynamically adjusts the model parameters according to the deviation. The deviation calculation uses the weighted root mean square error:
[0030] ,
[0031] where is the model deviation at time , is the total number of monitoring points, is the weight of the th monitoring point, is the prediction value of the digital twin model for the th monitoring point, is the actual measured value of the th monitoring point. The weight is determined according to the importance of the monitoring point, and the monitoring point weight of the key device is larger.
[0032] When the model deviation exceeds the preset threshold , the parameter adjustment mechanism is triggered, and the gradient descent method is used to update the model parameters:
[0033] ,
[0034] where is the model parameter of the th iteration, is the updated model parameter, is the learning rate, is the gradient of the loss function with respect to the parameter The loss function is defined as the mean squared error between the model prediction and the actual measurement. In the preferred embodiment, the learning rate is set to 0.001, the preset threshold is set to 5%, which can avoid model oscillation caused by excessive adjustment while ensuring model accuracy.
[0035] Through the above digital twin mapping mechanism, the application can construct a virtual energy model highly synchronized with the physical energy system, providing accurate system state information for subsequent data fusion and optimization decision-making. Experiments show that the synchronization accuracy of the digital twin model reaches more than 95%, the model update delay is less than 1s, and the dynamic changes of the energy system can be reflected in real time.
[0036] Referring to Figure 3 , the multi-source data fusion module 2 is used to collect and fuse multi-dimensional data to generate a unified energy state feature vector. The module includes an energy data acquisition unit, a business data acquisition unit, an environmental data acquisition unit, a price data acquisition unit, a time alignment unit, a spatial consistency calculation unit, and an adaptive fusion unit.
[0037] The energy data acquisition unit is responsible for collecting the enterprise's energy consumption data. Energy consumption data includes power consumption, gas consumption, heat consumption, and water resource consumption, etc. For power consumption, parameters such as three-phase voltage, three-phase current, active power, reactive power, power factor, and cumulative energy are collected. For gas consumption, parameters such as instantaneous flow, cumulative flow, pipeline pressure, and temperature are collected. For heat consumption, parameters such as supply water temperature, return water temperature, flow, and heat cumulative value are collected. For water resource consumption, parameters such as instantaneous flow and cumulative water consumption are collected. In an embodiment of the application, the energy data sampling period is set to 1min to 15min, and is differentiated according to the characteristics of different energy types. The power data sampling period is 1min, the gas and heat data sampling period is 5min, and the water resource data sampling period is 15min.
[0038] The business data acquisition unit is responsible for collecting the enterprise's business operation data. Business operation data is closely related to energy consumption and can reflect the driving factors of energy demand. Business operation data includes production plans, production work orders, production, equipment utilization, personnel attendance, shift arrangements, and order information, etc. In an embodiment of the application, business operation data is automatically obtained by interfacing with the enterprise's ERP system, MES system, and human resource system. The business data update period is set according to the business characteristics. Production plans and work order information are updated every hour, production and equipment utilization are updated every 30min, and personnel attendance and shift information are updated every day.
[0039] The environmental data collection unit is responsible for collecting environmental parameter data that affect energy consumption. Environmental parameters have a significant impact on energy demand, especially for systems such as air conditioning, lighting, and ventilation. Environmental parameter data includes outdoor temperature, indoor temperature, relative humidity, light intensity, air quality index (AQI), and wind speed, among others. In one embodiment of the invention, environmental parameters are collected in real-time through the deployment of weather stations and indoor environmental monitoring devices. Weather stations are set up within the enterprise campus and can accurately reflect the microclimate characteristics of the location where the enterprise is located. Indoor environmental monitoring devices are distributed in different functional areas such as office areas, production workshops, and storage areas. The environmental data sampling period is set to 5-30 minutes, outdoor temperature and humidity are collected every 15 minutes, indoor temperature is collected every 5 minutes, and light intensity is dynamically adjusted according to the sampling period of day and night changes.
[0040] The price data collection unit is responsible for collecting real-time electricity price data. Electricity price is an important component of energy cost, and real-time electricity price information has a key impact on energy procurement and usage strategies. Price data includes peak-valley electricity price, real-time market electricity price, tiered electricity price, and demand electricity price, among others. In one embodiment of the invention, real-time electricity price information is obtained through data interfaces with power trading platforms and grid companies. For enterprises participating in power market transactions, real-time electricity price data is updated every 15 minutes. For enterprises implementing peak-valley electricity prices, electricity price data is set according to time period division, with peak period electricity price usually being 1.5-2 times the flat period electricity price, and valley period electricity price being 0.3-0.5 times the flat period electricity price.
[0041] The time alignment unit is used to handle the time synchronization problem of different data sources. Due to different sampling frequencies and update periods of different data sources, it is necessary to unify multi-source data to the same time granularity. In one embodiment of the invention, a timestamp alignment method is used to unify the timestamps of all data sources to the standard UTC time. For data with high sampling frequency, a downsampling method is used to aggregate data to a unified time granularity. For data with low sampling frequency, an interpolation method is used to fill in the missing data at the time points. The target time granularity of time alignment is set to 5 minutes, which can balance data accuracy and computational efficiency.
[0042] The time alignment degree is calculated as follows: for data source , calculate the deviation of its timestamp from the standard time grid, the smaller the deviation, the higher the time alignment degree. The time alignment degree is defined as:
[0043] ,
[0044] where is the time alignment degree of data source , is the time alignment degree of data source The total number of data points in the statistical period, is the data source The first The time difference between the timestamp of the data point and the nearest standard time grid point, is the standard time granularity. The value of time alignment ranges from 0 to 1, and the closer the value is to 1, the higher the time alignment.
[0045] The spatial consistency calculation unit is used to determine the degree of association between different data sources. Spatial consistency reflects the correlation between the physical processes or business processes described by different data sources. In an embodiment of the present application, the spatial consistency is calculated using the mutual information method. Mutual information can measure the statistical dependence between two random variables, not limited to linear relationships, and is suitable for complex nonlinear correlation analysis.
[0046] For data source and data source , the spatial consistency is defined as:
[0047] ,
[0048] wherein, is the spatial consistency between data source and data source , is the mutual information between the variable of data source and the variable of data source , is the information entropy of variable , is the information entropy of variable , indicates taking the minimum value. The calculation formula of mutual information is:
[0049] ,
[0050] wherein, is the joint probability distribution of variable and , is the marginal probability distribution of variable , is the marginal probability distribution of variable , indicates the logarithm with base 2. The calculation formula of information entropy is:
[0051] ,
[0052] in, For variables The probability distribution. Spatial consistency. The value ranges from 0 to 1, and the closer the value is to 1, the higher the degree of correlation between the two data sources.
[0053] The adaptive fusion unit assigns fusion weights to each data source based on time alignment and spatial consistency, generating a unified energy state feature vector. The allocation of fusion weights must consider the time synchronization of the data sources, the correlation between data sources, and the reliability and importance of the data sources.
[0054] In one embodiment of the present invention, an innovative adaptive fusion weighting algorithm is employed. This algorithm comprehensively considers time alignment, spatial consistency, and data source quality to dynamically calculate the fusion weights. For the data source... Its fusion weight The calculation formula is:
[0055] ,
[0056] in, For data source The fusion weight, For data source Time alignment For data source Other data sources Spatial consistency between them For data source Quality rating Total number of data sources , and For the weighting coefficients to satisfy Time alignment coefficient This reflects the importance of data synchronization; the spatial consistency coefficient This reflects the importance of data correlation, and the quality score coefficient. This reflects the importance of data reliability. In a preferred embodiment, Set to 0.4, Set to 0.3, Setting it to 0.3 achieves a balance between data synchronization, relevance, and reliability.
[0057] Data source quality rating The quality score is determined by comprehensively considering data completeness, accuracy, and freshness. Completeness refers to the missing data rate, accuracy to the error rate, and freshness to the timeliness of data updates. The calculation formula is:
[0058] ,
[0059] wherein, is the missing rate of the data source , is the error rate of the data source , is the freshness index of the data source . The missing rate is calculated by the proportion of missing data points in the statistical period, the error rate is calculated by comparison with the reference value, and the freshness index is calculated according to the interval between the data update time and the current time.
[0060] The final fusion weight needs to be normalized to ensure that the sum of the fusion weights of all data sources is 1:
[0061] ,
[0062] wherein, is the normalized fusion weight, is the original fusion weight, is the total number of data sources.
[0063] Based on the normalized fusion weight, a unified energy state feature vector is generated:
[0064] ,
[0065] wherein, is the fused energy state feature vector, is the normalized fusion weight of the data source , is the feature vector of the data source after standardization processing, is the total number of data sources. The feature vector contains all monitoring variables of the data source , which needs to be standardized before fusion to map variables of different dimensions to the same numerical range.
[0066] The multi-source data fusion method of the application can fully exploit the complementary information between different data sources, and realize dynamic selection and optimized combination of data sources through an adaptive fusion weight mechanism. Experiments show that the fused energy state feature vector has an information completeness degree improved by 45% and a prediction accuracy improved by 25% compared with a single data source, providing high-quality input for subsequent reinforcement learning optimization.
[0067] The reinforcement learning optimization module 3 optimizes energy allocation based on the energy state feature vector generated by the multi-source data fusion module, and uses the Deep Q-Network (DQN) algorithm to learn the optimal energy allocation strategy. This module includes a state space definition unit, an action space definition unit, a reward function definition unit, a value network unit, a target network unit, an experience replay unit, and a strategy selection unit.
[0068] The state space definition unit is responsible for defining the state space of reinforcement learning. The state space describes the complete state information of the energy system at a given moment and is the basis for the agent's decision-making. In one embodiment of this invention, the state variables include an energy state feature vector, the current moment, the remaining budget, and historical energy consumption. The energy state feature vector is generated by a multi-source data fusion module and includes multi-dimensional information such as energy consumption, business operations, environmental parameters, and electricity prices. The current moment reflects information in the time dimension; energy demand and electricity prices differ significantly across different time periods. The remaining budget represents the amount of funds available for energy procurement in the current period, reflecting cost constraints. Historical energy consumption records energy usage over a past period, reflecting trends and cyclical patterns in energy consumption.
[0069] State vector At any moment Defined as:
[0070] ,
[0071] in, For a moment The state vector, For a moment The energy state feature vector, The time characteristics of the current moment include hour, weekday, and month. For a moment The remaining budget, The historical energy consumption vector contains energy consumption statistics for the past 24 hours, the past 7 days, and the past 30 days. The dimension of the state vector is typically 50 to 100, with the specific dimension determined based on the complexity of the enterprise's energy system and the richness of the data.
[0072] The action space definition unit is responsible for defining the action space of reinforcement learning. The action space describes all possible actions that the agent can take. In the energy allocation optimization scenario, actions include the usage ratio of each energy type, the timing of procurement, and the quantity procured. In one embodiment of the present invention, action variables include the electricity usage ratio, gas usage ratio, heat usage ratio, energy storage discharge power, energy procurement quantity, and procurement time. The electricity usage ratio determines the ratio of grid power supply to self-generated power; the gas usage ratio determines the operating power of the gas boiler; the heat usage ratio determines the ratio of centralized heating to electric heating; the energy storage discharge power determines the discharge strategy of the energy storage system; the energy procurement quantity determines the amount of electricity purchased in the electricity market; and the procurement time determines the specific time point of electricity purchase.
[0073] Action vectors At any moment Defined as:
[0074] ,
[0075] in, For a moment The action vector, The electricity usage ratio is set to a value between 0 and 1. The gas usage ratio is set to a value between 0 and 1. The heat usage ratio ranges from 0 to 1. The energy storage discharge power ranges from 0 to the rated power of the energy storage system. The energy purchase quantity ranges from 0 to the maximum purchase quantity. The procurement time range is all time points within the scheduling cycle. To simplify the action space, the continuous action space can be discretized into a finite number of action options, such as discretizing the power usage ratio into five levels: 0%, 25%, 50%, 75%, and 100%.
[0076] The reward function definition unit is responsible for defining the reward function for reinforcement learning. The reward function is the learning objective of the agent, used to evaluate the quality of a state-action pair. In energy allocation optimization scenarios, the reward function needs to comprehensively consider energy cost, energy efficiency indicators, and carbon emissions. In one embodiment of the invention, immediate reward... Defined as:
[0077] ,
[0078] in, For a moment Instant rewards For a moment Energy costs, For a moment Energy efficiency indicators Carbon emissions , , and are weight coefficients. The reward function adopts a negative form, and the optimization goal is to minimize cost, maximize energy efficiency and minimize emissions. The weight coefficients reflect the relative importance of different goals, and in the preferred embodiment, is set to 0.6, is set to 0.2, is set to 0.2, emphasizing cost optimization while taking into account energy efficiency and environmental protection.
[0079] Energy cost includes electricity cost, gas cost, heat cost and demand charge. Electricity cost is calculated according to real-time electricity price and electricity consumption, gas cost is calculated according to gas unit price and gas consumption, heat cost is calculated according to heat price and heat consumption, and demand charge is calculated according to maximum demand and demand price. Energy efficiency is defined as the ratio of useful energy output to total energy input, reflecting energy utilization efficiency. Carbon emissions are calculated according to the emission factor and usage of different energy types, with the carbon emission factor of electricity being about 0.6 kgCO / kWh, and the carbon emission factor of gas being about 2.0 kgCO / m .
[0080] Value network unit and target network unit are the core components of deep Q network algorithm. The value network is used to fit the state-action value function , i.e. the expected cumulative reward that can be obtained by taking action in state . The target network is used to generate target Q values during training, and the training process is stabilized by fixing the target network parameters for a period of time.
[0081] In one embodiment of the present application, the value network adopts a deep neural network structure, including an input layer, multiple hidden layers and an output layer. The input layer receives the state vector , the hidden layer adopts a fully connected layer and an activation function for nonlinear transformation, and the output layer outputs the Q value of each action. The network structure is designed as follows: the input layer dimension is the state vector dimension, the first hidden layer contains 128 neurons, the second hidden layer contains 256 neurons, the third hidden layer contains 128 neurons, and the output layer dimension is the action space size. The hidden layer activation function adopts ReLU function, and the output layer does not use activation function to support full range output of Q value.
[0082] The parameters of the value network are denoted as , and for the input state ,, the value network outputs a Q-value vector of each action:
[0083] ,
[0084] wherein, is the Q-value vector output by the value network, is the Q-value of taking action in state , is the size of the action space, is the parameter of the value network.
[0085] The target network has the same structure as the value network, and the parameter is denoted as . The parameters of the target network are not updated by gradient descent, but are periodically copied from the value network. In an embodiment of the present application, the parameter copying operation is performed every 100 training steps:
[0086] ,
[0087] wherein, is the parameter of the target network, is the parameter of the value network, represents the assignment operation. By delaying the update of the target network parameters, the fluctuation of the target Q-value in the training process can be reduced, and the convergence stability of the algorithm can be improved.
[0088] The experience replay unit is used to store the experience data generated by the interaction between the agent and the environment, and to randomly sample training samples from the experience pool. The experience replay mechanism can break the temporal correlation between data and improve the sample utilization efficiency, and is an important part of the deep Q network algorithm.
[0089] In an embodiment of the present application, the capacity of the experience replay pool is set to 10,000 experience samples. Each experience sample contains a five-tuple , wherein is the current state, is the action taken, is the immediate reward obtained, is the next state, is a termination flag indicating whether the task is completed. After the agent interacts with the environment at each time step, the generated experience sample is stored in the replay pool. When the replay pool is full, the oldest experience sample is deleted using a first-in, first-out strategy.
[0090] During training, a batch of experience samples is randomly sampled from the experience replay pool, and the batch size is set to 64. For each experience sample in the batch, the target Q-value is calculated as:
[0091] ,
[0092] wherein, Qtargetis the target Q-value, Rtis the immediate reward, γ is the discount factor, denotes the next state, all possible actions, the action with the maximum Q-value, Qtargetis the Q-value predicted by the target network for the next state and action, γ is the discount factor. The discount factor is used to balance the immediate reward and long-term return, and is usually set to 0.95-0.99. In a preferred embodiment, is set to 0.97, which can focus on long-term returns while avoiding excessive discounting.
[0093] The loss function of the value network is defined as the mean squared error between the predicted Q-value and the target Q-value:
[0094] ,
[0095] wherein, L is the loss function, B is the batch size, Qtargetis the predicted Q-value of the value network for the i-th sample, Qtargetis the target Q-value of the i-th sample, Qtargetis the target Q-value of the i-th sample, Qtargetis the target Q-value of the i-th sample, Qtargetis the target Q-value of the i-th sample,
[0096] The value network parameters are updated using the gradient descent algorithm:
[0097] ,
[0098] wherein, Qtargetis the target Q-value of the i-th sample, η is the learning rate, Qtargetis the target Q-value of the i-th sample, Qtargetis the target Q-value of the i-th sample, Qtargetis the target Q-value of the i-th sample, Qtargetis the target Q-value of the i-th sample,
[0099] The policy selection unit selects the optimal action based on the trained value network. In the training phase, the following strategies are used: - The greedy policy balances exploration and exploitation. - The greedy policy selects the action with the maximum Q-value with a probability randomly select an action to explore with probability select the action with the maximum current Q value:
[0100] ,
[0101] wherein, is the selected action at time , is the exploration rate, denotes the action that maximizes the Q value, is the Q value prediction of the value network for state and action , is the value network parameter. The exploration rate is adopted with a decreasing strategy, set to 1.0 at the beginning of training to maintain sufficient exploration, gradually decaying to 0.01 as training progresses, and the decay rate is set to 0.995, that is, after each training step .
[0102] In the test phase, no longer need to explore, directly select the action with the maximum Q value:
[0103] ,
[0104] wherein, is the optimal action selected at time , denotes the action that maximizes the Q value, is the Q value prediction of the value network for state and action , is the value network parameter obtained by training.
[0105] The energy configuration strategy output by the reinforcement learning optimization module includes the use ratio of each energy type and the procurement scheme. In an embodiment of the present application, for an enterprise including three types of energy, electricity, gas and heat, the output example of the optimization strategy is: in the peak period (08:00-11:00, 18:00-21:00), preferentially use the discharge of the energy storage system and self-generation, reduce the electricity grid purchase ratio to 30%, the gas use ratio is set to 40%, and the central heating ratio is set to 60%; in the flat period (07:00-08:00, 11:00-18:00, 21:00-23:00), the electricity grid purchase ratio is increased to 60%, the gas use ratio is set to 50%, and the central heating ratio is set to 70%; in the valley period (23:00-07:00), make full use of low-price electricity, the electricity grid purchase ratio is increased to 90%, and at the same time, the energy storage system is charged, the gas use ratio is reduced to 20%, and the central heating ratio is set to 80%. The strategy makes full use of the peak-valley electricity price difference, and realizes the optimization of energy cost.
[0106] The reinforcement learning optimization method of the present application has stronger adaptive ability and optimization performance compared with traditional rule-based or model-based predictive control methods. Experiments show that this method can reduce energy cost by 18%, improve comprehensive energy efficiency by 15%, and reduce carbon emissions by 12% in a complex dynamic environment. Through online learning and continuous optimization, the system can continuously adapt to changes in the energy system and fluctuations in the market environment, maintaining long-term stable optimization results.
[0107] Reference Figure 4 The closed-loop feedback adjustment module 4 evaluates the energy configuration effect in real time, generates feedback adjustment parameters according to the actual energy efficiency deviation, and transmits the adjustment parameters to the front-end module to realize closed-loop feedback and system adaptation. This module includes an actual energy efficiency collection unit, a deviation calculation unit, an adjustment intensity determination unit, a feedback parameter generation unit, and a parameter transmission unit.
[0108] The actual energy efficiency collection unit is responsible for collecting actual energy efficiency data after the execution of the energy configuration strategy. Actual energy efficiency data includes actual energy cost, actual energy consumption, actual carbon emissions, device operating efficiency, and user comfort, etc. In an embodiment of the present application, actual energy efficiency data is automatically obtained by interfacing with the energy metering system, the financial system, and the environmental monitoring system. Actual energy cost is extracted from the energy bill of the financial system, including electricity, gas, and heat fees, etc. Actual energy consumption is collected from the smart meters of the energy metering system, including electricity meters, gas meters, and heat meters, with a precision of 0.1 kWh, 0.01 m and 0.1 GJ. Actual carbon emissions are calculated based on actual energy consumption and emission factors. Device operating efficiency is obtained through the device monitoring system, reflecting the actual working state of the device. User comfort is obtained through indoor environmental monitoring and user feedback, including indoor temperature, humidity, and satisfaction survey, etc.
[0109] The deviation calculation unit is used to calculate the deviation of actual energy efficiency indicators from expected energy efficiency indicators. Expected energy efficiency indicators are the energy efficiency targets predicted by the reinforcement learning optimization module based on the energy configuration strategy, and actual energy efficiency indicators are the actual performance after the execution of the strategy. Deviation reflects the difference between prediction and reality, which is the key signal for closed-loop feedback.
[0110] In an embodiment of the present application, three types of deviation indicators are defined: cost deviation rate, energy consumption deviation rate, and emission deviation rate. The cost deviation rate is defined as:
[0111] ,
[0112] wherein, is the cost deviation rate, is the actual energy cost, The energy consumption deviation rate is defined as:
[0113] The energy consumption deviation rate is defined as: The energy consumption deviation rate is defined as:
[0114] The energy consumption deviation rate is defined as:
[0115] The energy consumption deviation rate is defined as: The energy consumption deviation rate is defined as: The energy consumption deviation rate is defined as: The energy consumption deviation rate is defined as: The energy consumption deviation rate is defined as:
[0116] The emission deviation rate is defined as: The emission deviation rate is defined as: The emission deviation rate is defined as:
[0117] The emission deviation rate is defined as: The emission deviation rate is defined as: The emission deviation rate is defined as:
[0118] The emission deviation rate is defined as: The emission deviation rate is defined as: The emission deviation rate is defined as: The emission deviation rate is defined as: The emission deviation rate is defined as:
[0119] The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations:
[0120] The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations:
[0121] The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations:
[0122] The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations:
[0123] The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations:
[0124] The comprehensive deviation index is defined as the weighted sum of the three types of deviations: The comprehensive deviation index is defined as the weighted sum of the three types of deviations: This is a comprehensive deviation index. The adjustment intensity ranges from 0 to 1, with a larger value indicating a stronger adjustment. When the comprehensive deviation is less than 5%, the forecast is considered to be basically consistent with the actual situation, and a fine-tuning strategy is adopted. When the comprehensive deviation is between 5% and 10%, a mild adjustment is adopted. When the comprehensive deviation is between 10% and 20%, a moderate adjustment is adopted. When the comprehensive deviation exceeds 20%, the forecast is considered to have a significant difference from the actual situation, and a large-scale adjustment strategy is adopted.
[0125] The feedback parameter generation unit generates feedback adjustment parameters based on the bias and adjustment intensity. These parameters include the update step size of the digital twin model, the adjustment coefficients for the data fusion weights, and the learning rate adjustment coefficients for reinforcement learning.
[0126] Digital twin model update step size adjustment parameters Defined as:
[0127] ,
[0128] in, To update the step size for the adjusted digital twin model, For the initial update step size, To adjust the intensity, This is a sign function of the energy consumption deviation rate. When actual energy consumption is higher than expected, the model update step size is increased to accelerate the adjustment of model parameters. When actual energy consumption is lower than expected, the model update step size is decreased to maintain model stability.
[0129] Adjustment coefficient for data fusion weights Defined as:
[0130] ,
[0131] in, For data fusion weight adjustment coefficients, To adjust the intensity, This represents the overall deviation for the current cycle. This represents the overall deviation from the previous period. When the overall deviation continues to increase, the adjustment coefficient is greater than 1, increasing the data fusion weight and strengthening the role of the data source. When the overall deviation decreases, the adjustment coefficient is less than 1, appropriately reducing the data fusion weight. The update formula for the data fusion weight is:
[0132] ,
[0133] in, For the updated data source The fusion weight, The fusion weights before the update To adjust the coefficients, the updated weights need to be re-normalized.
[0134] learning rate adjustment coefficient of reinforcement learning is defined as:
[0135]
[0136] wherein, is a learning rate adjustment coefficient, is an adjustment intensity, is a cost deviation rate. When the cost deviation is large, the learning rate is increased to speed up the strategy learning speed. The update formula of the learning rate of reinforcement learning is:
[0137]
[0138] wherein, is an updated learning rate, is a learning rate before updating, is an adjustment coefficient. To avoid the learning rate being too large to cause unstable training, the upper limit of the learning rate is set to 0.001, that is, .
[0139] The parameter transmission unit is responsible for transmitting the generated feedback adjustment parameters to the corresponding front-end module. The parameter transmission is realized through the interface between the modules, ensuring that the adjustment parameters can take effect in time. The update step adjustment parameter of the digital twin model is transmitted to the digital twin modeling module 1, which is used to adjust the speed and frequency of model parameter updating. The adjustment coefficient of the data fusion weight is transmitted to the multi-source data fusion module 2, which is used to dynamically adjust the fusion weight of each data source. The learning rate adjustment coefficient of reinforcement learning is transmitted to the reinforcement learning optimization module 3, which is used to adjust the learning speed of the value network.
[0140] Through the closed-loop feedback adjustment mechanism, the present application can dynamically optimize the system parameters and strategies according to the actual execution effect, realize the self-adaptation and continuous improvement of the system. Experiments show that after introducing the closed-loop feedback, the long-term stability of the system is improved by 35%, the prediction accuracy is improved by 15%, and the optimization effect is improved by 20%. The closed-loop feedback enables the system to have self-learning and self-correction capabilities, and can maintain high performance in long-term operation.
[0141] The deep coupling closed-loop collaborative system constructed by the present application organically combines the digital twin modeling module 1, the multi-source data fusion module 2, the reinforcement learning optimization module 3 and the closed-loop feedback adjustment module 4 to form a complete closed-loop structure. The core feature of the system lies in the deep coupling, closed-loop feedback and collaborative effect between the modules.
[0142] The deep coupling between modules is reflected in the close connection of data flow, control flow and feedback flow. In terms of data flow, the virtual energy model state generated by the digital twin modeling module 1 is one of the important inputs of the multi-source data fusion module 2, and the energy state feature vector generated by the multi-source data fusion module 2 is the state input of the reinforcement learning optimization module 3. The energy configuration strategy output by the reinforcement learning optimization module 3 guides the operation of the actual system and affects the model update of the digital twin modeling module 1. In terms of control flow, the closed-loop feedback adjustment module 4 sends adjustment instructions to the digital twin modeling module 1, the multi-source data fusion module 2 and the reinforcement learning optimization module 3 according to the actual energy efficiency evaluation results, realizing the unified coordination of the whole system. In terms of feedback flow, the optimization results of the reinforcement learning optimization module 3 reversely affect the model parameters of the digital twin modeling module 1 and the fusion weights of the multi-source data fusion module 2, forming a backward feedback path from back to front.
[0143] The closed-loop feedback is reflected in the complete forward transmission → performance evaluation → reverse feedback → parameter adjustment cycle. In the forward transmission process, data flows from the digital twin modeling module 1 to the multi-source data fusion module 2, and then to the reinforcement learning optimization module 3, and finally outputs the energy configuration strategy and executes. In the performance evaluation process, the closed-loop feedback adjustment module 4 monitors and evaluates the strategy execution effect in real time, calculates the deviation between the actual energy efficiency and the expected energy efficiency. In the reverse feedback process, adjustment parameters are generated according to the deviation and transmitted to the front-end module. In the parameter adjustment process, each module updates its parameters and strategies according to the received adjustment parameters. This closed-loop process runs continuously, enabling the system to continuously optimize and improve itself.
[0144] The synergistic effect is reflected in the mutual promotion and superimposed efficiency between modules. The high-precision energy system model provided by the digital twin modeling module 1 enables the multi-source data fusion module 2 to more accurately evaluate the reliability and consistency of different data sources, thereby optimizing the fusion weights. The high-quality energy state feature vector generated by the multi-source data fusion module 2 provides rich decision-making information for the reinforcement learning optimization module 3, speeding up the convergence speed of policy learning. The optimal strategy learned by the reinforcement learning optimization module 3 can guide the actual system to run more efficiently, reducing energy costs and carbon emissions. The continuous optimization of the closed-loop feedback adjustment module 4 makes the digital twin model more accurate, the data fusion more reasonable, and the reinforcement learning more effective, realizing the spiral rise of the performance of the whole system.
[0145] The quantitative evaluation of the synergistic effect is performed by comparing the performance differences between single-module operation and multi-module synergistic operation. In one embodiment of the present application, the use of digital twin modeling alone can improve energy management accuracy by 12%, the use of multi-source data fusion alone can improve prediction accuracy by 10%, the use of reinforcement learning optimization alone can reduce energy costs by 15%, and the use of closed-loop feedback adjustment alone can improve system stability by 8%. After the four modules are combined into a deep-coupled closed-loop synergistic system, the energy management accuracy is improved by 40%, the prediction accuracy is improved by 25%, the energy cost is reduced by 18%, the system stability is improved by 35%, and the overall energy efficiency is improved by 20%. The overall performance of the system is much higher than the sum of the individual contributions of each module, fully demonstrating the nonlinear synergistic effect of 1+1>2.
[0146] The running period of the deep-coupled closed-loop synergistic system is usually set to 15 minutes to 1 hour. In each running period, the digital twin modeling module 1 updates the energy system state in real time, the multi-source data fusion module 2 fuses multi-dimensional data, the reinforcement learning optimization module 3 generates energy configuration strategies, the system executes energy scheduling according to the strategies, and the closed-loop feedback adjustment module 4 evaluates the execution effect and generates adjustment parameters. The adjustment parameters are transmitted to each pre-module and the parameters are updated. Through periodic closed-loop iteration, the system continuously learns and optimizes, gradually approaching the optimal operating state.
[0147] In one embodiment of the present application, a 6-month practical application test was conducted on the energy management system of a certain manufacturing enterprise. The enterprise includes production workshops, office buildings, warehouses, and canteens, and is equipped with three energy supplies: electricity, gas, and heat. The total installed capacity is 5 MW, and the annual energy consumption is about 10 million yuan. During the test, the deep-coupled closed-loop synergistic energy management method of the present application was used for energy optimization.
[0148] The test results show that, compared with traditional energy management methods, the present application method has achieved significant improvement in multiple indicators. In terms of energy cost, 18.5% of energy fees are saved in 6 months, equivalent to saving 1.85 million yuan. In terms of energy consumption, the overall energy efficiency is improved by 20.3%, and the unit energy consumption is reduced by 16.8%. In terms of carbon emissions, the cumulative carbon emissions are reduced by 18.2%, equivalent to reducing 450 tons of CO . In terms of system reliability, the energy configuration prediction accuracy reaches 92%, the average deviation between actual energy efficiency and expected energy efficiency is less than 5%, and the system operation stability is improved by 35%. In terms of user satisfaction, the indoor environmental comfort score is improved from 75 to 88, and there is no production interruption or environmental discomfort caused by energy optimization.
[0149] The test also verifies the adaptive ability of the system. During the test, the enterprise experienced various scene changes such as production off-season to peak season conversion, equipment upgrading and reconstruction, seasonal climate change, and electricity price policy adjustment. The deep coupling closed-loop collaborative system of the present application can quickly adapt to these changes, automatically adjust the optimization strategy, and maintain the persistence of the optimization effect. For example, when the production peak season load increases by 30%, the system completes the strategy adaptation within 3 days, and the energy cost still remains at the optimized level. During the summer air conditioning load peak, the system successfully copes with the load challenge by optimizing the air conditioning operation strategy and using the energy storage system to cut the peak and fill the valley.
[0150] Referring to Figure 5 The present application also provides an energy management system based on big data. The system includes a memory and a processor, the memory stores computer executable instructions, and the computer executable instructions are executed by the processor to realize the above-mentioned energy management method based on big data.
[0151] The memory can be any type of non-volatile or volatile storage medium such as read-only memory (ROM), random access memory (RAM), flash memory, hard disk, solid state disk (SSD), etc. The memory stores program codes, configuration parameters, historical data, and model parameters of the energy management system. The program codes include the implementation of digital twin modeling algorithm, multi-source data fusion algorithm, reinforcement learning optimization algorithm, and closed-loop feedback adjustment algorithm. The configuration parameters include sampling frequency, time granularity, fusion weight, reward function coefficient, network structure parameter, and learning rate, etc. The historical data includes energy consumption history, business operation history, environmental parameter history, and electricity price history, etc. The model parameters include digital twin model parameters, value network parameters, and target network parameters, etc.
[0152] The processor can be any type of computing unit such as central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), field programmable gate array (FPGA), application specific integrated circuit (ASIC), etc. The processor executes the computer executable instructions stored in the memory to realize the functions of the energy management method. For computationally intensive tasks such as training and inference of deep neural networks, it is preferred to use GPU for accelerated computation, which can increase the computing speed by more than 10 times.
[0153] The system also includes peripheral interfaces such as a communication interface, a data acquisition interface, and a control output interface. The communication interface is used for data exchange with external systems such as enterprise information systems ERP, manufacturing execution systems MES, human resource systems HR, power trading platforms, and weather data platforms. The data acquisition interface is used to connect monitoring devices such as smart meters, gas meters, heat meters, flow meters, temperature sensors, and pressure sensors to collect energy and environmental data in real time. The control output interface is used to send control instructions to energy devices to implement energy configuration strategies. Interface protocols support common industrial communication protocols such as Modbus, OPC UA, MQTT, and HTTP.
[0154] The system can be deployed on a local server, an edge computing device, or a cloud computing platform. For small and medium-sized enterprises, a local deployment method can be used, and an industrial computer or a server can be used to carry the system. For large enterprises or multi-park enterprises, a cloud-edge collaborative deployment method can be used, with the edge end responsible for data acquisition and real-time control, and the cloud end responsible for model training and strategy optimization. Cloud computing platforms can provide powerful computing capabilities and massive storage spaces, supporting more complex algorithm models and larger-scale data processing.
[0155] The system also includes a human-computer interaction interface to provide managers with an intuitive monitoring and operation interface. The human-computer interaction interface is developed using Web technology and supports PC and mobile access. Interface functions include real-time monitoring, historical playback, report analysis, strategy configuration, and alarm management. The real-time monitoring interface displays the current operating status of the energy system, including real-time power, energy consumption accumulation, energy costs, and carbon emissions of each device, as well as three-dimensional visualization of the digital twin model. The historical playback interface supports querying and playing back historical data to analyze the temporal regularity and spatial distribution of energy consumption. The report analysis interface provides various report forms such as daily, weekly, monthly, and annual reports to calculate key indicators such as energy consumption, cost savings, and carbon emission reductions. The strategy configuration interface allows managers to set optimization targets, constraint conditions, and operating modes. The alarm management interface displays system alarm information in real time, including energy consumption anomalies, device failures, cost overruns, and carbon emissions exceeding standards, and supports alarm pushing and handling records.
[0156] Through the above hardware structure and software implementation, the energy management system of the present application can provide a one-stop intelligent energy management solution for enterprises, realize real-time monitoring, accurate prediction, and optimal configuration of energy consumption, help enterprises reduce energy costs, improve energy efficiency, reduce carbon emissions, and provide strong support for the green development and sustainable development of enterprises.
Claims
1. A big data-based energy management method, applied to an enterprise energy management system, characterized in that: include: By constructing a digital twin model of the enterprise's energy system through the digital twin modeling module, the parameters, topology, and operating status of the enterprise's energy equipment are collected, and a real-time mapping relationship between the physical energy system and the virtual energy model is established. Multi-source data fusion modules collect and integrate multi-dimensional data to obtain energy consumption data, business operation data, environmental parameter data, and real-time electricity price data. The energy consumption data includes electricity consumption, gas consumption, heat consumption, and water resource consumption. The business operation data includes production plans, business volume, equipment utilization rate, and personnel mobility. The environmental parameter data includes ambient temperature, humidity, light intensity, and air quality. The real-time electricity price data includes peak-valley electricity prices, real-time market electricity prices, and tiered electricity prices. The time alignment and spatial consistency of each data source are calculated. The time alignment... Defined as: , in, For data source Time alignment For data source The total number of data points within the statistical period. For data source No. The time difference between the timestamp of each data point and the nearest standard time grid point Standard time granularity; Space consistency is calculated using the mutual information method. Defined as: , in, For data source and data source Spatial consistency between them For data source variables With data source variables Mutual information between them For variables Information entropy For variables Information entropy This indicates taking the minimum value; Based on the adaptive fusion weight algorithm, fusion weights are assigned to each data source. : , in, For data source The fusion weight, For data source Time alignment For data source Other data sources Spatial consistency between them For data source Quality rating Total number of data sources , and For the weighting coefficients to satisfy ; The fusion weights are normalized. , in, The normalized fusion weights For the original fusion weights, Total number of data sources; Based on normalized fusion weights, a unified energy state feature vector is generated. : , in, This is the fused energy state feature vector. For data source Normalized fusion weights, For data source The feature vector after standardization Total number of data sources; The energy allocation optimization module optimizes energy configuration based on the energy state feature vector through reinforcement learning, constructing a Markov decision process including a state space, action space, and reward function. A deep Q-network algorithm is then used to learn the optimal energy allocation strategy, outputting the usage ratio and procurement plan for each energy type. The target Q-value of the deep Q-network algorithm is... The calculation formula is: , in, For the target Q value, For instant rewards, As a discount factor, Indicates the next state All possible actions The action with the largest Q value, For the target network to predict the Q-values of the next state and action, For target network parameters; The closed-loop feedback adjustment module performs real-time evaluation of energy configuration effectiveness, collecting actual energy efficiency data after energy configuration execution, including actual energy costs, actual energy consumption, and actual carbon emissions; it also calculates the deviation between actual and expected energy efficiency indicators, including the cost deviation rate. Energy consumption deviation rate and emission deviation rate The adjustment intensity is determined based on the deviation; if the deviation exceeds a preset threshold, the adjustment intensity is increased; feedback adjustment parameters are generated, including the update step size of the digital twin model, the adjustment coefficient of the data fusion weights, and the learning rate adjustment coefficient of reinforcement learning; the feedback adjustment parameters are passed to the digital twin modeling module to update the model parameters, to the multi-source data fusion module to dynamically adjust the fusion weights of each data source, and to the reinforcement learning optimization module to adjust the exploration rate and learning rate to improve the policy learning effect; The digital twin modeling module, multi-source data fusion module, reinforcement learning optimization module, and closed-loop feedback adjustment module form a deeply coupled closed-loop collaborative system. The output of the closed-loop feedback adjustment module directly serves as the key input parameter of the digital twin modeling module. The energy configuration result of the reinforcement learning optimization module inversely affects the model update frequency of the digital twin modeling module and the fusion weight of the multi-source data fusion module, achieving mutual promotion and synergistic effect among the modules. The closed-loop collaborative system achieves the following synergistic effects: the optimization result of the reinforcement learning optimization module inversely affects the digital twin modeling module, adjusting the sensitive parameter weights of the digital twin model according to the actual execution effect of energy configuration; the closed-loop feedback adjustment module simultaneously adjusts the fusion weight of the multi-source data fusion module and the learning parameters of the reinforcement learning optimization module, achieving joint optimization of multi-module parameters; through deep coupling among modules, triple synergy of data flow, control flow, and feedback flow is achieved, ensuring a non-linear improvement in the overall system performance.
2. The energy management method based on big data according to claim 1, characterized in that, The digital twin modeling module constructs a digital twin model of the enterprise's energy system, including: Collect physical equipment parameters of the enterprise's energy system, including equipment type, rated power, operating efficiency, and equipment status; Obtain the network topology of the energy system and determine the connection relationships and energy flow paths between various energy devices; Real-time acquisition of operating status parameters of each energy device, including real-time power, energy consumption, operating temperature and fault signals; Based on the physical device parameters, network topology, and operating status parameters, a digital twin mapping relationship is established to generate a virtual energy model; Based on the deviation between the virtual energy model and the physical system, the model parameters are dynamically adjusted to maintain the synchronization between the digital twin model and the physical system.
3. The energy management method based on big data according to claim 1, characterized in that, The reinforcement learning optimization module optimizes energy allocation based on the energy state feature vector, including: Define a state space, with the energy state feature vector, current time, remaining budget, and historical energy consumption as state variables; Define the action space, and use the usage ratio of each energy type, the timing of procurement, and the quantity of procurement as action variables; Define a reward function to calculate an instant reward value based on energy costs, energy efficiency indicators, and carbon emissions; A deep Q-network algorithm is used to construct a value network and a target network, and a state-action value function is fitted through a neural network. Historical interaction data is stored using an experience replay mechanism, and training samples are randomly sampled from the experience pool. The value network parameters are updated using the gradient descent algorithm to minimize the mean square error between the predicted Q value and the target Q value. The action is selected based on the ε-greedy strategy, balancing exploration and exploitation, and outputting the optimal energy allocation strategy.
4. The energy management method based on big data according to claim 1, characterized in that, The feedback adjustment parameters are generated using an adaptive adjustment mechanism. The adjustment direction and magnitude are determined based on historical deviation trends. If the deviation continues to increase for several consecutive periods, a large-scale adjustment strategy is adopted. If the deviation fluctuation is stable, a fine-tuning strategy is adopted.
5. The energy management method based on big data according to claim 1, characterized in that, It also includes an early warning mechanism that triggers an early warning signal and automatically adjusts the energy configuration strategy when abnormal energy consumption, equipment failure risk, or cost overrun risk is detected.
6. A big data-based energy management system, comprising a memory and a processor, wherein the memory stores computer-executable instructions, characterized in that, When the computer-executable instructions are executed by the processor, they implement the energy management method based on big data according to any one of claims 1 to 5.
Citation Information
Patent Citations
Big Data-Based Energy Management Methods and Devices
CN115994628B
Energy optimization type motion control system for unmanned factory
CN120630646A
Intelligent operation and maintenance method and system for energy management platform
CN121072843A