Electric energy control method and device, electronic equipment and readable storage medium
Optimizing the power control strategy through the target digital twin model and Q learning algorithm, the problem of low power control accuracy in traditional hotels is solved, and efficient, safe distribution of electricity and reduced waste are achieved.
Patent Information
- Application Number
- CN202510433675.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional hotel power control methods have low accuracy and cannot perform intelligent and accurate power control based on actual occupancy and equipment operating status, resulting in serious waste of electricity.
The target digital twin model is used to predict electricity consumption, adjust the initial electricity consumption control strategy based on the difference, and obtain the target electricity consumption control strategy through Q learning algorithm optimization, combining real-time monitoring and safety protection mechanisms to ensure the rationality and effectiveness of electricity distribution.
It improves the accuracy of power control, reduces power waste, ensures the safety and stability of the power system, and realizes the rational utilization and efficient allocation of power.
Smart Images

Figure CN120338548A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of electric power, and more particularly, relates to a method and device for controlling electric energy, an electronic device, and a readable storage medium. Background Art
[0002] With the development of the hotel industry, the problem of energy consumption has become increasingly prominent. As a comprehensive service place, a hotel has a large number of electrical devices, such as lighting systems, air-conditioning systems, elevator systems, in-room appliances, etc. These devices consume a large amount of electric energy during operation, resulting in an increase in the hotel's operating costs. At the same time, unreasonable use of electric energy also causes great pressure on the environment. Traditional hotel electric energy control methods have low accuracy, often only operating by manually turning on and off devices at regular intervals or through simple time control programs, and cannot perform intelligent and accurate electric energy regulation according to factors such as the actual occupancy situation of the hotel, real-time electricity demand, and device operating status, resulting in serious problems of electric energy waste. Summary of the Invention
[0003] The purpose of the present disclosure is to provide a method and device for controlling electric energy, an electronic device, and a readable storage medium to improve the accuracy of electric energy control, thereby reducing the occurrence of electric energy waste.
[0004] In the first aspect of the embodiments of the present disclosure, a method for controlling electric energy is provided, including: Predicting the first power consumption of a power supply main circuit based on a target digital twin model to obtain the second power consumption of the power supply main circuit, where the first power consumption is the power consumption of a target area within a current preset time period, and the second power consumption is the predicted power consumption of the target area within a future target time period; Determining an initial power consumption control strategy for the power supply main circuit to distribute electric energy to multiple power supply branches based on a first difference, where the first difference is the difference between the second power consumption of the power supply main circuit and the rated power consumption; Adjusting the initial power consumption control strategy based on the Q-learning algorithm to obtain a target power consumption control strategy.
[0005] In the second aspect of the embodiments of the present disclosure, a device for controlling electric energy is provided, including: A power consumption prediction module, configured to predict the first power consumption of a power supply main circuit based on a target digital twin model to obtain the second power consumption of the power supply main circuit, where the first power consumption is the power consumption of a target area within a current preset time period, and the second power consumption is the predicted power consumption of the target area within a future target time period; An initial power consumption control module, configured to determine an initial power consumption control strategy for the power supply main circuit to distribute electric energy to multiple power supply branches based on a first difference, where the first difference is the difference between the second power consumption of the power supply main circuit and the rated power consumption; The target power consumption control module is used to adjust the initial power consumption control strategy based on the Q - learning algorithm to obtain the target power consumption control strategy.
[0006] In the third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above - mentioned power control method are implemented.
[0007] In the fourth aspect of the embodiments of the present disclosure, a computer - readable storage medium is provided. The computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above - mentioned power control method are implemented.
[0008] The beneficial effects of the power control method, device, electronic device, and readable storage medium provided by the embodiments of the present disclosure are as follows: By predicting the second power consumption, the present disclosure provides reliable data support for power distribution, which helps to achieve the rational use of electric energy and avoid waste. Then, the present disclosure flexibly adjusts the initial power consumption control strategy for the power supply trunk to distribute electric energy to multiple power supply branches through the first difference, ensuring the rationality and effectiveness of power distribution. At the same time, the Q - learning algorithm is introduced to intelligently adjust the initial power consumption control strategy to obtain the target power consumption control strategy, further improving the accuracy and adaptability of power control. Therefore, the present disclosure can improve the accuracy of power control, thereby reducing the occurrence of power waste. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0010] Figure 1 It is a schematic flowchart of the power control method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of the power control device provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of the electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0012] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.
[0013] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the power control method provided for an embodiment of the present disclosure. The method includes: S101: Predict the first power consumption of the power supply main circuit based on the target digital twin model to obtain the second power consumption of the power supply main circuit. The first power consumption is the power consumption of the target area within the current preset time period, and the second power consumption is the predicted power consumption of the target area within the future target time period.
[0014] In this embodiment, the target digital twin model is a digital simulation of the power system in the target area, which can be constructed by integrating information such as the physical structure, operation rules, and historical data of the power system. This model can simulate the operation state of the power system in a virtual environment and make predictions based on the input data (i.e., the first power consumption).
[0015] The target area can be a hotel area, and the power supply main circuit is a set of multiple power branches that supply power to the target area. The power of all electrical equipment in this area is transmitted through the power supply main circuit. The first power consumption is the actual power consumption of the target area within a preset time period, which can be 1 day, 1 week, or 1 month, etc., and can be directly measured by power metering equipment.
[0016] The second power consumption is the result obtained by predicting the future power consumption of the target area through the target digital twin model, which can reflect the possible future power consumption situation of the target area and is used for the control of subsequent electrical equipment. Among them, the future target time period is the next moment of the current preset time period. Both the first power consumption and the second power consumption can be power.
[0017] Specifically, the target digital twin model is a virtual model that is highly similar to the real power system in the target area. The first power consumption provides the actual operation result data of the power system at the current moment for this model. Combining with environmental characteristics, the target digital twin model can simulate the future operation status of the power system under the influence of various factors, that is, predict the second power consumption, providing a decision-making basis for power control.
[0018] S102: Determine the initial power consumption control strategy for the power supply trunk to distribute electric energy to multiple power supply branches based on the first difference. The first difference is the difference between the second power consumption of the power supply trunk and the rated power consumption.
[0019] In this embodiment, the rated power consumption is the upper limit of the power that the power supply trunk can continuously and stably supply under normal operating conditions. This value is determined by factors such as the specifications of the power supply line and the capacity of the power equipment. The rated power consumption can be the rated power.
[0020] The first difference is the difference between the second power consumption of the power supply trunk and the rated power consumption. This difference can reflect the relationship between the predicted power consumption and the rated power, and is the key basis for formulating the power consumption control strategy.
[0021] A power supply branch is a branch of the power supply trunk, that is, a line that branches out from the power supply trunk and can provide electric energy for different electrical equipment or power consumption areas. For example, in a hotel area, the power supply lines for different power consumption areas can be regarded as power supply branches branched out from the power supply trunk of the hotel area.
[0022] The initial power consumption control strategy is a preliminary plan for the power supply trunk to distribute electric energy to multiple power supply branches formulated based on the first difference. The distributed electric energy determines the number of electrical equipment used in the power consumption area.
[0023] Specifically, determining the initial power consumption control strategy for the power supply trunk to distribute electric energy to multiple power supply branches based on the first difference includes: If the first difference is less than the first difference threshold, the initial power consumption control strategy is to increase the number of electrical equipment on multiple power supply branches; If the first difference is greater than or equal to the first difference threshold, the initial power consumption control strategy is to reduce the number of electrical equipment on the power supply branches.
[0024] In this embodiment, the first difference threshold is a preset reference value, used as the boundary for judging whether the power supply situation is abundant or tense. When comparing the first difference with this threshold, the initial power consumption control strategy can be determined.
[0025] Electrical equipment is equipment connected to the power supply branch that consumes electric energy to achieve operation, and can be equipment such as washing machines, lighting systems, elevators, or air conditioners.
[0026] The initial power consumption control strategy is to adjust the number of electrical equipment on the power supply branch according to the comparison result of the first difference and the first difference threshold.
[0027] Specifically, if the first difference is less than the first difference threshold, it means that the predicted power consumption is relatively small compared to the rated power consumption, and the power supply is relatively abundant. At this time, the initial power consumption control strategy is to increase the number of electrical devices on multiple power supply branches to improve the utilization rate of electric energy. If the first difference is greater than or equal to the first difference threshold, it indicates that the predicted power consumption is close to or exceeds the rated power consumption, and the power supply is relatively tight. At this time, the initial power consumption control strategy is to reduce the number of electrical devices on the power supply branch to avoid overload and ensure the stable operation of the power supply system.
[0028] S103: Adjust the initial power consumption control strategy based on the Q-learning algorithm to obtain the target power consumption control strategy.
[0029] In this embodiment, the Q-learning algorithm is a model-free reinforcement learning algorithm. Reinforcement learning enables the agent to interact with the environment and continuously try different actions to obtain the maximum cumulative reward. The Q-learning algorithm achieves the above goal by learning a value function.
[0030] The target power consumption control strategy is a more reasonable power consumption control strategy that can achieve better performance in power consumption situations after being adjusted and optimized by the Q-learning algorithm for the initial power consumption control strategy.
[0031] Specifically, the Q-learning algorithm allows the agent to select the optimal action in different states through continuous attempts and learning to maximize the long-term cumulative reward. In the power control scenario, the agent interacts with the power system environment and adjusts the Q value according to the system feedback (reward) to find the power consumption control strategy that can optimize the system performance in various states.
[0032] It can be concluded from the above that the present disclosure provides reliable data support for power distribution by predicting the second power consumption, which helps to achieve the rational use of electric energy and avoid waste. Then, the present disclosure flexibly adjusts the initial power consumption control strategy for distributing electric energy from the power supply main circuit to multiple power supply branches through the first difference, ensuring the rationality and effectiveness of power distribution. At the same time, the Q-learning algorithm is introduced to intelligently adjust the initial power consumption control strategy to obtain the target power consumption control strategy, further improving the accuracy and adaptability of power control. Therefore, the present disclosure can improve the accuracy of power control, thereby reducing the occurrence of power waste.
[0033] In an embodiment of the present disclosure, the power control method further includes: Process the first power consumption parameters on multiple power supply branches to obtain a first monitoring result. The first power consumption parameters are the power consumption parameters of the electrical devices on each power supply branch; the first monitoring result includes normal and abnormal; When the first monitoring result is abnormal, output a safety protection instruction; the safety protection instruction is used to control the protection device of the corresponding power supply branch.
[0034] In this embodiment, the first electrical parameter is various parameters related to electricity consumption involved in electrical equipment on each power supply branch, such as current, voltage, power, electricity consumption, etc. The above parameters can reflect the operating state of the electrical equipment.
[0035] The first monitoring result is the result obtained after monitoring the first electrical parameters of electrical equipment on multiple power supply branches. This result is divided into two situations: "normal" and "abnormal", and is used to determine whether the operation of the power supply branch and the electrical equipment is in a normal state.
[0036] The safety protection instruction is an instruction issued when the first monitoring result is "abnormal". Its function is to control the protection device of the corresponding power supply branch to perform corresponding actions to ensure the safety of the power supply system and electrical equipment.
[0037] The protection device is a device installed on the power supply branch, such as a circuit breaker, a leakage protector, etc. It can take measures such as cutting off the circuit after receiving the safety protection instruction to prevent safety accidents such as equipment damage and fire caused by abnormal electricity consumption.
[0038] Specifically, processing the first electrical parameters on multiple power supply branches to obtain the first monitoring result includes: For each power supply branch: Determine multiple degrees of matching between the first electrical parameter and the standard electrical parameter; each standard electrical parameter corresponds to a monitoring result; Compare the multiple degrees of matching to determine the maximum degree of matching, and use the monitoring result corresponding to the maximum degree of matching as the monitoring result corresponding to the first electrical parameter; Based on the monitoring results corresponding to multiple first electrical parameters, obtain the monitoring result of each power supply branch.
[0039] Among them, the degree of matching can be calculated based on the Pearson correlation coefficient.
[0040] In this embodiment, the monitoring results corresponding to multiple first electrical parameters can be normal current, abnormal current, normal voltage, abnormal voltage, normal power, abnormal power, etc. Based on the above results, the monitoring result of each power supply branch can be obtained, that is, when the monitoring results corresponding to multiple first electrical parameters are all normal, the monitoring result of this power supply branch is normal; when the monitoring results corresponding to multiple first electrical parameters are at least one abnormal, the monitoring result of this power supply branch is abnormal. And when the monitoring result of at least one power supply branch is abnormal, the first monitoring result is abnormal; when the monitoring results of each power supply branch are all normal, the first monitoring result is normal.
[0041] As can be seen from the above, in this embodiment, by monitoring the first electrical parameters on multiple power supply branches and promptly responding to abnormal situations, the safety and stability of the power system are effectively improved. When abnormal power consumption is detected, this embodiment can quickly output a safety protection instruction to control the protection device of the corresponding power supply branch to act, prevent the spread of faults, and reduce losses. This real-time monitoring and rapid response mechanism not only ensures the safe operation of electrical equipment but also reduces the risk of safety accidents caused by electrical faults, providing a strong guarantee for the efficient and safe utilization of electric energy.
[0042] In an embodiment of the present disclosure, the power control method further includes: The target digital twin model includes a physical model and a target virtual model; The training process of the target digital twin model includes: Construct a physical model and a virtual model of the main power supply line based on historical power consumption data; there is a mapping relationship between the virtual model and the physical model; Iteratively optimize the parameters of the virtual model until the error between the output result of the virtual model and the output result of the physical model is less than the first error threshold, and obtain the first target parameters corresponding to the virtual model; Determine the target virtual model based on the first target parameters; Integrate the physical model and the target virtual model to determine the target digital twin model.
[0043] In this embodiment, using historical power consumption data and combining the physical characteristics of the main power supply line, such as electrical characteristics, line parameters, and equipment parameters, etc., to construct a physical model, that is, to mathematically describe the operation law of the main power supply line. At the same time, construct a virtual model based on historical power consumption data. The structure of the virtual model corresponds to the physical model, and there is a mapping relationship between them, that is, the variables in the virtual model correspond to the actual physical quantities in the physical model.
[0044] The target virtual model is a virtual model that, after being trained and optimized, corresponds to the physical model and can accurately simulate the operation state of the main power supply line. It and the physical model together constitute the target digital twin model.
[0045] The historical power consumption data is the power consumption-related data of the main power supply line and each power supply branch over a long period of time in the past, such as the power consumption at different time points, the on / off time of electrical equipment, and the change of power load, etc. The above data reflects the past operation situation of the power system.
[0046] The virtual model is a preliminarily constructed model that simulates the operation state of the main power supply line, and its parameters are continuously adjusted and optimized during the training process.
[0047] The first error threshold is a pre-set allowable error range, which is used to judge the closeness of the output results of the virtual model and the physical model. When the error between the two is less than this threshold, it is considered that the virtual model has achieved an acceptable accuracy. The first target parameter is the parameter value corresponding to the situation where the output result of the virtual model is less than the first error threshold after the virtual model is iteratively optimized. Moreover, in this embodiment, the number of times the error is continuously less than the first error threshold is greater than or equal to the first number threshold.
[0048] Specifically, in this embodiment, the virtual model is continuously optimized to enable it to simulate the behavior of the physical model as accurately as possible. The model foundation is constructed using historical power consumption data, and the virtual model parameters are adjusted by comparing the errors of the output results. After the virtual model meets certain accuracy requirements, it is integrated with the physical model to form the target digital twin model.
[0049] It can be concluded from the above that this embodiment uses historical power consumption data to construct the physical model and the virtual model, and iteratively optimizes the virtual model parameters to ensure a high degree of consistency between the virtual model and the physical model, thereby improving the accuracy of prediction and control. The application of the target digital twin model makes the power management more refined and efficient, helps to detect and solve potential problems in a timely manner, and optimize the power resource allocation.
[0050] In an embodiment of the present disclosure, the parameters of the virtual model are iteratively optimized, including: Determine the fitness function of the particle swarm optimization algorithm based on the dimensions of the particles in the particle swarm optimization algorithm. The dimensions of the particles in the particle swarm optimization algorithm include the learning rate, batch value, and neuron dropout rate; Determine the inertia weight of the particle swarm optimization algorithm based on the historical environmental characteristics of the power supply trunk; Iteratively optimize the parameters of the virtual model based on the fitness function and the inertia weight.
[0051] In this embodiment, the particle swarm optimization algorithm is used to optimize the parameters of the virtual model. A particle is the basic individual in the particle swarm optimization algorithm, and each particle represents a set of possible solutions, that is, a set of parameters of the virtual model. The dimensions of the particles include the learning rate, batch value, and neuron dropout rate, all of which are important parameters affecting the training effect of the virtual model. For example, the learning rate can affect the convergence speed of the virtual model, the batch value can affect the generalization ability of the virtual model, and the neuron dropout rate can prevent the virtual model from overfitting during training.
[0052] The fitness function is used to evaluate the quality of the solution represented by each particle. In the particle swarm optimization algorithm, the particles will move in the direction of a better fitness value to find the global optimal solution.
[0053] Specifically, the calculation formula of the fitness function is:
[0054] Among them, is the value of the fitness function, are all weight coefficients, and are greater than or equal to 0, , is the mean square error, is the learning rate, is the optimal learning rate, is the batch value, is the optimal batch value, is the neuron dropout rate, is the optimal neuron dropout rate. At the same time, the optimal learning rate, the optimal batch value, and the optimal neuron dropout rate are all determined through a large number of experiments.
[0055] The design logic of the above formula is as follows: The mean square error can directly reflect the difference between the output of the virtual model and the output of the physical model, and measure the prediction accuracy of the virtual model. The smaller this value is, the closer the prediction result of the virtual model is to the real situation. In the fitness function, this part is expected to be as small as possible to reflect the good performance of the model.
[0056] This is because the learning rate affects the speed and effect of model training. Taking the square of the reciprocal difference can make the influence of the learning rate on the fitness function smaller when it is close to the optimal value; while when the learning rate deviates greatly from the optimal value, it will have a greater influence, thus guiding the particles to search in the direction of the optimal learning rate.
[0057] This is because the size of the batch value will affect the stability and efficiency of model training. Using the form of the square of the logarithmic difference can measure the proportional relationship between the batch value and the optimal batch value, making the fitness function have a certain sensitivity and stability to the change of the batch value, neither overreacting to small deviations nor effectively distinguishing large deviations, and prompting the particles to find the appropriate batch value.
[0058] simply and intuitively reflects the degree of deviation of the current neuron dropout rate from the ideal value. The neuron dropout rate is used to prevent model overfitting. In this way, the fitness function can guide the particles to find the optimal neuron dropout rate to improve the generalization ability of the model.
[0059] In this embodiment, by performing weighted summation with the above parameters, the relationship between the mean square error and the deviation of each parameter from the optimal value can be balanced. According to the requirements of the actual problem, adjusting these weight coefficients can highlight the degree of attention to the prediction accuracy or parameter rationality of the model, so as to more effectively guide the particle swarm optimization algorithm to search for the optimal model parameter combination.
[0060] In the particle swarm optimization algorithm, the inertia weight controls the influence degree of the particle's previous velocity on the current velocity. An appropriate inertia weight can balance the global search and local search capabilities of the particle.
[0061] The historical environmental features are the information related to the past operating environment of the power supply main circuit, such as the electricity consumption load and weather conditions in different seasons and different time periods. The above information can reflect the operating rules and characteristics of the power supply main circuit.
[0062] Specifically, determining the inertia weight of the particle swarm optimization algorithm based on the historical environmental features of the power supply main circuit includes: Determine the initial inertia weight; Determine that the historical environmental feature of the power supply main circuit is the fluctuation value of the electricity consumption load; In response to the fluctuation value being greater than or equal to the first fluctuation threshold, control the initial inertia weight to increase with the third step size; In response to the fluctuation value being less than the first fluctuation threshold, control the initial inertia weight to increase with the fourth step size.
[0063] Among them, the initial inertia weight, the first fluctuation threshold, the third step size, and the fourth step size are all set according to experience. The fluctuation value of the electricity consumption load can be calculated based on the absolute value of the difference between the maximum fluctuation value and the minimum fluctuation value.
[0064] In this embodiment, the fitness function is used to evaluate the pros and cons of the virtual model parameters represented by each particle, the inertia weight is used to control the search direction and speed of the particle, and through continuous iteration, the particle gradually approaches the optimal solution, thereby realizing the optimization of the virtual model parameters.
[0065] It can be concluded from the above that in this embodiment, the particle swarm optimization algorithm is used to iteratively optimize the virtual model parameters, improving the tuning efficiency and accuracy of the model parameters. In this embodiment, by reasonably setting the particle dimension and determining the inertia weight based on the historical environmental features of the power supply main circuit, the close association between the optimization process and the actual application scenario is ensured. This refined parameter adjustment strategy not only improves the prediction accuracy of the virtual model but also enhances its generalization ability.
[0066] In an embodiment of the present disclosure, adjusting the initial electricity consumption control strategy to obtain the target electricity consumption control strategy based on the Q - learning algorithm includes: Determine the initial action space vector of the Q - learning algorithm corresponding to the initial electricity consumption control strategy; Determine multiple standard electricity consumption control strategies, and determine the action space vectors of the Q - learning algorithm corresponding to each standard electricity consumption control strategy; Determine the state space vector of the Q - learning algorithm based on the second electricity consumption; Determine the value function of the Q-learning algorithm based on the action space vector, state space vector, reward function, and penalty factor; Determine the target power consumption control strategy based on the initial action space vector and the feedback value of the value function.
[0067] In this embodiment, the initial action space vector is a set representation of the actions that the agent can take under the corresponding initial power consumption control strategy, and each action can change the power consumption control situation. For example, the initial action space vector can be [turn off the washing machine operation, turn off the lighting in unoccupied areas, pause the standby of small appliances in the guest room], or the initial action space vector can be [turn on the washing machine operation, turn on the standby of small appliances in the guest room].
[0068] The standard power consumption control strategy is a pre-set power consumption control scheme with reference value, which is used to optimize the initial power consumption control strategy.
[0069] The action space vector is a mathematical representation of all the actions that the agent can take for each standard power consumption control strategy, and the action spaces of different strategies are different. For example, the action space vector can be s = [control the operation time period of the washing machine, control the number of turned-on lighting fixtures, control the standby duration of small appliances in the guest room].
[0070] The state space vector is a vector that describes the environmental state in the Q-learning algorithm and can reflect the current power consumption demand state of the power system. For example, the state space vector can be a = [number of lighting fixtures turned on, occupancy rate of users, operating state of washing machines in the laundry room, operating state of dryers in the laundry room].
[0071] The reward function is a function that gives the agent corresponding rewards or penalties according to the state change of the system after the agent takes an action. If the action makes this embodiment tend to a better power consumption control effect, a positive reward is given; otherwise, a negative reward is given.
[0072] Specifically, this embodiment can determine the reward function based on weighted calculation of power consumption cost, power consumption load, and user occupancy rate.
[0073] The power consumption cost is the cost generated during the process of using electric energy, covering electricity bills, equipment loss costs, etc. The power consumption load is the total power supplied by the power system within a certain period of time, which reflects the magnitude of the user's demand for electric energy. The size of the power consumption load directly affects the operating state of the power supply system. The user occupancy rate is the ratio of the actual number of occupied users to the total number of users in the hotel area. This ratio is closely related to the power consumption demand. The higher the occupancy rate, the greater the potential power consumption demand.
[0074] The calculation formula of the reward function is:
[0075] Where, is the value of the reward function, is the weight coefficient of the electricity cost, is the weight coefficient of the electricity load, is the weight coefficient of the user occupancy rate, , is the electricity cost, is the electricity load, is the user occupancy rate, is the sub - function of the electricity cost, is the sub - function of the electricity load, is the sub - function of the user occupancy rate.
[0076]
[0077] Among them, , are constants, set according to experience, and are used to adjust the scale and offset of the sub - function of the electricity cost. For example, when is very large, will be very small, reflecting that a high - cost strategy gets a lower reward.
[0078]
[0079] Among them, , are constants, set according to experience, controls the opening size of the sub - function of the electricity load, controls the maximum value of the sub - function of the electricity load, is the average value of the electricity load. When is close to , gets the maximum value, reflecting that a strategy close to the reasonable load gets a higher reward.
[0080]
[0081] Among them, is a constant, used to adjust the slope of the sub - function of the user occupancy rate. When is the case, gets the maximum value, reflecting that a high - occupancy - rate strategy gets a higher reward.
[0082] The penalty factor is a coefficient used to increase the penalty intensity in the value function. When the agent takes a bad action, it increases the negative feedback to prompt it to learn better actions.
[0083] The value function is used to evaluate the long-term value of taking a certain action in a certain state. It comprehensively considers the immediate reward and the rewards that may be obtained in the future, and guides the agent to select the optimal action. The feedback value is the result calculated by the value function based on the current state and action, reflecting the quality of the action in the current state and providing a basis for policy adjustment.
[0084] The calculation formula of the value function is:
[0085] Where is the value function, is the value of the reward function, is the penalty factor, is the action space vector, is the state space vector, is the new action space vector, is the new state space vector.
[0086] The target power consumption control strategy is a power consumption control strategy that has been optimized by the Q-learning algorithm and can make the power system achieve better performance.
[0087] It can be concluded from the above that this embodiment can dynamically optimize the power consumption control strategy and improve the energy utilization efficiency. By defining the initial and standard action space vectors, combining with the real-time state space vector, and using the value function feedback, this embodiment realizes the intelligent regulation of power consumption behavior. This embodiment not only considers the reward mechanism to encourage energy-saving behaviors, but also introduces a penalty factor to restrict unreasonable power consumption, thus promoting the rationality of power consumption behavior.
[0088] In an embodiment of the present disclosure, the power control method further includes: Adjust the penalty factor based on the power consumption load; In response to the power consumption load being greater than the first load threshold, increase the penalty factor by the first step size; In response to the power consumption load being less than the second load threshold, decrease the penalty factor by the second step size; Where the first load threshold is greater than the second load threshold.
[0089] In this embodiment, the penalty factor is a coefficient used to increase the penalty intensity for bad behaviors. When the agent's action causes bad situations such as excessive power consumption load, the negative feedback is increased through the penalty factor to prompt the agent to adjust its behavior. Both the first load threshold and the second load threshold are pre-set power consumption load boundary values.
[0090] The first step length is the amplitude by which the penalty factor increases when the electricity consumption load is greater than the first load threshold, and this step length determines the rate of increase in the penalty intensity. The second step length is the amplitude by which the penalty factor decreases when the electricity consumption load is less than the second load threshold, and it is used to adjust the penalty intensity to adapt to different electricity consumption load situations. When the electricity consumption load is less than or equal to the first load threshold and greater than or equal to the second load threshold, the penalty factor can be controlled to remain unchanged.
[0091] As can be seen from the above, in this embodiment, by carefully considering the comprehensive cost-benefit of electricity consumption behavior and flexibly adjusting the penalty intensity in combination with the load situation, it not only effectively motivates energy-saving behaviors but also ensures the reasonable allocation of power resources. Especially during peak load periods, timely increasing the penalty factor helps to suppress excessive electricity consumption, while reducing the penalty during off-peak load periods encourages reasonable electricity consumption, thus achieving the dual goals of power supply-demand balance and energy conservation and emission reduction.
[0092] Corresponding to the electric energy control method in the above embodiment, Figure 2 is a structural block diagram of an electric energy control device provided by an embodiment of the present disclosure. For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The electric energy control device 20 includes: a power consumption prediction module 21, an initial power consumption control module 22, and a target power consumption control module 23.
[0093] Among them, the power consumption prediction module 21 is used to predict the first power consumption of the power supply main circuit based on the target digital twin model to obtain the second power consumption of the power supply main circuit. The first power consumption is the power consumption of the target area within the current preset time period, and the second power consumption is the predicted power consumption of the target area within the future target time period; The initial power consumption control module 22 is used to determine the initial power consumption control strategy for the power supply main circuit to distribute electric energy to multiple power supply branches based on the first difference. The first difference is the difference between the second power consumption of the power supply main circuit and the rated power consumption; The target power consumption control module 23 is used to adjust the initial power consumption control strategy based on the Q-learning algorithm to obtain the target power consumption control strategy.
[0094] In an embodiment of the present disclosure, the electric energy control device 20 further includes: a safety protection module, which is used to process the first power consumption parameters on multiple power supply branches to obtain a first monitoring result. The first power consumption parameters are the power consumption parameters of the electrical equipment on each power supply branch; the first monitoring result includes normal and abnormal; In response to the first monitoring result being abnormal, a safety protection instruction is output; the safety protection instruction is used to control the protection device of the corresponding power supply branch.
[0095] In an embodiment of the present disclosure, the target digital twin model includes a physical model and a target virtual model; The power control device 20 further includes: a model training module, configured to construct a physical model and a virtual model of the power supply main circuit based on historical power consumption data; there is a mapping relationship between the virtual model and the physical model; Iteratively optimize the parameters of the virtual model until the error between the output result of the virtual model and the output result of the physical model is less than the first error threshold, and obtain the first target parameter corresponding to the virtual model; Determine the target virtual model based on the first target parameter; Integrate the physical model and the target virtual model to determine the target digital twin model.
[0096] In an embodiment of the present disclosure, the model training module is specifically configured to determine the fitness function of the particle swarm optimization algorithm based on the dimensions of the particles in the particle swarm optimization algorithm. The dimensions of the particles in the particle swarm optimization algorithm include the learning rate, the batch value, and the neuron dropout rate; Determine the inertia weight of the particle swarm optimization algorithm based on the historical environmental characteristics of the power supply main circuit; Iteratively optimize the parameters of the virtual model based on the fitness function and the inertia weight.
[0097] In an embodiment of the present disclosure, the initial power consumption control module 22 is specifically configured to, if the first difference is less than the first difference threshold, the initial power consumption control strategy is to increase the number of electrical devices on multiple power supply branches; If the first difference is greater than or equal to the first difference threshold, the initial power consumption control strategy is to reduce the number of electrical devices on the power supply branch.
[0098] In an embodiment of the present disclosure, the target power consumption control module 23 is specifically configured to determine the initial action space vector of the Q-learning algorithm corresponding to the initial power consumption control strategy; Determine multiple standard power consumption control strategies, and determine the action space vectors of the Q-learning algorithm corresponding to each standard power consumption control strategy; Determine the state space vector of the Q-learning algorithm based on the second power consumption; Determine the value function of the Q-learning algorithm based on the action space vector, the state space vector, the reward function, and the penalty factor; Determine the target power consumption control strategy based on the initial action space vector and the feedback value of the value function.
[0099] In an embodiment of the present disclosure, the target power consumption control module 23 is specifically further configured to adjust the penalty factor based on the power consumption load; In response to the power consumption load being greater than the first load threshold, increase the penalty factor by the first step size; In response to the power consumption load being less than the second load threshold, decrease the penalty factor by the second step size; Among them, the first load threshold is greater than the second load threshold.
[0100] See Figure 3 , Figure 3 which is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, for example Figure 2 the functions of the power consumption prediction module 21, the initial power consumption control module 22, and the target power consumption control module 23 shown.
[0101] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0102] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0103] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0104] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first and second embodiments of the power control method provided by the embodiments of the present disclosure, and may also execute the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.
[0105] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the foregoing embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0106] The computer-readable storage medium may be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.
[0107] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0108] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0109] In several embodiments provided by this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings, direct couplings, or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical, or other forms of connection.
[0110] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.
[0111] In addition, the functional units in each embodiment of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0112] The above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by this disclosure, and these modifications or substitutions should be covered by the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.
Claims
1. A method for controlling electric energy, characterized in that, Including: Predicting the second power consumption of the power supply main line based on the target digital twin model to obtain the second power consumption of the power supply main line, where the first power consumption is the power consumption of the target area within the current preset time period, and the second power consumption is the predicted power consumption of the target area within the future target time period; Determining the initial power consumption control strategy for the power supply main line to distribute electric energy to multiple power supply branches based on the first difference, where the first difference is the difference between the second power consumption of the power supply main line and the rated power consumption; Adjusting the initial power consumption control strategy based on the Q-learning algorithm to obtain the target power consumption control strategy.
2. The electric energy control method according to claim 1, characterized in that Also including: Processing the first power consumption parameters on the multiple power supply branches to obtain the first monitoring result, where the first power consumption parameters are the power consumption parameters of the electrical equipment on each power supply branch; the first monitoring result includes normal and abnormal; When the first monitoring result is abnormal, outputting a safety protection instruction; the safety protection instruction is used to control the protection device of the corresponding power supply branch.
3. The electric energy control method according to claim 1, characterized in that Also including: The target digital twin model includes a physical model and a target virtual model; The training process of the target digital twin model includes: Constructing the physical model and the virtual model of the power supply main line based on historical power consumption data; there is a mapping relationship between the virtual model and the physical model; Iteratively optimizing the parameters of the virtual model until the error between the output result of the virtual model and the output result of the physical model is less than the first error threshold, and obtaining the first target parameters corresponding to the virtual model; Determining the target virtual model based on the first target parameters; Integrating the physical model and the target virtual model to determine the target digital twin model.
4. The power control method according to claim 3, wherein The iteratively optimizing the parameters of the virtual model includes: Determining the fitness function of the particle swarm optimization algorithm based on the dimensions of the particles in the particle swarm optimization algorithm, where the dimensions of the particles in the particle swarm optimization algorithm include the learning rate, batch value, and neuron dropout rate; Determining the inertia weight of the particle swarm optimization algorithm based on the historical environmental characteristics of the power supply main line; Iteratively optimizing the parameters of the virtual model based on the fitness function and the inertia weight.
5. The electric energy control method according to claim 1, wherein The determining the initial power consumption control strategy for the power supply main line to distribute electric energy to multiple power supply branches based on the first difference includes: If the first difference is less than the first difference threshold, the initial power consumption control strategy is to increase the number of electrical equipment on the multiple power supply branches; If the first difference is greater than or equal to the first difference threshold, the initial power consumption control strategy is to reduce the number of electrical equipment on the power supply branches.
6. The electric energy control method according to claim 1, characterized in that The adjusting the initial power consumption control strategy based on the Q-learning algorithm to obtain the target power consumption control strategy includes: Determining the initial action space vector of the Q-learning algorithm corresponding to the initial power consumption control strategy; Determining multiple standard power consumption control strategies and determining the action space vectors of the Q-learning algorithm corresponding to each standard power consumption control strategy; Determining the state space vector of the Q-learning algorithm based on the second power consumption; Determine the value function of the Q - learning algorithm based on the action space vector, the state space vector, the reward function, and the penalty factor; Determine the target power consumption control strategy based on the initial action space vector and the feedback value of the value function.
7. The electric energy control method according to claim 6, wherein It further includes: Adjust the penalty factor based on the power consumption load; In response to the power consumption load being greater than the first load threshold, increase the penalty factor by the first step size; In response to the power consumption load being less than the second load threshold, decrease the penalty factor by the second step size; Wherein, the first load threshold is greater than the second load threshold.
8. An electric energy control device, characterized in that, It includes: A power consumption prediction module, configured to predict the first power consumption of the power supply main circuit based on the target digital twin model to obtain the second power consumption of the power supply main circuit, where the first power consumption is the power consumption of the target area within the current preset time period, and the second power consumption is the predicted power consumption of the target area within the future target time period; An initial power consumption control module, configured to determine the initial power consumption control strategy for the power supply main circuit to distribute electric energy to multiple power supply branches based on the first difference, where the first difference is the difference between the second power consumption of the power supply main circuit and the rated power consumption; A target power consumption control module, configured to adjust the initial power consumption control strategy based on the Q - learning algorithm to obtain the target power consumption control strategy.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for predicting building energy consumption based on depth reinforcement learning
CN109063903A
Energy scheduling method and device based on digital twinning
CN116757328A
Power dispatching method, system and equipment for production enterprise and medium
CN118211797A
Park-level adjustable load-oriented resource regulation and control method, device and equipment and medium
CN119539424A
Coordination and optimization method and system for comprehensive electric-thermal energy system, and device, medium and program
WO2023082697A1