Smart home energy supply and demand collaborative optimization system based on reinforcement learning
By applying a collaborative optimization system for energy supply and demand based on reinforcement learning in smart homes, the problem that traditional systems cannot make intelligent decisions is solved, efficient energy utilization and cost control are achieved, and energy waste and carbon emissions are reduced.
Patent Information
- Application Number
- CN202510511065.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional smart home energy management systems cannot make intelligent decisions based on real-time energy supply and demand conditions and environmental changes, resulting in energy waste, high costs and environmental protection problems.
The smart home energy supply and demand collaborative optimization system based on reinforcement learning is adopted, and efficient energy regulation and optimization is achieved through energy data acquisition module, reinforcement learning decision module, supply and demand collaborative optimization module, real-time regulation and execution module and system performance monitoring module.
It realizes efficient utilization and cost control of energy, reduces energy waste and carbon emissions, and improves the system's intelligent decision-making capabilities and the scientific nature of energy management.
Smart Images

Figure CN120049522A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart home, and particularly to a smart home energy supply and demand collaborative optimization system based on reinforcement learning. Background Art
[0002] With the rapid development of technology, smart homes have gradually entered people's lives. The wide application of various smart home appliances and devices has brought challenges in energy management while improving the convenience of life. The traditional home energy management method is relatively extensive, lacking precise regulation and optimization of energy supply and demand, resulting in widespread energy waste and high energy costs.
[0003] From the perspective of the energy supply side, the application of renewable energy sources such as solar energy and wind energy in homes is increasing day by day. However, such energy sources are unstable and intermittent. For example, solar power generation depends on lighting conditions, and the power generation will decrease significantly or even stop at night or on rainy and cloudy days; wind power generation is affected by wind speed and stability, making it difficult to provide continuous and stable power supply. This makes it difficult for home energy supply to meet real-time demand, unable to effectively store and rationally distribute energy during energy surplus, and unable to replenish energy in time during energy shortage, resulting in waste of energy resources.
[0004] On the energy demand side, there is a wide variety of home electrical equipment and complex usage patterns. The power requirements of different devices vary greatly. High-power devices such as air conditioners and electric water heaters consume a large amount of electric energy when operating; while some small devices such as smart lights and routers, although with small power, have a large number and run for a long time, and the cumulative energy consumption cannot be ignored. In addition, users' electricity consumption habits are also different. Some users are accustomed to using a large number of electrical appliances at night, while others concentrate on using electricity during the day, which makes the home electricity load fluctuate significantly at different times, increasing the difficulty of energy management.
[0005] Most of the existing smart home energy management systems are controlled based on simple rules or preset programs, and cannot make intelligent decisions according to real-time energy supply and demand conditions and environmental changes. For example, some systems can only control the start and stop of devices according to a fixed schedule, and cannot flexibly respond to sudden changes in energy demand or fluctuations in renewable energy power generation. Moreover, these systems often do not fully consider the balance among energy cost, user comfort, and environmental protection factors. When pursuing to reduce energy cost, user comfort may be sacrificed; while overly focusing on user comfort may lead to energy waste and increased carbon emissions.
[0006] At the same time, traditional systems also have deficiencies in data collection and analysis. The types of data they collect are limited, usually only focusing on basic information such as the electricity consumption of devices, ignoring data that is crucial for energy management, such as environmental parameters and the status of energy storage devices. Moreover, there is a lack of in-depth analysis and mining of the collected data, and it is unable to provide a comprehensive and accurate basis for energy regulation. In the face of abnormal situations such as tight energy supply or equipment failures, the response ability of traditional systems is also weak, and they cannot take effective countermeasures in a timely manner, affecting the stability and reliability of household energy supply. Summary of the Invention
[0007] The purpose of the present invention is to provide a smart home energy supply and demand collaborative optimization system based on reinforcement learning to solve the problems raised in the above-mentioned background technology.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A smart home energy supply and demand collaborative optimization system based on reinforcement learning, the system includes: An energy data collection module, a reinforcement learning decision-making module, a supply and demand collaborative optimization module, a real-time regulation execution module, and a system performance monitoring module; The energy data collection module is used to collect household energy supply and demand data, including the power of electrical equipment, the power generation of renewable energy, the status of energy storage devices, and environmental parameters; The reinforcement learning decision-making module constructs a state space based on the collected data, and generates an energy regulation strategy in combination with a preset action space and a reward function; The supply and demand collaborative optimization module performs multi-objective optimization on energy distribution according to the energy regulation strategy, and generates a real-time regulation instruction; The real-time regulation execution module adjusts the operating state of electrical equipment, the charge and discharge strategy of energy storage devices, and the energy transmission path according to the real-time regulation instruction; The system performance monitoring module real-time tracks the operating states of each module, collects abnormal deviation signals during the regulation process, and sends the abnormal deviation signals to the reinforcement learning decision-making module for policy iteration and update.
[0009] Preferably, the specific operation process of the energy data collection module is as follows: In the data collection stage, the instantaneous power of electrical equipment, the power generation of renewable energy, the remaining power of the energy storage device, and the environmental temperature and humidity data are obtained through a sensor network; the deviation value between the instantaneous power and a preset power threshold is marked as a power deviation value, the deviation value between the remaining power and a preset power threshold is marked as an energy storage margin deviation value, and the difference between the environmental temperature and humidity and a preset comfort interval is marked as an environmental deviation value; if the power deviation value, the energy storage margin deviation value, or the environmental deviation value exceeds the corresponding preset threshold, a data abnormality signal is generated and sent to the system performance monitoring module.
[0010] Preferably, the specific decision-making process of the reinforcement learning decision-making module is as follows: Define the state space as a vector set containing power consumption demand, energy supply volume, energy storage state, and time characteristics; the action space includes device start / stop instructions, energy storage charge / discharge rate, and energy trading strategy; the reward function is calculated by multi-dimensional weighting based on energy cost, user comfort, and carbon emissions; update the policy network parameters through the Q-learning algorithm. If the cumulative reward value is lower than the preset convergence threshold during the policy iteration process, generate a policy optimization signal and trigger the expansion of the action space.
[0011] Preferably, the specific optimization process of the supply-demand collaborative optimization module is as follows: Model the energy distribution problem as a multi-objective optimization model. The objective functions include minimizing the grid power purchase cost, maximizing the utilization rate of renewable energy, and balancing the service life of energy storage devices; use the non-dominated sorting genetic algorithm to generate the Pareto optimal solution set, and select the optimal distribution plan from the solution set according to the real-time electricity price signal and user preferences; if there is no feasible solution in the solution set, generate an optimization failure signal and switch to the standby scheduling mode.
[0012] Preferably, the specific execution process of the real-time regulation execution module is as follows: Classify the instructions into emergency regulation instructions, regular regulation instructions, and delayed regulation instructions according to the priority of the regulation instructions; the emergency regulation instructions directly interrupt the operation of high-power-consuming devices, the regular regulation instructions achieve progressive optimization by adjusting the device working mode, and the delayed regulation instructions are executed during off-peak hours in combination with user behavior prediction; if the actual energy consumption deviates from the expectation by more than the preset fault tolerance threshold after the instruction is executed, generate an execution exception signal and feedback it to the supply-demand collaborative optimization module.
[0013] Preferably, the system performance monitoring module is communicatively connected to the dynamic policy evaluation module. The dynamic policy evaluation module collects the policy execution efficiency, abnormal signal frequency, and user satisfaction indicators in the historical regulation data, and weights and fuses them into a policy evaluation value; if the policy evaluation value is lower than the preset evaluation threshold, trigger the policy reset function of the reinforcement learning decision-making module.
[0014] Preferably, the specific evaluation process of the dynamic policy evaluation module is as follows: Define the policy execution efficiency as the ratio of the actual energy saving rate to the theoretical maximum energy saving rate, mark the abnormal signal frequency as the number of abnormal signal generations per unit time, and score the user satisfaction based on device availability and environmental comfort; calculate the weights of each index through the entropy weight method, and sum them up by weighting to obtain the policy evaluation value; if the policy evaluation value is lower than the threshold for N consecutive evaluation periods, it is determined that the current policy fails, where N is a preset integer constant.
[0015] Preferably, it further includes an abnormal pattern recognition module. The abnormal pattern recognition module constructs a normal energy consumption pattern library by analyzing the periodic fluctuations and sudden peaks in the historical energy consumption data; the real-time collected energy consumption data is matched with the pattern library for similarity. If the matching degree is lower than the preset similarity threshold, an energy consumption abnormal signal is generated and an artificial intervention process is triggered.
[0016] Preferably, the specific recognition process of the abnormal pattern recognition module is as follows: The dynamic time warping algorithm is used to calculate the distance between the real-time energy consumption curve and each template curve in the pattern library, and the minimum distance value is marked as the matching degree; if the matching degree is lower than the threshold, abnormal features including peak duration, fluctuation frequency, and equipment relevance are extracted, and an abnormal type code is generated according to the feature combination and associated with a preset processing strategy.
[0017] Preferably, it further includes a multi-agent cooperation module. The multi-agent cooperation module models each energy device in the home as an independent agent, and realizes the coordination of local decision-making and global optimization through a distributed reinforcement learning algorithm; if the local strategy of a single agent conflicts with the global goal, the optimization weight is reallocated through a game theory model.
[0018] Compared with the prior art, the beneficial effects of the present invention are: In the energy data collection link, the energy data collection module obtains rich information such as the instantaneous power of electrical equipment, the power generation power of renewable energy, the remaining power of energy storage equipment, and environmental temperature and humidity data through a sensor network, and deeply processes the data. By comparing the power, power consumption, and environmental data with preset thresholds, data anomalies can be detected in a timely manner and signals can be sent to ensure that the energy data obtained by the system is accurate and reliable, providing a solid data basis for subsequent decision-making and regulation. This comprehensive and accurate data collection method greatly improves the integrity and effectiveness of the data compared with the traditional data collection method that only focuses on single electrical consumption data, enabling the system to more comprehensively understand the home energy supply and demand situation.
[0019] The reinforcement learning decision-making module constructs a state space covering electricity demand, energy supply, energy storage status, and time characteristics based on the collected data. Combining an action space containing various flexible regulation methods and a reward function calculated by multi-dimensional weighting of energy cost, user comfort, and carbon emissions, it uses the Q-learning algorithm to generate an energy regulation strategy. This process enables the system to have powerful intelligent decision-making capabilities, comprehensively consider various factors, and make optimal decisions according to different energy supply and demand scenarios. Different from the traditional simple preset rule-based decision-making method, it can dynamically adapt to environmental changes, continuously optimize the strategy, and thus achieve efficient energy utilization and effective cost control. For example, when renewable energy generation is sufficient, it preferentially uses renewable energy to power equipment and charge energy storage devices, reducing power purchase from the grid and lowering energy costs. At the same time, according to the user's comfort requirements at different times, it reasonably adjusts the operating state of the equipment, minimizing energy consumption and carbon emissions while ensuring user comfort.
[0020] The supply-demand collaborative optimization module models the energy allocation problem as a multi-objective optimization model, aiming to minimize the grid power purchase cost, maximize the utilization rate of renewable energy, and balance the lifespan of energy storage devices. It uses the non-dominated sorting genetic algorithm to generate a Pareto optimal solution set and selects the optimal allocation scheme in combination with real-time electricity price signals and user preferences. This optimization process fully considers multiple aspects of energy supply and demand, realizing the scientific allocation of energy. When the real-time electricity price is low, the system will reasonably increase the power purchase from the grid and store it in the energy storage device, and give priority to using the energy storage device and renewable energy when the price is high, effectively reducing energy costs. By optimizing the utilization method of renewable energy, it increases the proportion of renewable energy in household energy consumption, promotes the use of clean energy, reduces dependence on traditional grid power, and reduces carbon emissions, with significant environmental benefits. Moreover, in the process of optimizing energy allocation, it evenly considers the lifespan of energy storage devices, avoiding damage to energy storage devices caused by overcharging and over-discharging, extending the service life of energy storage devices, and reducing equipment replacement costs.
[0021] The real-time control execution module classifies and executes according to the priority of the control instructions, and adopts different control methods for instructions with different priorities. Emergency control instructions can quickly interrupt the operation of high-power-consuming devices in case of sudden energy supply tension to ensure the stability of energy supply; conventional control instructions achieve progressive optimization by adjusting the device working mode, reducing energy consumption without affecting the normal life of users; delayed control instructions are executed during off-peak hours in combination with user behavior prediction, making full use of the power resources during low electricity price periods to further reduce energy costs. At the same time, if the actual energy consumption deviates from the expected value by more than the preset fault tolerance threshold after the instruction is executed, it will be timely feedback and re-evaluated and optimized to ensure the accuracy and reliability of the control effect. This refined control execution method not only ensures the emergency handling ability of the system in case of emergencies, but also realizes the optimized management of daily energy use, improves energy utilization efficiency, and reduces energy waste.
[0022] The system performance monitoring module tracks the running status of each module in real time, collects abnormal deviation signals and feeds them back to the reinforcement learning decision-making module for policy iteration and update. The dynamic policy evaluation module collects indicators such as policy execution efficiency, abnormal signal frequency and user satisfaction, and fuses them into a policy evaluation value through the entropy weight method. When the evaluation value is lower than the preset threshold, it triggers policy reset. This closed-loop monitoring and evaluation mechanism enables the system to continuously self-optimize and improve, and continuously improve performance. By analyzing historical control data, problems existing in the policy are timely discovered and adjusted to ensure that the system always operates in the optimal state and provides stable and efficient energy management services for users.
[0023] The abnormal mode recognition module constructs a normal energy consumption mode library by analyzing historical energy consumption data, collects real-time energy consumption data for similarity matching, timely discovers abnormal energy consumption situations and triggers the manual intervention process. This function effectively guarantees the safety of household energy use, can timely discover potential equipment failures or energy waste problems, and avoid potential safety hazards and economic losses caused by abnormal energy use.
[0024] The multi-agent collaboration module models each energy device in the home as an independent agent, and realizes the collaboration of local decision-making and global optimization through a distributed reinforcement learning algorithm. When there is a conflict between the local strategy of a single agent and the global goal, the game theory model is used to reallocate the optimization weights, improving the overall collaboration and adaptability of the system. This collaboration mechanism gives full play to the autonomy and intelligence of each energy device, enabling them to serve the global optimization goal of household energy supply and demand while meeting their own operation requirements, and further improving the efficiency and effect of energy management. Description of the Drawings
[0025] Figure 1 It is the working principle diagram of the smart home energy supply and demand collaborative optimization system described in the present invention; Figure 2 It is the working principle diagram of the energy data acquisition module; Figure 3 It is the working principle diagram of the supply-demand collaborative optimization module; Figure 4 It is the working principle diagram of the specific recognition process of the abnormal mode recognition module. Specific implementation manners
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] Please refer to Figures 1-4 , the present invention provides a smart home energy supply-demand collaborative optimization system based on reinforcement learning, and its overall implementation solution is as follows: This system mainly consists of an energy data acquisition module, a reinforcement learning decision-making module, a supply-demand collaborative optimization module, a real-time regulation execution module, and a system performance monitoring module.
[0028] The energy data acquisition module is responsible for collecting data related to household energy supply and demand. By deploying sensors at various key positions in the home, information such as the power of electrical appliances, the power generation of renewable energy, the state of energy storage devices, and environmental parameters is collected. For example, a power sensor is installed at the power access point of each electrical appliance to obtain the power consumption of the device in real time; a power monitoring device supporting the solar panel is used to collect the power generation of renewable energy; the state of the energy storage device is mastered through the built-in power detection chip of the energy storage device; a temperature and humidity sensor is used to collect environmental temperature and humidity data. These collected data will serve as the basis for subsequent system decisions.
[0029] The reinforcement learning decision-making module constructs a state space based on the collected data. The state space covers information in multiple dimensions such as electricity demand, energy supply, energy storage state, and time characteristics, forming a vector set. At the same time, an action space is preset, which includes various executable actions such as device start-stop instructions, energy storage charge-discharge rates, and energy trading strategies. A reward function is designed according to factors such as energy cost, user comfort, and carbon emissions, and the policy network parameters are continuously updated through the Q-learning algorithm to generate an energy regulation strategy.
[0030] Based on the energy regulation strategy generated by the reinforcement learning decision-making module, the supply-demand collaborative optimization module performs multi-objective optimization on energy distribution. This module models the energy distribution problem as a multi-objective optimization model, considering objectives such as minimizing the power purchase cost of the power grid, maximizing the utilization rate of renewable energy, and balancing the service life of energy storage devices. Through a specific algorithm, a Pareto optimal solution set is generated, and combined with real-time electricity price signals and user preferences, the optimal distribution plan is selected from the solution set, and finally real-time regulation instructions are generated.
[0031] The real-time regulation execution module receives the real-time regulation instructions issued by the supply-demand collaborative optimization module. According to the priority of the instructions, the instructions are divided into emergency regulation instructions, regular regulation instructions, and delayed regulation instructions. Different execution methods are adopted for instructions with different priorities. Emergency regulation instructions will directly interrupt the operation of high-power-consuming devices to quickly respond to emergencies such as tight energy supply; regular regulation instructions achieve progressive optimization by adjusting the device working mode; delayed regulation instructions are executed during off-peak hours in combination with user behavior prediction to avoid affecting the normal life of users. During the execution of the instructions, if the actual energy consumption deviates from the expected value by more than the preset fault tolerance threshold, an execution exception signal is generated and fed back to the supply-demand collaborative optimization module.
[0032] The system performance monitoring module continuously tracks the operating status of each module. During the data collection process of the energy data collection module, if the power deviation value, energy storage margin deviation value, or environmental deviation value exceeds the corresponding preset threshold, the system performance monitoring module will receive a data exception signal. At the same time, when the real-time regulation execution module executes instructions, if an execution exception signal appears, the system performance monitoring module will also receive it. In addition, the system performance monitoring module will also collect abnormal deviation signals during the regulation process and send these signals to the reinforcement learning decision-making module for iterative update of the strategy to continuously optimize the system performance.
[0033] The implementation of the present invention will be further described below in conjunction with Embodiments 1 to 6. Embodiment 1
[0034] In the data collection stage, high-precision instantaneous power sensors are installed on each electrical device in the household. These sensors can accurately measure the real-time power of the device during operation. Taking a common air conditioner device as an example, the sensor collects instantaneous power data at regular intervals (such as every 1 minute) to ensure that the power changes of the device can be captured in a timely manner. For renewable energy power generation devices, such as solar panels, the generated power is collected in real time through a supporting power generation monitoring device. This device can convert the electrical energy generated by the solar panels into electrical signals for measurement and recording. In terms of energy storage devices, the remaining power is obtained by using the built-in battery capacity detection chip, accurate to the percentage of the battery capacity. The environmental temperature and humidity data are collected by temperature and humidity sensors distributed at different positions in the room to obtain more comprehensive and accurate environmental information.
[0035] After the data is collected, it needs to be further processed. The instantaneous power of the electrical equipment is compared with the preset power threshold, and the deviation value between the two is calculated. This deviation value is marked as the power deviation value. The preset power threshold is set according to the rated power of the equipment and the power fluctuation range during normal operation. For example, for an electric kettle with a rated power of 1000W, the preset power threshold is set between 900W - 1100W. If the collected instantaneous power is 850W, then the power deviation value is 850 - 900 = -50W. For the remaining power of the energy storage device, it is compared with the preset power threshold, and the deviation value, that is, the energy storage margin deviation value, is calculated. The preset power threshold is determined according to the total capacity of the energy storage device and the optimal power usage range. For example, for an energy storage device with a total capacity of 10 kWh, the optimal power usage range is set between 3 - 8 kWh. If the remaining power is 2 kWh, the energy storage margin deviation value is 2 - 3 = -1 kWh. The environmental temperature and humidity data also need to be compared with the preset comfort interval, and the difference is calculated to obtain the environmental deviation value. The preset comfort interval is set according to the comfortable feeling range of the human body for temperature and humidity. Generally, the temperature comfort interval is set between 22°C - 26°C, and the humidity comfort interval is set between 40% - 60%. Assuming the collected temperature is 20°C and the humidity is 35%, then the temperature deviation value is 20 - 22 = -2°C, and the humidity deviation value is 35 - 40 = -5%.
[0036] When the power deviation value, the energy storage margin deviation value, or the environmental deviation value exceeds the corresponding preset threshold, it indicates that the data is abnormal. At this time, the energy data acquisition module will generate a data anomaly signal and send this signal to the system performance monitoring module. For example, if the power deviation value of the electric kettle exceeds the preset range of ±100W, or the energy storage margin deviation value of the energy storage device is lower than -2 kWh, or the environmental temperature deviation value exceeds ±3°C, it will trigger the generation and sending of the data anomaly signal. After receiving the signal, the system performance monitoring module will take corresponding measures, such as prompting the user to check the equipment operation status, further analyzing the data, etc., to ensure that the system can timely detect and handle potential problems and guarantee the stable operation of the smart home energy supply and demand collaborative optimization system. Example 2
[0037] The reinforcement learning decision-making module is the core part of the entire system to achieve intelligent decision-making, and the accuracy and effectiveness of its decision-making directly affect the energy optimization effect of the system.
[0038] First, clarify the definition of the state space. The state space is defined as a set of vectors containing electricity demand, energy supply, energy storage state, and time characteristics. Electricity demand is obtained by collecting power data of each electrical device and conducting statistics and predictions based on the device's operating time and usage frequency. For example, during dinner time, the usage frequency of kitchen appliances such as microwave ovens and rice cookers increases. By analyzing historical data and monitoring the current operating conditions of the devices, the electricity demand at this time can be estimated. The energy supply mainly comes from renewable energy generation and grid power supply, which is determined by collecting the power generation of solar panels, the power of wind power generation equipment, and the real-time power supply data of the grid. The energy storage state is represented by obtaining information such as the remaining power and charge / discharge status of the energy storage device. Time characteristics include specific time points, dates, seasons, etc., because the energy supply and demand situations vary greatly at different time points and seasons. For example, in summer, the solar power generation is large during the day, and at the same time, the electricity demand for cooling equipment such as air conditioners is also large; in winter, the electricity demand is high at night, but the renewable energy generation is relatively small. Integrate this information into a vector as the input state of the reinforcement learning algorithm.
[0039] The action space includes device start / stop commands, energy storage charge / discharge rates, and energy trading strategies. Device start / stop commands are used to control the turning on and off of electrical devices. According to the energy supply and demand situation and user needs, it is decided which devices need to be turned on or off. For example, when the renewable energy generation is sufficient and the energy storage device is fully charged, some non-urgent but power-consuming devices such as washing machines can be turned on for laundry operations. The energy storage charge / discharge rate determines the speed of charging and discharging of the energy storage device, which is adjusted according to the energy cost and the state of the energy storage device. If the current electricity price is low and the energy storage device has insufficient power, a higher charging rate can be set; when it is the peak electricity consumption period and the energy storage device is fully charged, the discharge rate is increased to supply power to the household. The energy trading strategy involves power trading with the grid, including purchasing electricity from the grid at a low electricity price and storing it in the energy storage device, and selling the electricity in the energy storage device back to the grid at a high electricity price to reduce the energy cost.
[0040] The reward function is a key part of the reinforcement learning decision-making module, which is calculated by multi-dimensional weighting based on energy cost, user comfort, and carbon emissions. Energy cost mainly considers the grid power purchase cost and renewable energy generation cost. Assume the grid power purchase unit price is yuan per degree, the grid power purchase quantity is degrees, the renewable energy generation cost is yuan per degree, the renewable energy generation quantity is degrees, the calculation formula of the energy cost is . User comfort is measured by a comprehensive evaluation of factors such as environmental temperature and humidity, and equipment availability. For example, when the indoor temperature is maintained within a comfortable range and the user's commonly used equipment can operate normally, the user comfort is high. The carbon emissions are calculated based on the type and quantity of energy consumption, and the carbon emission coefficients of different energy sources are different. Assume that the carbon emission coefficient of using grid electricity is kg / kWh, and the carbon emission coefficient of using renewable energy is kg / kWh, the grid electricity consumption is kWh, and the renewable energy power generation is kWh. The carbon emissions can be calculated by the formula . By setting different weights , , , a weighted sum of the energy cost, user comfort, and carbon emissions is obtained to get the reward function , that is , where , , are the maximum values of the energy cost, user comfort, and carbon emissions respectively, and is the user comfort score.
[0041] In the decision-making process, the parameters of the policy network are updated through the Q-learning algorithm. The Q-learning algorithm continuously learns and updates the Q value based on the current state and the executed action to find the optimal policy. If during the policy iteration process, the cumulative reward value is lower than the preset convergence threshold, it indicates that the current policy effect is not good and needs to be optimized. At this time, a policy optimization signal is generated and the action space expansion is triggered. For example, new device control strategies can be added or the rate options of energy storage charging and discharging can be adjusted to explore more possible decision-making methods and improve the decision-making ability and energy optimization effect of the system. Example 3
[0042] This example focuses on introducing the specific optimization process of the supply-demand collaborative optimization module. The supply-demand collaborative optimization module plays a crucial role in the entire smart home energy supply-demand collaborative optimization system, which further refines the energy regulation strategy generated by the reinforcement learning decision-making module into executable real-time regulation instructions.
[0043] First, the energy allocation problem is modeled as a multi-objective optimization model. The objective function of this model includes multiple aspects, and one of them is to minimize the grid power purchase cost. In actual household electricity use, the grid power purchase cost is an important part of the energy expenditure. Assume that the real-time electricity price of the grid is yuan / kWh, and the electricity purchased from the grid within the time period is kWh. Then the grid power purchase cost within a period of time can be expressed as To reduce this part of the cost, it is necessary to reasonably arrange the time and amount of power purchased from the power grid according to the real-time electricity price signal and the household energy supply and demand situation.
[0044] Another goal is to maximize the utilization rate of renewable energy. With the enhancement of environmental awareness and the development of renewable energy technologies, making full use of renewable energy is of great significance for reducing energy costs and carbon emissions. Let the power generation of renewable energy in the time period be kWh, and the total electricity consumption of the household in this time period be kWh. The calculation formula for the utilization rate of renewable energy is. To improve the utilization rate of renewable energy, it is necessary to give priority to using renewable energy to meet household electricity demand. When the power generation of renewable energy is in excess, consider storing the excess electric energy in energy storage devices or making other reasonable uses.
[0045] Balancing the life of energy storage devices is also one of the important goals. The charge and discharge times and charge and discharge depth of energy storage devices have a significant impact on their life. Assume that the maximum charge and discharge times of the energy storage device is, and the charge and discharge times in a period of time is , and the charge and discharge depth is ( ). The life loss of the energy storage device can be expressed by a certain functional relationship, such as . To balance the life of the energy storage device, it is necessary to reasonably control the charge and discharge strategy of the energy storage device to avoid overcharging and over-discharging.
[0046] The non-dominated sorting genetic algorithm is used to generate the Pareto optimal solution set. The non-dominated sorting genetic algorithm is an effective multi-objective optimization algorithm. It searches for the optimal solution in the solution space by simulating the selection, crossover, and mutation operations in the natural evolution process. In this system, the algorithm continuously optimizes the energy distribution plan to generate a series of non-dominated solutions, and these solutions constitute the Pareto optimal solution set. Each plan in this solution set balances the goals such as the power grid power purchase cost, the utilization rate of renewable energy, and the life of energy storage devices to varying degrees.
[0047] Select the optimal allocation plan from the solution set according to the real-time electricity price signal and user preferences. The real-time electricity price signal reflects the real-time cost of electricity in the power grid. When the electricity price is low, the electricity purchased from the power grid can be appropriately increased and stored in the energy storage device; when the electricity price is high, the energy storage device or renewable energy is preferentially used for power supply. User preferences take into account the user's emphasis on aspects such as energy cost and environmental protection. For example, some users pay more attention to environmental protection and are willing to give priority to using renewable energy, even if this may increase a certain amount of energy cost; while some users are more concerned about energy cost and hope to minimize the electricity bill. According to these factors, select the plan that best suits the current situation from the Pareto optimal solution set as the final energy allocation plan.
[0048] If there is no feasible plan in the solution set, it means that the current energy supply and demand situation is relatively complex and a reasonable allocation plan cannot be obtained through the existing optimization strategies. At this time, generate an optimization failure signal and switch to the standby scheduling mode. The standby scheduling mode can adopt some simple rules for energy allocation, such as giving priority to ensuring the electricity demand of important equipment, or obtaining electricity from the power grid and renewable energy according to a fixed ratio, etc., to ensure the basic stability of the home energy supply and wait for the system to re-adjust the optimization strategy before resuming normal operation. Embodiment 4
[0049] The real-time regulation execution module converts the energy regulation instruction into actual device operations, and its execution process directly affects the regulation effect of the system on smart home energy. After receiving the regulation instruction from the supply-demand collaborative optimization module, the real-time regulation execution module will classify the instruction into an emergency regulation instruction, a regular regulation instruction, and a delayed regulation instruction according to factors such as the urgency of the instruction and the impact on energy optimization, and then execute these instructions according to different priorities and methods.
[0050] For emergency control instructions, they mainly respond to sudden situations that may threaten the stability of household energy supply. For example, there are serious faults in the grid power supply, large fluctuations in voltage or a sharp drop in the power supply; or sudden abnormalities occur in the renewable energy generation equipment in the household, and the energy storage equipment cannot replenish energy in time, etc. When such emergencies occur, the emergency control instructions will quickly come into play and directly interrupt the operation of high-power-consuming devices. Take the electric heater in the household as an example. During the peak electricity consumption period in winter, if the grid power supply is severely insufficient, in order to ensure the stability of the overall energy supply and maintain the normal operation of key devices (such as refrigerators, medical devices, etc.), the system will send an emergency stop operation instruction to the electric heater. When executing this instruction, the system will quickly cut off the power supply of the electric heater, making it stop working immediately. At the same time, the system will record the relevant information of this emergency control, including the execution time, the devices involved, the type of emergency situation, etc., for subsequent analysis and summary to optimize the emergency control strategy. And the system will timely inform the user of the reason for the electric heater to stop running through the intelligent home interaction interface (such as voice prompts of smart speakers, push messages of mobile phone APPs, etc.), to avoid the user's doubts and troubles caused by the sudden stop of the device.
[0051] Regular control instructions aim to achieve the gradual optimization of energy utilization. By adjusting the working modes of devices, on the premise of not affecting the normal life of users, the energy consumption is gradually reduced. For example, for a smart air conditioner, the regular control instructions will intelligently adjust the operation mode of the air conditioner according to information such as the indoor and outdoor temperatures, humidity, and the comfort parameters set by the user. If the current indoor temperature is relatively close to the temperature set by the user and the outdoor temperature is relatively low, the system will adjust the air conditioner from the cooling mode to the ventilation mode, using the natural outdoor wind to adjust the indoor temperature, thereby reducing the working time of the air conditioner compressor and lowering the energy consumption. For a washing machine with multiple working modes, the system will select the appropriate washing mode according to the quantity and material of the clothes and the current energy supply and demand situation. If the energy supply is relatively tight, and the quantity of clothes is small and the material is not very dirty, the system will select the energy-saving mode, by appropriately extending the washing time, reducing the dehydration speed, etc., to reduce the energy consumption while ensuring the washing effect. During the execution of the regular control instructions, the system will continuously monitor the operation status and energy consumption changes of the devices, and evaluate the control effect in real time. If it is found that the actual energy consumption does not reach the expected optimization goal, the system will fine-tune the control parameters or re-select a more appropriate working mode to ensure the continuous improvement of the energy utilization optimization effect.
[0052] The delayed regulation instruction is an important means to achieve energy optimization by combining user behavior prediction. The system constructs a user behavior model by analyzing the user's historical electricity consumption data and behavior habits at different time periods, so as to accurately predict the user's electricity demand at different times. For example, through long-term data monitoring, it is found that a certain user is usually in a resting state between 10 pm and 6 am, and only a few low-power devices (such as routers, night lights, etc.) are running at home during this period, which belongs to the low electricity consumption period. For some electricity-consuming tasks that are not time-sensitive, such as electric vehicle charging and smart washing machine laundry, the system will set the corresponding regulation instructions as delayed regulation instructions and arrange them to be executed during the low electricity consumption period. Taking electric vehicle charging as an example, the system will preferentially choose the night time period with lower electricity prices to charge within the charging time range set by the user. Before executing the delayed regulation instruction, the system will plan in advance the start time and charging power of the charging task to avoid excessive local electricity load caused by multiple devices charging at the same time. At the same time, the system will interact with the user and display the execution plan of the delayed regulation instruction to the user through means such as a mobile APP, including specific execution time, estimated completion time and other information, so that the user can understand in advance and reasonably arrange their lives. If the user has special needs, such as needing to use the electric vehicle in advance, they can also manually adjust the charging time through the APP, and the system will adjust the execution plan of the delayed regulation instruction in time according to the user's operation.
[0053] During the entire process of instruction execution, the system will monitor the deviation between the actual energy consumption and the expected energy consumption in real time. If the deviation between the actual energy consumption and the expected value exceeds the preset fault tolerance threshold, it indicates that the instruction execution is abnormal, and there may be problems such as equipment failure and unreasonable regulation strategy. For example, after executing the regulation instruction for the smart washing machine, it is expected that the energy consumption of the washing machine during a complete washing process is 0.5 kWh, but the actual energy consumption reaches 0.7 kWh, and exceeds the preset fault tolerance threshold of 0.1 kWh. At this time, the real-time regulation execution module will immediately generate an execution abnormal signal and feedback this signal to the supply-demand collaborative optimization module. After receiving the feedback, the supply-demand collaborative optimization module will re-evaluate and adjust the current energy distribution plan. It may re-analyze the operation data of the washing machine to check for equipment failures; or re-optimize the regulation strategy, adjust the working mode and operation parameters of the washing machine, and then generate a new regulation instruction and send it to the real-time regulation execution module to ensure that the system can correct the abnormal situation in time, achieve the expected energy optimization effect, and ensure the stable and efficient operation of the smart home energy supply-demand collaborative optimization system. Example 5
[0054] The system performance monitoring module continuously collects various types of data during the system operation and communicates with the dynamic policy evaluation module in real time. The dynamic policy evaluation module, based on this data, collects the policy execution efficiency, abnormal signal frequency, and user satisfaction indicators in the historical regulation data to comprehensively evaluate the advantages and disadvantages of the current energy regulation policy.
[0055] The policy execution efficiency is used to measure the closeness between the actual energy-saving effect of the system and the theoretical optimal energy-saving effect. Among them, the actual energy-saving rate is obtained by comparing the household energy consumption before and after implementing the regulation policy. Suppose that within a certain period of time before implementing the regulation policy, the total household energy consumption is kWh; within the same period after implementing the regulation policy, the total household energy consumption is kWh. Then the actual energy-saving rate is calculated by the formula: . The theoretical maximum energy-saving rate is the maximum energy-saving ratio that can be achieved under ideal conditions by making full use of household energy resources and optimizing the equipment operation mode, denoted as . Therefore, the policy execution efficiency is the ratio of the actual energy-saving rate to the theoretical maximum energy-saving rate, that is . For example, after a period of monitoring, it is found that the monthly electricity consumption of a certain household before optimization is kWh, and the monthly electricity consumption after implementing the regulation policy drops to kWh, while the household can theoretically achieve a maximum energy saving. Then the actual energy-saving rate , and the policy execution efficiency .
[0056] The abnormal signal frequency reflects the frequency of abnormal situations occurring during the system operation. It is marked as the number of abnormal signals generated per unit time. For example, during a day of operation, the system generated a total of abnormal signals. Then the abnormal signal frequency on this day is times / day. Abnormal signals may come from data anomalies detected by the energy data acquisition module, such as a sudden large fluctuation in equipment power beyond the normal range; or they may result from execution anomalies feedback by the real-time regulation execution module, such as a too large deviation between the actual energy consumption and the expected value after the regulation instruction is executed. The higher the abnormal signal frequency, the more problems there are during the system operation, indicating that there may be defects in the current policy.
[0057] User satisfaction is an important indicator to measure whether the system meets user needs, and it is scored based on device availability and environmental comfort. Device availability mainly examines whether the commonly used devices of users can operate normally when needed. If the devices frequently fail to start or operate unstably due to energy regulation, the user's evaluation of device availability will decrease. Environmental comfort is related to environmental parameters such as indoor temperature and humidity. When the indoor temperature and humidity deviate from the human comfort range for a long time, users will feel uncomfortable, thus affecting the evaluation of environmental comfort. Assume that the full score of user satisfaction is points, and according to the comprehensive performance of device availability and environmental comfort, the user's score for the system is points. For example, during the use process, a certain user found that the air conditioner and water heater at home occasionally failed to start normally, and the indoor temperature was often too high in summer. After comprehensive consideration, the user gave the system a score of points, that is .
[0058] The dynamic policy evaluation module calculates the weights of each index through the entropy weight method, and then performs weighted summation on these indexes to obtain the policy evaluation value . The entropy weight method is an objective weighting method, which determines the weights according to the dispersion degree of each index data. The greater the dispersion degree, the greater the role of this index in the comprehensive evaluation, and the higher the weight. Assume that the weight of policy execution efficiency is , the weight of the abnormal signal frequency is , the weight of user satisfaction is , and , then the calculation formula of the policy evaluation value is: . For example, after calculation by the entropy weight method, , , , combined with the previously calculated , times per day (assume that after dimensionless processing, it is ), , then the policy evaluation value .
[0059] If the policy evaluation value is lower than the preset evaluation threshold , it indicates that the current energy regulation policy has poor effects and may not be able to meet the optimization goals of the system. At this time, the dynamic policy evaluation module will trigger the policy reset function of the reinforcement learning decision module. The system will re-examine and adjust the policy network parameters, and explore new energy regulation policies to improve the overall performance of the system. If for consecutive evaluation cycles ( is a preset integer constant, for example If the policy evaluation values are all lower than the threshold, it is determined that the current policy fails, and the system will re-plan and optimize the policy more comprehensively to ensure the continuous and stable operation of the smart home energy supply and demand collaborative optimization system. Embodiment 6
[0060] The main task of the abnormal mode recognition module is to identify abnormal situations in the household energy consumption process and ensure the safety and stability of energy use. It constructs a normal energy consumption pattern library by analyzing the periodic fluctuations and sudden peaks in historical energy consumption data. During the historical energy consumption data collection stage, the system continuously records the electricity consumption data of various household energy devices, and the time span can be several months or even years. For example, record the electricity consumption of devices such as air conditioners, refrigerators, and TVs at different times of each day. By deeply analyzing these data, the periodic patterns among them are mined. Taking the air conditioner as an example, in summer, from noon to evening every day, due to the high temperature, the usage frequency and electricity consumption of the air conditioner will show an obvious periodic upward trend; while at night, as the temperature drops, the electricity consumption will gradually decrease. Sudden peak data cannot be ignored either. For example, when using high-power devices such as electric heaters in the household, it may cause a sudden large increase in electricity consumption, forming a sudden peak.
[0061] Using these analysis results, the abnormal mode recognition module constructs a normal energy consumption pattern library. Each template curve in the library represents a normal energy consumption pattern, including the energy consumption range and change trend in different time periods. During the operation of the system, the abnormal mode recognition module real-time collects energy consumption data and performs similarity matching with the pattern library. The dynamic time warping algorithm is used to calculate the distance between the real-time energy consumption curve and each template curve in the pattern library, and the minimum distance value is marked as the matching degree. Suppose the real-time energy consumption curve is , and a certain template curve in the pattern library is , the dynamic time warping algorithm finds a time warping path such that and After non-linearly aligning on the time axis, the distance between them is the smallest. This minimum distance is the matching degree. For example, after calculation, the minimum distance between the real-time energy consumption curve and a certain template curve in the pattern library is , if the preset similarity threshold is , , then the matching degree is lower than the threshold, and it is determined that an abnormal energy consumption situation has occurred.
[0062] When the matching degree is lower than the threshold, the abnormal pattern recognition module will extract abnormal features, including peak duration, fluctuation frequency, and device correlation. The peak duration refers to the length of time that the abnormal peak lasts when the energy consumption shows an abnormal peak. For example, in a certain abnormal power consumption situation, when a high-power device is turned on, the power consumption peak lasts for [X] minutes, and these [X] minutes are the peak duration. The fluctuation frequency reflects the frequency of fluctuations in the energy consumption data during the abnormal period, which is determined by counting the number of fluctuations in the energy consumption data per unit time. The device correlation analyzes which devices' power consumption changes are related to the abnormal situation. For example, during a certain abnormal period, it is found that the air conditioner and the electric water heater simultaneously show abnormal power consumption fluctuations, indicating that there may be a certain correlation between them. Based on these feature combinations, an abnormal type code is generated, and different abnormal type codes correspond to different preset processing strategies. For example, if the peak duration is long and the fluctuation frequency is low, it may be caused by a failure of a certain high-power device. The system will generate the corresponding abnormal type code, and send a related notification to the user to check the device, and at the same time limit the use of the device and other processing strategies.
[0063] The multi-agent collaboration module models each energy device in the home as an independent agent to achieve the collaboration of local decision-making and global optimization. For example, devices such as air conditioners, refrigerators, and washing machines are regarded as different agents respectively, and each agent has its own decision-making ability and goal. The goal of the air conditioner agent may be to keep the indoor temperature comfortable while minimizing its own energy consumption; the refrigerator agent needs to maintain a low temperature environment inside to ensure food freshness, and reasonably arrange its own start and stop times to save energy. These agents carry out information interaction and collaborative decision-making through a distributed reinforcement learning algorithm. The distributed reinforcement learning algorithm allows each agent to collect environmental information, execute actions, and obtain rewards locally, while sharing some information with other agents, so as to achieve the optimization of the overall system.
[0064] During the actual operation process, there may be a situation where the local strategy of a single agent conflicts with the global goal. For example, during the peak power consumption period, the air conditioner agent may continuously operate at a high power to keep the indoor temperature comfortable, which conflicts with the global goal of the system to reduce the overall power consumption load. At this time, the multi-agent collaboration module will reallocate the optimization weights through a game theory model. The game theory model considers the interests and decisions of each agent, and by establishing a game scenario, allows the agents to make strategy choices in it. During this process, each agent needs to balance its own goal and the global goal. For example, through the calculation of the game theory model, the optimization weight of the air conditioner agent is adjusted, so that it appropriately reduces the power operation on the premise of ensuring the basic indoor comfort, to meet the overall energy optimization requirements of the system, achieve better collaboration between local decision-making and global optimization, and improve the overall performance of the smart home energy supply and demand collaborative optimization system.
[0065] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0066] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A smart home energy supply and demand collaborative optimization system based on reinforcement learning, characterized in that: It includes energy data acquisition module, reinforcement learning decision module, supply and demand collaborative optimization module, real-time control execution module and system performance monitoring module; The energy data collection module is used to collect household energy supply and demand data, including power of electrical equipment, renewable energy generation, energy storage equipment status and environmental parameters; The reinforcement learning decision module constructs a state space based on the collected data and generates an energy control strategy in combination with the preset action space and reward function; The supply-demand collaborative optimization module performs multi-objective optimization on energy distribution according to the energy control strategy and generates real-time control instructions; The real-time control execution module adjusts the operating state of the power-consuming equipment, the charging and discharging strategy of the energy storage equipment and the energy transmission path according to the real-time control instruction; The system performance monitoring module tracks the operating status of each module in real time, collects abnormal deviation signals during the control process, and sends the abnormal deviation signals to the reinforcement learning decision module for strategy iterative update.
2. According to claim 1, a smart home energy supply and demand collaborative optimization system based on reinforcement learning is characterized in that: The specific operation process of the energy data acquisition module is as follows: During the data collection stage, the instantaneous power of electrical equipment, the power generated by renewable energy, the remaining power of energy storage equipment, and the ambient temperature and humidity data are obtained through the sensor network; the deviation value between the instantaneous power and the preset power threshold is marked as the power deviation value, the deviation value between the remaining power and the preset power threshold is marked as the energy storage surplus deviation value, and the difference between the ambient temperature and humidity and the preset comfort range is marked as the environmental deviation value; if the power deviation value, the energy storage surplus deviation value or the environmental deviation value exceeds the corresponding preset threshold, a data abnormality signal is generated and sent to the system performance monitoring module.
3. According to the reinforcement learning-based smart home energy supply and demand collaborative optimization system of claim 1, it is characterized in that: The specific decision-making process of the reinforcement learning decision-making module is as follows: The state space is defined as a set of vectors including electricity demand, energy supply, energy storage status and time characteristics; the action space includes equipment start and stop instructions, energy storage charging and discharging rates and energy trading strategies; the reward function is based on multi-dimensional weighted calculations based on energy costs, user comfort and carbon emissions; the strategy network parameters are updated through the Q learning algorithm. If the accumulated reward value during the strategy iteration is lower than the preset convergence threshold, a strategy optimization signal is generated and the action space expansion is triggered.
4. According to claim 1, a smart home energy supply and demand collaborative optimization system based on reinforcement learning is characterized in that: The specific optimization process of the supply-demand collaborative optimization module is as follows: The energy distribution problem is modeled as a multi-objective optimization model. The objective functions include minimizing the power purchase cost of the power grid, maximizing the utilization rate of renewable energy and balancing the life of energy storage equipment. A non-dominated sorting genetic algorithm is used to generate a Pareto optimal solution set, and the optimal allocation plan is selected from the solution set based on real-time electricity price signals and user preferences. If there is no feasible solution in the solution set, an optimization failure signal is generated and the system switches to the backup scheduling mode.
5. According to the reinforcement learning-based smart home energy supply and demand collaborative optimization system of claim 1, it is characterized in that: The specific execution process of the real-time control execution module is as follows: According to the priority of the control instructions, the instructions are divided into emergency control instructions, regular control instructions and delayed control instructions; emergency control instructions directly interrupt the operation of high-power consumption equipment, regular control instructions achieve progressive optimization by adjusting the equipment working mode, and delayed control instructions are executed during non-peak hours based on user behavior prediction; if the actual energy consumption after the execution of the instruction deviates from the expected value by more than the preset fault tolerance threshold, an execution exception signal is generated and fed back to the supply and demand collaborative optimization module.
6. According to the reinforcement learning-based smart home energy supply and demand collaborative optimization system of claim 1, it is characterized in that: The system performance monitoring module is communicatively connected with the dynamic strategy evaluation module. The dynamic strategy evaluation module collects strategy execution efficiency, abnormal signal frequency and user satisfaction indicators in historical control data, and weights and fuses them into a strategy evaluation value. If the strategy evaluation value is lower than a preset evaluation threshold, the strategy reset function of the reinforcement learning decision module is triggered.
7. The smart home energy supply and demand collaborative optimization system based on reinforcement learning according to claim 6 is characterized in that: The specific evaluation process of the dynamic strategy evaluation module is as follows: The strategy execution efficiency is defined as the ratio of the actual energy saving rate to the theoretical maximum energy saving rate. The abnormal signal frequency is marked as the number of abnormal signal generation per unit time. The user satisfaction is scored based on equipment availability and environmental comfort. The weight of each indicator is calculated by the entropy weight method, and the strategy evaluation value is obtained by weighted summation. If the strategy evaluation value is lower than the threshold for N consecutive evaluation cycles, the current strategy is deemed invalid, where N is a preset integer constant.
8. According to claim 1, a smart home energy supply and demand collaborative optimization system based on reinforcement learning is characterized in that: It also includes an abnormal pattern recognition module, which constructs a normal energy consumption pattern library by analyzing the periodic fluctuations and sudden peaks in the historical energy consumption data; the real-time collected energy consumption data is matched with the pattern library for similarity. If the matching degree is lower than the preset similarity threshold, an energy consumption abnormality signal is generated and the manual intervention process is triggered.
9. The smart home energy supply and demand collaborative optimization system based on reinforcement learning according to claim 8, characterized in that: The specific recognition process of the abnormal pattern recognition module is as follows: The dynamic time warping algorithm is used to calculate the distance between the real-time energy consumption curve and each template curve in the pattern library, and the minimum distance value is marked as the matching degree; if the matching degree is lower than the threshold, the abnormal features including peak duration, fluctuation frequency and equipment correlation are extracted, and the abnormal type code is generated according to the feature combination and associated with the preset processing strategy.
10. The smart home energy supply and demand collaborative optimization system based on reinforcement learning according to claim 1, characterized in that: It also includes a multi-agent collaboration module, which models each energy device in the home as an independent agent, and realizes the coordination of local decision-making and global optimization through a distributed reinforcement learning algorithm; if the local strategy of a single agent conflicts with the global goal, the optimization weight is redistributed through a game theory model.
Citation Information
Patent Citations
Micro-grid energy-saving scheme generation method and system based on energy storage optimization scheduling
CN119209504A
Plateau multi-energy complementary agricultural energy management system and method
CN119543159A
Dynamic balance system for wind, light, fire and nuclear storage integrated regulation and control of power grid in cold region
CN119695852A
Virtual power plant peak regulation optimization scheduling method and system, electronic equipment and medium
CN119783997A
Automatic cooperative regulation and control system and method for smart home equipment
CN119846984A
Cited By
Intelligent park source network load storage and charging integrated scheduling method based on AI
CN120474006A
An AI-based intelligent park source-grid-load-storage-charging integrated scheduling method
CN120474006B
Intelligent flower shed intelligent monitoring method based on multi-dimensional environment perception and cooperative control
CN120508176A
Networked loom intelligent control method and system based on process knowledge software
CN120972828A
Smart home energy consumption optimization control method and system
CN121523082A