Household energy low-carbon operation strategy based on block chain and hybrid evolution reinforcement learning

Through the home energy system model of blockchain and hybrid evolutionary reinforcement learning, combined with active and reactive power optimization, the problems of reactive power neglect and local optimality in home energy systems are solved, low-carbon and efficient energy management is achieved, and data transparency and global optimization capabilities are improved.

CN120595572APending Publication Date: 2025-09-05SOUTH CHINA UNIV OF TECH +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510540966.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies fail to effectively regulate reactive power in home energy systems, ignore carbon emissions and user comfort costs, and reinforcement learning algorithms are prone to falling into local optimality and lack exploration capabilities.

Method used

A home energy system model based on blockchain and hybrid evolutionary reinforcement learning is adopted, combined with active and reactive power optimization. Data is collected in real time through smart meters and uploaded to the blockchain ledger. A Markov decision process and hybrid evolutionary reinforcement learning algorithm are designed to optimize equipment scheduling to minimize electricity costs, carbon emissions and improve user comfort.

Benefits of technology

It achieves low-carbon and efficient management of home energy systems, takes into account multiple optimization objectives, solves the problems of reactive power neglect, carbon emissions and local optimality, and improves data transparency and global optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595572A_ABST
    Figure CN120595572A_ABST
Patent Text Reader

Abstract

According to the household energy low-carbon operation strategy based on the block chain and the hybrid evolution reinforcement learning, output of the household equipment is controlled through the intelligent contract, and low-carbon and efficient energy management is achieved. The system obtains environment and power data of each device through an intelligent electric meter, and ensures that the data is transparent and cannot be tampered based on blockchain account book storage. A home system model covering multiple load types and optimization targets is provided, an energy scheduling problem is converted into a Markov decision process, and automatic optimization scheduling of equipment is realized in combination with a block chain smart contract and hybrid evolution reinforcement learning so as to minimize power consumption cost and carbon emission and improve comfort and power factors. In addition, low-carbon power utilization of the user is motivated through a carbon emission integral and cost reward mechanism. The result shows that the power utilization cost and carbon emission can be effectively reduced, the comfort level and the power factor are considered, and a safe, transparent and intelligent solution is provided for carbon emission reduction and household energy management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power systems, and in particular to a low-carbon strategy for home energy systems based on blockchain and hybrid evolutionary reinforcement learning. Background Art

[0002] With energy shortages and environmental concerns, carbon emission reduction and energy conservation have become core goals of global sustainable development. As a key vehicle for distributed energy management, home energy systems have incorporated photovoltaics, energy storage, and other devices, enabling energy conservation and emission reduction by optimizing appliances and energy scheduling. Traditional energy management methods are often criticized for their cumbersome explicit mathematical models and their inability to adapt to dynamic changes and fluctuations in distributed energy resources, resulting in limited optimization results. Reinforcement learning, a significant shift in approach, has emerged as a key tool for energy optimization due to its lack of prior knowledge and its ability to autonomously learn under uncertainty. However, in practical applications, it is prone to falling into local optima and lacks exploration capabilities. Therefore, it can be combined with evolutionary strategies to enhance global exploration capabilities and improve optimization results. Furthermore, home energy management processes involve large amounts of user data, posing significant challenges to data privacy and security. Blockchain technology, with its decentralized, immutable nature and the support of smart contracts, provides technical support for data transparency and security, enabling distributed electricity usage logging, carbon emissions monitoring, carbon credit rewards, and energy trading. The low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning achieves dual optimization of electricity costs and carbon emissions by optimizing equipment scheduling, load management and energy supply and demand balance, providing an innovative path for the intelligent and low-carbon development of household energy systems.

[0003] Chinese invention patent: A method for optimizing and controlling the energy of a community household microgrid based on deep reinforcement learning (publication number: CN11828705B) discloses a method for optimizing and controlling the energy of a community household microgrid based on blockchain and deep reinforcement learning. The specific contents are as follows: 1) Establish a blockchain community household energy management system network model and collect the first energy consumption of the household microgrid that has been put on the chain. 2) Based on the blockchain community household energy management system network model, construct a Markov decision process; use deep reinforcement learning to solve the Markov decision process to obtain the optimal intelligent agent strategy of the household microgrid that has been put on the chain. 3) Obtain the second energy consumption of the household microgrid to be optimized and regulated, calculate the similarity between the second energy consumption and all the first energy consumptions, and migrate the optimal intelligent agent strategy of the household microgrid corresponding to the first energy consumption with the highest similarity to the household microgrid to be optimized and regulated, and the household microgrid to be optimized and regulated uses the migrated optimal intelligent agent strategy to optimize and control the energy of the household microgrid. It has the following technical deficiencies:

[0004] 1) This invention focuses on optimizing the active power of the home energy system while neglecting the regulation of reactive power, which will lead to the degradation of the power factor;

[0005] 2) This invention uses blockchain to conduct private energy transactions to reduce electricity costs, but does not address carbon emission costs and user comfort costs, nor does it establish a corresponding blockchain transaction framework;

[0006] 3) This invention uses conventional reinforcement learning as an optimization algorithm and does not discuss the local optimality of the reinforcement learning algorithm or the problem of exploring diversity; Summary of the Invention

[0007] To overcome the shortcomings of existing technologies, this paper proposes a low-carbon strategy for home energy systems based on blockchain and hybrid evolutionary reinforcement learning. This paper proposes a home energy system model that considers both active and reactive power and transforms it into a Markov decision process. A blockchain-based home energy system optimization framework is designed, incorporating electricity costs, carbon credits, and user comfort, and providing a path for private data transmission. To address the local optimality problem in reinforcement learning, an optimization algorithm based on hybrid evolutionary reinforcement learning is proposed to address the challenges of long-term reporting, local optimality for multiple types of home devices and optimization objectives, and insufficient strategy exploration diversity.

[0008] In order to solve the above problems, the present invention provides the following technical solutions:

[0009] The low-carbon operation strategy of household energy based on blockchain and hybrid reinforcement learning includes the following steps:

[0010] S1: Analyze the operating characteristics of household appliances and classify them into four types: fixed appliances, time-adjustable appliances, power-adjustable appliances, and energy storage appliances, and then construct a home energy system model;

[0011] S2: Intrusive smart meters collect real-time environmental data and power consumption data of each device, including the active and reactive power of the device, as well as the start and stop status of time-adjustable devices.

[0012] S3: Calculate household electricity carbon emissions (carbon credits) based on collected electricity consumption data and combined with the grid carbon emission factor. Upload electricity consumption data and carbon emission data to the blockchain ledger through the IoT gateway to ensure data transparency and immutability.

[0013] S4: Considering the characteristics of each household device, the home energy system optimization problem in S1 is transformed into a Markov decision process. The reward function in the Markov decision process includes electricity cost, comfort level, and carbon credits.

[0014] S5: Under the guidance of the reward function, a soft actor-critic (SAC) optimization model based on hybrid evolutionary reinforcement learning is trained as the decision model for the home energy system;

[0015] S6: Predefine smart contract trigger conditions, including reincarnation triggers and abnormal event triggers. When the trigger conditions are met, the current ledger data is accessed and the home energy management system is activated to run the decision model and generate the latest scheduling plan.

[0016] Preferably, the models of different types of household appliances in step S1 are as follows:

[0017] S11. Fixed devices: Fixed devices refer to devices with fixed power and operating time. They are rigid demands of users and cannot be scheduled by the home energy management system. They are represented by rooftop photovoltaics, lights, televisions, and refrigerators. At time t, the rooftop photovoltaic power output is defined as The active power of the i-th fixed household appliance is The reactive power of the i-th fixed household appliance is Among them, i∈[1,…,N f ],N f is the number of fixed devices.

[0018] S12. Time-adjustable devices: Time-adjustable devices can postpone the start time of the device without affecting the user's comfort. Once turned on, such devices will continue to run at a fixed power until the designated task is completed, and they cannot be interrupted. Common time-adjustable devices are dishwashers and washing machines. In addition, in order to ensure user comfort in actual scenarios, it is necessary to define the prohibited time period for such devices. For example, a dishwasher cannot be run during breakfast time. Let the active power of the jth time-adjustable device be Reactive power is The prohibited driving period is The continuous running time required to complete the task is Among them, j∈[1,…,N d ],N d is the number of time-adjustable devices, The starting time of the device prohibition time, is the end time of the device prohibition time; therefore, at time t, the operating state of the jth time-adjustable device can be defined as:

[0019]

[0020] Where, is the operating status of the time-adjustable equipment; Δt is the scheduling time interval, T P is the scheduling period; is the start and stop status of the device at time t, The start and stop status of the device at time t-1. When the device is on, it is 1, otherwise it is 0. Considering that such devices cannot be interrupted, The following constraints must be met:

[0021]

[0022] S13. Power-adjustable devices: Taking air conditioners as an example, power-adjustable devices can adjust power output to maintain a comfortable indoor temperature. Their operating status at time t can be characterized by the outdoor temperature and indoor temperature:

[0023]

[0024] in, The operating status of the air conditioning equipment. is the outdoor temperature, is the indoor temperature. Based on the thermal dynamic model, the indoor temperature adjusted by the air conditioner is expressed as:

[0025]

[0026] in, is the indoor temperature at time t+1; η ac represents the heat transfer coefficient, ξ ac represents the coefficient of inertia, and F represents the average thermal conductivity; is the active power output of the air conditioner, and the reactive power of the air conditioner is recorded as Both need to satisfy the following constraints:

[0027]

[0028] in, is the rated active power of the air conditioner, is the rated reactive power of the air conditioner.

[0029] S14. Energy storage devices: Energy storage devices are an important component of home energy systems. They can switch roles between energy providers and consumers based on electricity price fluctuations to reduce electricity bills. Energy storage devices include home batteries and electric vehicles, both of which are equipped with inverters to provide both active and reactive power. The operating status of both can be characterized by the energy state of the battery. The operating state of the home battery at time t is:

[0030]

[0031] in, The operating status of the household battery; is the energy state of the household battery at time t, is the energy state of the household battery at time t+1; is the charging and discharging power of household batteries; For the charging efficiency of household batteries, is the discharge efficiency of the household battery. In order to extend the service life of the household battery, the following constraints must also be met:

[0032]

[0033] in, The reactive power of the household battery and is the apparent power of the household battery; is the maximum apparent power of the household battery; Corresponding to the minimum battery energy state, Corresponding to the maximum battery energy state. As a mobile energy storage device, the dynamic model of electric vehicles is similar to that of household batteries, except that electric vehicles only operate at specific times. Connect to home system. The time it takes for the electric car to arrive home. is the time when the electric vehicle leaves home. The operating state model of the electric vehicle is expressed as:

[0034]

[0035] Where, is the operating status of the electric vehicle; is the energy state of the electric vehicle at time t, is the energy state of the electric vehicle at time t+1; is the charging and discharging power of the electric vehicle; For the charging efficiency of electric vehicles, is the discharge efficiency of electric vehicles. is the reactive power of the electric vehicle, is the apparent power of the electric vehicle; is the maximum apparent power of the electric vehicle; Corresponding to the minimum battery energy state of electric vehicles, Corresponding to the maximum battery energy state of electric vehicles.

[0036] Preferably, the environmental data collected in real time in step S2 includes the outdoor temperature and indoor temperature Photovoltaic power output Electricity consumption data including active power of fixed equipment and reactive power Active power of time-adjustable equipment and reactive power Start / Stop Status No-driving time and the duration required to complete the task Active power of power-adjustable equipment and reactive power Active power of household batteries and reactive power of household batteries and the state of charge of household batteries Electric car home time Active power of electric vehicles and reactive power of electric vehicles and the state of charge of electric vehicles

[0037] Preferably, the calculation formula of carbon emissions and the process of uploading electricity consumption data and carbon emissions data in step S3 are as follows:

[0038] Total household active load at time S31.t and total reactive load Expressed as:

[0039]

[0040] The carbon emissions of a household energy system are converted from the household’s net active power:

[0041]

[0042] Where, is the net active power, is the carbon emission factor, For carbon emissions.

[0043] S32. The IoT gateway encrypts and signs the formatted electricity consumption and carbon emission data, uploads it to the blockchain via the HTTP protocol, and calls the smart contract to store the data in the blockchain ledger, achieving tamper-proof data records.

[0044] Preferably, the optimization scheduling problem of the home energy system model in step S4 is regarded as a Markov decision process in a finite time, which can be represented by a four-tuple [S t ,A t ,S t+1 ,R t ] to describe. In the quaternion, S t is the current state space (representing the operating state of the system at time t); A t represents the current action space; S t+1 is the state space at the next moment; R t is the immediate reward of the system after the current action. The detailed definition of the elements in the four-tuple is as follows:

[0045] S41. Current state space: To fully describe the current system state, the operating states of all devices are included in the state space set. Considering the non-dispatchability of fixed devices, the sum of the active power and reactive power of all fixed devices is used to represent their operating state. In addition, as a stimulus signal, electricity price information is also included in the system state. Then, S t Can be formulated as:

[0046]

[0047] Among them, p t is the electricity price at time t; and It is the sum of active power and reactive power of fixed equipment; This is the operating status of the first time-adjustable device. For Nth d The operating status of a time-adjustable device;

[0048] S42. Action space: The time-adjustable device includes the controllable parameters of the adjustable device, such as the start and stop status of the time-adjustable device and the active and reactive power of the power-adjustable device.

[0049]

[0050] in, It is the start and stop status of the first time-adjustable device; For Nth d The start and stop status of the device can be adjusted at a time.

[0051] S43. Next-moment state space: The elements of the next-moment state space are the same as the elements of the current state space, but are obtained after the current action is executed. It is represented as:

[0052]

[0053] Among them, the photovoltaic output at the next moment Active power of fixed equipment Reactive power and electricity price p t+1 Obtained by prediction; from the 1st time adjustable device to the Nth d The operating status of the time-adjustable device The operating state of the power adjustable device at the next moment is obtained by formula (1) and (2). Calculated by formulas (3) and (4); the operating state of the household battery at the next moment and the operating status of electric vehicles Calculated by formulas (6) and (10) respectively.

[0054] S44. Immediate Rewards: The goal of optimizing a home energy system is to minimize electricity costs, carbon emissions, and improve user comfort. Based on this, immediate rewards are defined as

[0055]

[0056] Where, is the electricity price cost, is the carbon emission cost (i.e. carbon credits), is the comfort cost; α price is the electricity price cost coefficient, α carbon is the carbon emission cost coefficient and α com is the comfort cost coefficient. Specifically, the electricity price cost, carbon emission cost and comfort cost are expressed as

[0057]

[0058] Among them, C t is the carbon price; The preset optimal indoor temperature.

[0059] In particular, to avoid the range anxiety of electric vehicles and maintain the power factor, a sufficiently large penalty is attached to R t :

[0060]

[0061] in, It is the energy state of the electric car when it leaves home; It is the preset minimum energy state to avoid range anxiety. PF t PF is the power factor; set is the preset minimum power factor.

[0062] Preferably, the SAC optimization process based on hybrid evolutionary reinforcement learning in step S5 is as follows:

[0063] The S51.SAC algorithm is implemented using five neural networks, including an action network, two value networks, and two target evaluation networks. The training process consists of five steps.

[0064] Step 1: Initialize parameters. Initialize the parameters θ of the action network actor and the parameters θ1 of the first value network and the parameters θ2 of the second value network, as well as the hyperparameters λ, γ, β and τ, where λ is the learning rate, γ is the discount factor, τ is the soft update coefficient; β is the regularization coefficient; the initialization parameters of the two target evaluation networks are copied θ1 and θ2;

[0065] Step 2: Interact with the environment; the action network interacts with the environment and generates the mean of the action and standard deviation For continuous action space, the current action is obtained by Gaussian sampling:

[0066]

[0067] tanh(·) is the hyperbolic tangent function; S is the state space set, A is the action space set; ε is Gaussian noise; For actor network function; represents the standard normal distribution; for discrete actions, a binary sampling strategy is used to select the current action:

[0068]

[0069] Step 3: Update the value network. Randomly select a batch of data B = {(S t ,A t ,S t+1 ,R t )} as training data. The action network generates the next action, and the target value network is responsible for calculating the Q value y corresponding to the current action:

[0070]

[0071] in, is the first target value network function, V θt2 (·) is the second target value network function, θ t1 is the parameter of the first target value network, θ t2 are the parameters of the second target value network. The optimization goal of the value network is to minimize the distance between the predicted Q values ​​of the two value networks and the Q value of the target value network. Therefore, the loss function can be defined as:

[0072]

[0073] is the loss function of the first value network, is the loss function of the second value network; is the first value network function, is the second value network function; gradient descent is used to update the value network parameters:

[0074]

[0075] Step 4: Update the action network. The goal of the action network is to maximize the value function and policy randomness, and its loss function is formulated as

[0076]

[0077] Where, is the expected function. Similar to the value network, the parameters of the action network are also updated using the gradient descent algorithm:

[0078]

[0079] Step 5: Update the target value network. The target value network is a reflection of the value network and is updated at a slower rate to ensure stable training.

[0080] θ t1 ←τθ t1 +(1-τ)θ t1

[0081] θ t2 ←τθ t2 +(1-τ)θ t2

[0082] S52. Hybrid evolutionary SAC uses evolutionary strategy to search for the global optimal value of the SAC model, and then uses the gradient descent algorithm to finely optimize the parameters. The process can be divided into two steps.

[0083] Step 1: Evolutionary Strategy Optimization. To avoid local optimality in gradient optimization, Gaussian noise is introduced into the current action network parameters, thereby generating multiple candidate parameter sets:

[0084]

[0085] Where k∈[1,…,M] is the index of the candidate parameter, and M is the population size of the generated action parameters; is the exploration scale of the kth candidate parameter, is the sampling noise of the kth candidate parameter.

[0086] Step 2: The cumulative reward of each candidate parameter is determined by the fitness F of the interaction with the environment k To calculate:

[0087]

[0088] Where r represents the state and action trajectory contained in a cycle. The evolutionary strategy optimization update relies on fitness ranking rather than gradient update. Its parameter update process is calculated as follows:

[0089]

[0090] Where, is the average fitness of the entire population.

[0091] Preferably, the triggering conditions of the smart contract in step S6 are as follows:

[0092] The triggering of the home energy system decision model includes two situations: cyclic triggering and abnormal event triggering. Cyclic triggering is a periodic operation, that is, the optimization scheduling is performed every 15 minutes to generate the latest scheduling plan. The preset household bill is The default carbon emissions are The preset user comfort limit is Based on this, abnormal events are divided into three types, including excessively high household bills, excessively high carbon emissions, and violations of user comfort. When the trigger conditions are met, the current ledger data is accessed, the home energy management program is activated, and the latest scheduling plan is generated.

[0093] This paper proposes a low-carbon home energy operation strategy based on blockchain and hybrid evolutionary reinforcement learning. This strategy uses smart contracts to control the output of household devices, achieving low-carbon and efficient energy management. The system uses smart meters to obtain environmental and device power data, which is stored in a blockchain-based ledger to ensure data transparency and immutability. A home system model covering multiple load types and optimization objectives is proposed. The energy scheduling problem is transformed into a Markov decision process. By combining blockchain smart contracts with hybrid evolutionary reinforcement learning, automatic optimal device scheduling is achieved to minimize electricity costs and carbon emissions while improving comfort and power factor. Furthermore, through a carbon emission credit and cost-reward mechanism, users are incentivized to use low-carbon electricity. Results show that this method can effectively reduce electricity costs and carbon emissions while also balancing comfort and power factor, providing a secure, transparent, and intelligent solution for carbon reduction and home energy management.

[0094] Beneficial effects: The present invention integrates the active and reactive power of the home energy system, solving the problem of neglecting reactive power in traditional methods. The output of home devices is controlled by smart contracts, and the environment and power data of each device are obtained by smart meters and stored in the blockchain account book to ensure that the data is transparent and cannot be tampered with. Combined with the data of the blockchain, through the carbon emission points and cost reward mechanism, users are encouraged to use low-carbon electricity, and the optimization goals of minimizing electricity costs, minimizing carbon emissions and maximizing user comfort are established, taking into account the multiple needs of the home energy system. At the same time, hybrid evolutionary reinforcement learning is introduced to optimize the model to manage dynamic and uncertain consumption behaviors, solving the problems of local optimality and insufficient exploration of traditional reinforcement learning. The proposal of this application enables the home energy management model to have the ability to consider reactive power, multiple optimization objectives, data privacy and global optimality, which can achieve more refined and comprehensive energy management. The present invention solves the lack of reactive power, carbon emissions and user comfort considerations in the existing home energy system model based on reinforcement learning, as well as the problem of local optimality. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0096] Figure 1 This is a flow chart of the low-carbon operation strategy of household energy based on blockchain and hybrid evolutionary reinforcement learning in the present invention. DETAILED DESCRIPTION

[0097] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0098] The low-carbon operation strategy of household energy based on blockchain and hybrid evolutionary reinforcement learning includes the following steps:

[0099] S1. Models of different types of home appliances are as follows:

[0100] S11. Fixed devices: Fixed devices refer to devices with fixed power and running time. They are rigid demands of users and cannot be scheduled by the home energy management system. They are represented by rooftop photovoltaics, lights, TVs, and refrigerators. Define time t as the rooftop photovoltaic power output The active power of the i-th fixed household appliance is The reactive power of fixed equipment is Among them, i∈[1,…,N f ],N f is the number of fixed devices.

[0101] S12. Time-adjustable devices: Time-adjustable devices can delay the start time of the device without affecting user comfort. Once turned on, such devices will continue to run at a fixed power until the designated task is completed, and they cannot be interrupted. Common time-adjustable devices are dishwashers and washing machines. In addition, to ensure user comfort in actual scenarios, it is necessary to define the prohibited time period for such devices, such as dishwashers cannot run at breakfast time. Let the active power of the jth time-adjustable device be The reactive power of the jth time adjustable device is The prohibited driving period is The continuous running time required to complete the task is Among them, j∈[1,…,N d ],Nd is the number of time-adjustable devices, The starting time of the device prohibition time, is the end time of the device prohibition time. Therefore, at time t, the operating state of the jth time-adjustable device can be defined as:

[0102]

[0103] Where, is the operating status of the time-adjustable equipment; Δt is the scheduling time interval, T P is the scheduling period; is the start and stop status of the device at time t, The start and stop status of the device at time t-1. When the device is on, it is 1, otherwise it is 0. Considering that such devices cannot be interrupted, The following constraints must be met:

[0104]

[0105] S13. Power-adjustable devices: Taking air conditioners as an example, power-adjustable devices can adjust power output to maintain a comfortable indoor temperature. Their operating status at time t can be characterized by the outdoor temperature and indoor temperature:

[0106]

[0107] in, The operating status of the air conditioning equipment. is the outdoor temperature, is the indoor temperature. Based on the thermal dynamic model, the indoor temperature adjusted by the air conditioner can be expressed as:

[0108]

[0109] in, is the indoor temperature at time t+1; η ac is the heat transfer coefficient, ξ ac is the inertia coefficient, F represents the average thermal conductivity; is the active power output of the air conditioner, and the reactive power is recorded as Both need to satisfy the following constraints:

[0110]

[0111] in, is the rated active power of the air conditioner and is the rated reactive power of the air conditioner.

[0112] S14. Energy storage devices: Energy storage devices are an important component of home energy systems. They can switch roles between energy providers and consumers based on electricity price fluctuations to reduce electricity bills. Energy storage devices include home batteries and electric vehicles, both of which are equipped with inverters to provide both active and reactive power. The operating status of both can be characterized by the energy state of the battery. The operating state of the home battery at time t is:

[0113]

[0114] in, The operating status of the household battery; is the energy state of the household battery at time t, is the energy state of the household battery at time t+1; is the charging and discharging power of household batteries; For the charging efficiency of household batteries, is the discharge efficiency of the household battery. In order to extend the service life of the household battery, the following constraints must also be met:

[0115]

[0116] in, is the reactive power of the household battery, is the apparent power of the household battery; is the maximum apparent power of the battery; is the minimum battery energy state, As a mobile energy storage device, the dynamic model of electric vehicles is similar to that of household batteries, except that electric vehicles only operate at specific times. Connect to home system. The time it takes for the electric car to arrive home. is the time when the electric vehicle leaves home. The operating state model of the electric vehicle can be expressed as:

[0117]

[0118] Where, is the operating status of the electric vehicle; is the energy state of the electric vehicle at time t, is the energy state of the electric vehicle at time t+1; is the charging and discharging power of the electric vehicle; Corresponding to the charging efficiency of electric vehicles, Corresponding to the discharge efficiency of electric vehicles. is the reactive power of the electric vehicle, is the apparent power of the electric vehicle; is the maximum apparent power of the electric vehicle; Corresponding to the minimum energy state of electric vehicles, Corresponding to the maximum energy state of electric vehicles.

[0119] S2. Real-time collected environmental data including outdoor temperature and indoor temperature Photovoltaic power output Electricity consumption data including active power of fixed equipment and reactive power of fixed equipment Active power of time-adjustable equipment and reactive power of time-adjustable devices Time-adjustable start and stop status of equipment No-travel time for time-adjustable devices and the duration required to complete the task Active power of power-adjustable equipment and reactive power of power-adjustable equipment Active power of household batteries and reactive power of household batteries and the energy status of household batteries Electric car home time Active power of electric vehicles and reactive power of electric vehicles and the energy status of electric vehicles

[0120] S3. The calculation formula for carbon emissions and the process for uploading electricity consumption data and carbon emissions data are as follows:

[0121] Total household active load at time S31.t and total reactive load It can be expressed as

[0122]

[0123] The carbon emissions of a household energy system are converted from the household’s net active power:

[0124]

[0125] Where, is the net active power, is the carbon emission factor, For carbon emissions.

[0126] S32. The IoT gateway encrypts and signs the formatted electricity consumption and carbon emission data, uploads it to the blockchain via the HTTP protocol, and calls the smart contract to store the data in the blockchain ledger, achieving tamper-proof data records.

[0127] Preferably, the optimization scheduling problem of the home energy system model in step S4 is regarded as a Markov decision process in a finite time, which can be represented by a four-tuple [S t ,A t ,S t+1 ,R t ] to describe. In the quaternion, S t is the current state space, representing the operating state of the system at time t; A t represents the current action space; S t+1 is the state space at the next moment; R t is the immediate reward of the system after the current action. The detailed definition of the elements in the four-tuple is as follows:

[0128] S41. Current state space: To fully describe the current system state, the operating states of all devices are included in the state space set. Considering the non-dispatchability of fixed devices, the sum of the active power and reactive power of all fixed devices is used to represent their operating state. In addition, as a stimulus signal, electricity price information is also included in the system state. Then, S t Can be formulated as:

[0129]

[0130] Among them, p t is the electricity price at time t; and It is the sum of active power and reactive power of fixed equipment; This is the operating status of the first time-adjustable device. For N d The operating status of a time-adjustable device;

[0131] S42. Action space: The time-adjustable device includes the controllable parameters of the adjustable device, such as the start and stop status of the time-adjustable device and the active and reactive power of the power-adjustable device.

[0132]

[0133] in, It is the start and stop status of the first time-adjustable device; For N d The start and stop status of the device can be adjusted at a time.

[0134] S43. Next moment state space: The elements of the next moment state space are the same as the elements of the current state space, but they are obtained after the current action is executed and can be expressed as

[0135]

[0136] Among them, the photovoltaic output at the next moment Active power of fixed equipment Reactive power and electricity price p t+1 Obtained by prediction; from the 1st time adjustable device to the Nth d The operating status of the time-adjustable device The operating state of the power adjustable device at the next moment is obtained by formula (1) and (2). Calculated by formulas (3) and (4); the operating state of the household battery at the next moment and the operating status of electric vehicles Calculated by formulas (6) and (10) respectively.

[0137] S44. Immediate Rewards: The goal of optimizing a home energy system is to minimize electricity costs, carbon emissions, and improve user comfort. Based on this, immediate rewards are defined as

[0138]

[0139] Where, is the electricity price cost, is the cost of carbon emissions, is the comfort cost; α price is the electricity price cost coefficient, α carbon is the carbon emission cost coefficient, α com is the comfort cost coefficient. Specifically, the electricity price cost, carbon emission cost and comfort cost can be expressed as

[0140]

[0141] Among them, C t is the carbon price; The preset optimal indoor temperature.

[0142] In particular, to avoid the range anxiety of electric vehicles and maintain the power factor, a sufficiently large penalty is attached to R t :

[0143]

[0144] in, It is the energy state of the electric car when it leaves home; It is the preset minimum energy state to avoid range anxiety. Is the power factor; PF set is the preset minimum power factor.

[0145] S5. The SAC optimization process based on hybrid evolutionary reinforcement learning is as follows:

[0146] The S51.SAC algorithm is implemented using five neural networks, including an action network, two value networks, and two target evaluation networks. The training process consists of five steps.

[0147] Step 1: Initialize parameters. Initialize the parameters θ of the action network actor The parameters θ1 of the first value network and θ2 of the second value network, as well as the hyperparameters λ, γ, β, and τ, are given. λ is the learning rate, γ is the discount factor, τ is the soft update coefficient, and β is the regularization coefficient. The initialization parameters of the two target evaluation networks are copied to θ1 and θ2.

[0148] Step 2: Interact with the environment. The action network interacts with the environment and generates the mean of the action and standard deviation For continuous action space, the current action is obtained by Gaussian sampling:

[0149]

[0150] tanh(·) is the hyperbolic tangent function; S is the state space set, A is the action space set; ε is Gaussian noise; represents the standard normal distribution; is the actor network function. For discrete actions, a binary sampling strategy is used to select the current action:

[0151]

[0152] Step 3: Update the value network. Randomly select a batch of data B = {(S t ,A t ,S t+1 ,R t )} as training data. The action network generates the next action, and the target value network is responsible for calculating the Q value y corresponding to the current action:

[0153]

[0154] in, is the first target value network function, is the second target value network function, θ t1 is the parameter of the first target value network, θ t2 is the parameter of the second target value network. The optimization goal of the value network is to minimize the distance between the predicted Q values ​​of the two value networks and the Q value of the target value network. Therefore, the loss function can be defined as:

[0155]

[0156] and are the loss functions of the first and second value networks respectively; is the first value network function, is the second value network function. Gradient descent is used to update the value network parameters:

[0157]

[0158] Step 4: Update the action network. The goal of the action network is to maximize the value function and policy randomness, and its loss function is formulated as

[0159]

[0160] Where, is the expected function. Similar to the value network, the parameters of the action network are also updated using the gradient descent algorithm:

[0161]

[0162] Step 5: Update the target value network. The target value network is a reflection of the value network and is updated at a slower rate to ensure stable training.

[0163] θ t1 ←τθ t1 +(1-τ)θ t1

[0164] θ t2 ←τθ t2 +(1-τ)θ t2

[0165] S52. Hybrid evolutionary SAC uses evolutionary strategy to search for the global optimal value of the SAC model, and then uses the gradient descent algorithm to finely optimize the parameters. The process can be divided into two steps.

[0166] Step 1: Evolutionary Strategy Optimization. To avoid local optimality in gradient optimization, Gaussian noise is introduced into the current action network parameters, thereby generating multiple candidate parameter sets:

[0167]

[0168] Where k∈[1,…,M] is the index of the candidate parameter, and M is the population size of the generated action parameters; is the exploration scale of the kth candidate parameter, is the sampling noise of the kth candidate parameter. In addition, the cumulative reward of each candidate parameter is determined by the fitness F of the interaction with the environment k To calculate.

[0169]

[0170] Where r represents the state and action trajectory contained in a cycle. The evolutionary strategy optimization update relies on fitness ranking rather than gradient update. Its parameter update process is calculated as follows:

[0171]

[0172] Where, is the average fitness of the entire population.

[0173] S6. The triggering conditions of the smart contract are as follows:

[0174] The triggering of the home energy system decision model includes two situations: cyclic triggering and abnormal event triggering. Cyclic triggering is a periodic operation, that is, the optimization scheduling is performed every 15 minutes to generate the latest scheduling plan. The preset household bill limit is The default carbon emission cap is The preset user comfort limit is Based on this, abnormal events are categorized into three types: excessive household bills, excessive carbon emissions, and user comfort violations. When the trigger conditions are met, the current ledger data is accessed, the home energy management program is activated, and the latest scheduling plan is generated.

[0175] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning, characterized by: The following steps are involved: S1: Analyze the operating characteristics of household appliances and classify them into four types: fixed appliances, time-adjustable appliances, power-adjustable appliances, and energy storage appliances, and then construct a home energy system model; S2: Intrusive smart meters collect real-time environmental data and power consumption data of each device, including active and reactive power of the device, as well as the start and stop status of time-adjustable devices; S3: Calculates household electricity carbon emissions based on collected electricity consumption data and the grid carbon emission factor. Uploads electricity consumption and carbon emission data to the blockchain ledger via the IoT gateway to ensure data transparency and immutability. S4: Considering the characteristics of each household device, the home energy system optimization problem in S1 is transformed into a Markov decision process; The reward function in the Markov decision process includes electricity cost, comfort and carbon credits; S5: Under the guidance of the reward function, a soft actor-critic optimization model based on hybrid evolutionary reinforcement learning is trained as a decision model for the home energy system; S6: Predefine smart contract trigger conditions, including reincarnation trigger and abnormal event trigger. When the trigger conditions are met, access the current ledger data and stimulate the home energy management system to run the decision model and generate the latest scheduling plan.

2. The low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning according to claim 1 is characterized in that: Models of different types of household appliances in step S1: Fixed devices: Fixed devices refer to devices with fixed power and running time. They are rigid demands of users and cannot be scheduled by the home energy management system. They are represented by rooftop photovoltaics, lights, TVs, and refrigerators. At time t, the rooftop photovoltaic power output is defined as The active power of the i-th fixed household appliance is The reactive power of the i-th fixed household appliance is Among them, i∈[1,…,N f ],N f is the number of fixed equipment; Time-adjustable devices: Time-adjustable devices delay the start time of the device without affecting user comfort. Once turned on, the time-adjustable device will continue to run at a fixed power until the specified task is completed, and it cannot be interrupted. Common time-adjustable devices include dishwashers and washing machines. In addition, to ensure user comfort in actual scenarios, it is necessary to define the prohibited time period for such devices, such as dishwashers cannot run during breakfast time. The active power of the jth time-adjustable device is and reactive power is The prohibited time period of the jth time-adjustable device is The continuous running time required to complete the task is Among them, j∈[1,…,N d ],N d is the number of time-adjustable devices, The starting time of the device prohibition time, is the end time of the device prohibition time; therefore, at time t, the operating state of the jth time-adjustable device can be defined as: Where, is the operating status of the time-adjustable equipment; Δt is the scheduling time interval, T P is the scheduling period; is the start and stop status of the device at time t, is the start / stop status of the device at time t-1. It is 1 when the device is on and 0 otherwise. Considering that such devices cannot be interrupted, The following constraints must be met: Power-adjustable devices: Taking air conditioners as an example, power-adjustable devices can adjust power output to maintain a comfortable indoor temperature. Their operating status at time t is characterized by the outdoor temperature and indoor temperature: in, The operating status of the air conditioning equipment. is the outdoor temperature, is the indoor temperature; based on the thermal dynamic model, the indoor temperature adjusted by the air conditioner is expressed as: in, is the indoor temperature at time t+1; η ac represents the heat transfer coefficient, ξ ac represents the coefficient of inertia, and F represents the average thermal conductivity; is the active power output of the air conditioner, and the reactive power is recorded as Both need to satisfy the following constraints: in, is the rated active power of the air conditioner and is the rated reactive power of the air conditioner; Energy storage devices: Energy storage devices are a crucial component of home energy systems, switching roles between energy providers and consumers based on electricity price fluctuations to reduce electricity bills. Energy storage devices include home batteries and electric vehicles, both equipped with inverters to provide both active and reactive power. The operating state of both can be characterized by the energy state of the battery. The operating state of the home battery at time t is: in, The operating status of the household battery; is the energy state of the household battery at time t, is the energy state of the household battery at time t+1; is the charging and discharging power of household batteries; For the charging efficiency of household batteries, is the discharge efficiency of household batteries. To extend the service life of household batteries, the following constraints must also be met: in, is the reactive power of the household battery, is the actual apparent power of the household battery; is the maximum apparent power of a household battery; is the minimum battery energy state, The maximum battery energy state; as a mobile energy storage device, the dynamic model of electric vehicles is similar to that of household batteries, except that electric vehicles only Connect to home systems; where The time it takes for the electric car to arrive home. is the time when the electric vehicle leaves home; the operating state model of the electric vehicle is expressed as: Where, is the operating status of the electric vehicle; is the energy state of the electric vehicle at time t, is the energy state of the electric vehicle at time t+1; is the charging and discharging power of the electric vehicle; For the charging efficiency of electric vehicles, is the discharge efficiency of electric vehicles; is the reactive power of the electric vehicle, is the apparent power of the electric vehicle; is the maximum apparent power of the electric vehicle; is the minimum battery energy state of an electric vehicle, It is the maximum battery energy state of an electric vehicle.

3. The low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning according to claim 2 is characterized in that: The environmental data collected in real time in step S2 includes the outdoor temperature and indoor temperature Photovoltaic power output Electricity consumption data including active power of fixed equipment and reactive power of fixed equipment Active power of time-adjustable equipment and reactive power of time-adjustable devices Start / Stop Status No-driving time and the duration required to complete the task Active power of power-adjustable equipment Reactive power of power-adjustable equipment Active power of household batteries and reactive power of household batteries and state of charge Electric car home time Active power of electric vehicles and reactive power of electric vehicles and the state of charge of electric vehicles 4. The low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning according to claim 3 is characterized in that: In step S3, the calculation formula for carbon emissions and the process of uploading electricity consumption data and carbon emissions data are as follows: Total household active load at time t and total reactive load Expressed as: The carbon emissions of a household energy system are converted from the household’s net active power: Where, is the net active power, is the carbon emission factor, is carbon emissions; The IoT gateway encrypts and signs the formatted electricity consumption and carbon emission data, uploads it to the blockchain via the HTTP protocol, and calls the smart contract to store the data in the blockchain ledger to achieve tamper-proof data records.

5. The low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning according to claim 4 is characterized in that: In step S4, the Markov decision process of the home energy system model is as follows: The optimal scheduling problem of home energy system is regarded as a finite-time Markov decision process, which can be represented by the four-tuple [S t ,A t ,S t+1 ,R t ] to describe; in the quaternion, S t is the current state space, representing the operating state of the system at time t; A t represents the current action space; S t+1 is the state space at the next moment; R t is the immediate reward of the system after the current action; the detailed definition of the elements in the quaternary is as follows: Current state space: In order to fully describe the current system state, the operating states of all devices are included in the state space set; considering the unschedulable nature of fixed devices, the sum of the active power and reactive power of all fixed devices is used to characterize their operating states; in addition, as a stimulus signal, electricity price information is also included in the system state; then, S t is formulated as: Among them, p t is the electricity price at time t; is the sum of the active power of fixed equipment, is the sum of the reactive power of fixed equipment; This is the operating status of the first time-adjustable device. For Nth d The operating status of a time-adjustable device; Action space: Time-adjustable devices include the controllable parameters of the adjustable devices, such as the start and stop status of time-adjustable devices and the active and reactive power of power-adjustable devices; in, It is the start and stop status of the first time-adjustable device; For Nth d The start and stop status of a time-adjustable device; Next-moment state space: The elements of the next-moment state space are the same as the elements of the current state space, but they are obtained after the current action is executed and are expressed as: Among them, the photovoltaic output at the next moment Active power of fixed equipment Reactive power of fixed equipment and electricity price p t+1 Obtained by prediction; from the 1st time adjustable device to the Nth d The operating status of the time-adjustable device Obtained by formulas (1) and (2); the operating state of the power adjustable device at the next moment Calculated by formulas (3) and (4); the operating state of the household battery at the next moment and the operating status of electric vehicles Calculated by formulas (6) and (10) respectively; Immediate rewards: The goal of home energy system optimization is to minimize electricity costs, carbon emissions, and improve user comfort; based on this, the immediate reward is defined as: Where, is the electricity price cost, is the cost of carbon emissions, is the comfort cost; α price is the coefficient corresponding to the electricity price cost, α carbon is the coefficient corresponding to the carbon emission cost, α com is the coefficient corresponding to the comfort cost; specifically, the electricity price cost, carbon emission cost and comfort cost are expressed as: Among them, C t is the carbon price; The preset optimal indoor temperature; In particular, to avoid the range anxiety of electric vehicles and maintain the power factor, a sufficiently large penalty is attached to R t : in, It is the energy state of the electric car when it leaves home; It is the preset minimum energy state to avoid range anxiety; PF is the power factor; set is the preset minimum power factor.

6. The low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning according to claim 1 is characterized in that: In step S5, the SAC optimization process based on hybrid evolutionary reinforcement learning is as follows: S51: The SAC algorithm is implemented using five neural networks: an action network, two value networks, and two goal evaluation networks. The training process consists of five steps. Step 1: Initialize parameters; Initialize the parameters θ of the action network actor And the parameters θ1 of the first value network and the parameters θ2 of the second value network, as well as the hyperparameters λ, γ, β and τ; where λ is the learning rate, γ is the discount factor, τ is the soft update coefficient; β is the regularization coefficient; the initialization parameters of the two target evaluation networks are copied θ1 and θ2. Step 2: Interact with the environment; the action network interacts with the environment and generates the mean of the action and standard deviation For continuous action space, the current action is obtained by Gaussian sampling: tanh(·) is the hyperbolic tangent function; S is the state space set, A is the action space set; ε is Gaussian noise; For actor network function; represents the standard normal distribution; for discrete actions, a binary sampling strategy is used to select the current action: Step 3: Update the value network; randomly select a batch of data B = {(S t ,A t ,S t+1 ,R t )} as training data; the action network generates the next action, and the target value network is responsible for calculating the Q value y corresponding to the current action: in, is the first target value network function, is the second target value network function, θ t1 are the parameters of the first target value network and θ t2 is the parameter of the second target value network; the optimization goal of the value network is to minimize the distance between the predicted Q value of the two value networks and the Q value of the target value network; the loss function of the value network is defined as: is the loss function of the first value network, is the loss function of the second value network; is the first value network function, is the second value network function; gradient descent is used to update the value network parameters: Step 4: Update the action network; the goal of the action network is to maximize the value function and policy randomness, and the loss function of the action network is formulated as: Where, is the expected function; similar to the value network, the parameters of the action network are also updated by the gradient descent algorithm: Step 5: Update the target value network; the target value network is a reflection of the value network and is updated at a slower rate to ensure training stability; Hybrid evolutionary SAC uses evolutionary strategies to search for the global optimal value of the SAC model, and then uses the gradient descent algorithm to finely optimize the parameters. The process can be divided into two steps; Step 1: Evolutionary Strategy Optimization: To avoid local optimality in gradient optimization, Gaussian noise is introduced into the current action network parameters, thereby generating multiple candidate parameter sets: Where k∈[1,…,M] is the index of the candidate parameter, and M is the population size of the generated action parameters; is the exploration scale of the kth candidate parameter, is the sampling noise of the kth candidate parameter; Step 2: The cumulative reward of each candidate parameter is determined by the fitness F of the interaction with the environment k To calculate; Where r represents the state and action trajectory contained in a cycle; the evolutionary strategy optimization update relies on fitness ranking rather than gradient update, and its parameter update process is calculated as follows: Where, is the average fitness of the entire population.

7. The low-carbon operation strategy for household energy based on blockchain and hybrid evolutionary reinforcement learning according to claim 6 is characterized in that: In step S6, the triggering conditions of the smart contract are as follows: The triggering of the home energy system decision model includes two situations: cyclic triggering and abnormal event triggering. Cyclic triggering is a periodic operation, that is, the optimization scheduling is performed every 15 minutes to generate the latest scheduling plan; The default household bill limit is The default carbon emission cap is The preset user comfort limit is There are three types of abnormal events, including excessively high household bills, excessively high carbon emissions, and violations of user comfort. When the trigger conditions are met, the current ledger data is accessed, the home energy management program is activated, and the latest scheduling plan is generated.

Citation Information

Cited By

  • Micro-grid-heating ventilation air conditioner coordinated optimization method based on deep reinforcement learning

    CN120975528A