A community-level energy utilization and multi-user privacy protection system
By using a multi-agent reinforcement learning algorithm to coordinate the control of multiple energy management units and energy storage devices in a community-level smart grid, the problems of energy utilization and privacy protection in a multi-user environment are solved, achieving the dual effects of electricity cost optimization and privacy protection.
Patent Information
- Application Number
- CN202510104563.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In community environments with multiple users and multiple devices, existing smart grid systems struggle to optimize energy utilization and protect privacy, especially when faced with complex electricity price fluctuations. They cannot fully leverage the synergistic effect of multiple meters and energy storage devices, and there is a risk of user privacy leaks.
Employing a multi-agent reinforcement learning algorithm, multiple energy management units, smart meters, and centralized energy storage devices are deployed. Global information optimization is achieved through a commentator network and a discriminator network, which collaboratively control the battery charging and discharging strategy, protecting user privacy and optimizing electricity costs.
Significantly reduces electricity costs in multi-user environments while effectively protecting user privacy, achieving stability and efficiency in community-level energy utilization, and adapting to complex electricity price changes.
Smart Images

Figure CN119906015B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an energy management system in the field of smart grids, and more particularly to a collaborative control system for community-level multi-user smart meters, energy management units, and energy storage devices. Background Technology
[0002] With the rapid development of smart grid technology in my country, more and more households and communities are widely adopting smart meters and energy storage devices (such as household batteries) to optimize electricity usage efficiency and reduce electricity bills. Existing smart meters typically provide billing services by monitoring and recording users' electricity consumption data, and to some extent help users adjust their electricity usage behavior according to electricity price fluctuations. However, most existing systems rely solely on real-time data from a single meter, failing to fully leverage the synergistic effect of multiple meters and energy storage devices within a community, thus making them unable to cope with complex electricity price changes.
[0003] For example, Chinese patent document application number 202410795815.7 discloses a smart meter control system that can realize basic single-user power usage plan adjustment functions. However, due to the lack of advanced optimization algorithms, it is difficult to achieve the same energy utilization effect, especially in a community environment with multiple users and multiple devices. In addition, existing systems also have shortcomings in multi-user privacy protection, such as the inability to accommodate users with different electricity consumption habits, which poses a risk of privacy leakage.
[0004] For example, while some existing collaborative control methods have been applied to local control of smart grids, they often fail to fully consider the overall optimization of the community and the special needs of user privacy protection. Existing technologies struggle to ensure the optimization of collaborative charging and discharging strategies among multiple users and devices, and also find it difficult to achieve energy utilization and electricity management while guaranteeing privacy.
[0005] Therefore, a new smart grid energy management system is needed to coordinate the control of smart meters and energy storage devices at the community level, which can not only optimize energy utilization but also effectively protect users' electricity privacy. Summary of the Invention
[0006] Existing smart meter and energy management systems in multi-user, single-energy storage device community environments suffer from the following technical problems: First, when dealing with complex electricity price changes, existing systems cannot fully utilize the synergistic effect of multiple meters and one energy storage device within a community, inevitably leading to unsatisfactory optimization results in electricity cost savings. Second, the lack of effective privacy protection mechanisms easily exposes users' electricity consumption patterns, posing a risk of privacy leaks. Third, existing multi-agent collaborative control algorithms fail to fully consider the privacy needs of individual users, making it difficult to achieve efficient privacy protection while saving electricity costs. This invention aims to effectively solve the above problems by deploying a multi-agent adversarial reinforcement learning algorithm in community and energy management solutions to achieve collaborative control of multiple smart meters and energy storage devices, thereby optimizing electricity costs and providing user privacy protection functions.
[0007] The technical solution adopted by this invention to solve its technical problem is:
[0008] A community-level energy utilization and multi-user privacy protection system includes: multiple energy management units, multiple smart meters, a centralized energy storage device, and processing equipment that assists the energy management units in performing privacy protection.
[0009] This system optimizes the charging and discharging strategies of multiple energy management units within a community in a multi-user collaborative environment by deploying a novel multi-agent reinforcement learning algorithm.
[0010] The energy management unit corresponds to the smart meter and deploys agents in the multi-agent reinforcement learning algorithm.
[0011] The smart meter is used to monitor the user's electricity load, electricity price, and battery charging and discharging status in real time, and interacts with the system.
[0012] The centralized energy storage device is a battery, and the intelligent agent protects user privacy and saves electricity costs by controlling its charging and discharging behavior.
[0013] The processing device that assists the energy management unit in completing privacy protection is deployed with a centralized network of critics to evaluate the strategies of each energy management unit. The advantage of this is that it can improve the coordination between multiple energy management units through global information, thereby effectively improving the stability and efficiency of the entire system when dealing with the complex coordination of multiple energy management units to complete energy utilization and privacy protection.
[0014] Furthermore, each of the energy management units deploys an independent actor network to process user meter data and formulate localized energy storage device control strategies to control meter readings.
[0015] Furthermore, the energy management unit also includes a discriminator network for measuring the mutual information between the smart meter reading and the user's actual load. In terms of algorithm structure, the discriminator network and the actor network correspond one-to-one and appear in pairs.
[0016] Furthermore, the system deploys a centralized commentator network to receive meter data and charging / discharging decisions from all users, and optimizes the overall battery scheduling strategy based on the global state. This network evaluates and provides feedback on the strategies of each energy management unit based on community-level global information, thereby achieving dual optimization at the community level for both community-level energy utilization management and user-level privacy protection.
[0017] Furthermore, the multi-agent reinforcement learning algorithm flow is as follows:
[0018] Step S1, Environment Initialization: Initialize the state of all smart meters, initial battery charge, power load, and electricity price information (s t Actor network V and its parameter φ, critic network π i and its parameter θ i and the discriminator network F i and its parameter ω i Initialize to random values.
[0019] Step S2, Strategy Execution:
[0020] Within each time step t, the actor network calculates based on the current partial observations. Select Action After interacting with the environment for T time slots, a joint reward sequence (r0, r1, ..., r) is obtained. T (The calculation method is shown in formula (1)).
[0021] The environment described is a community where multiple intelligent agents share a battery. The process involves protecting the privacy of individual users by controlling the battery's charging and discharging behavior, receiving feedback on energy costs and the effectiveness of privacy protection during interactions, and continuously optimizing its own strategies.
[0022] The observed values include the actual power and battery energy percentage of the user corresponding to the agent at time t.
[0023] The action referred to is the charging (discharging) power of the battery;
[0024] The centralized network of critics evaluates the value of each action according to formula (2);
[0025]
[0026] Wherein, the global state is s t The actions of all agents except agent i are as follows: (N represents the number of actor networks), discount factor γ = 1, generalized dominance parameter λ A =0.97, Based on the current critic network V φ Calculated TD-error This represents the value function estimated by the current critic network, which is output by the critic network.
[0027] Step S3, Privacy Feedback Calculation: Discriminator network F corresponding to the i-th actor network i The corresponding actor network (π) is calculated in real time according to formula (3). i Readings of smart meters under control Compared to real load Mutual information between them:
[0028]
[0029] This information is fed back into the reinforcement learning process, As the reward function r p The value is used to measure privacy risk, while Where g t Indicates the current electricity price, g t The settings can flexibly simulate electricity price models. t =λr c +(1-λ)r p This helps optimize the parameters of the critic network and its corresponding actor network.
[0030] Step S4, Loss Function Optimization: Integrating Actor Network The action selection, the evaluation results of the critic network, and the feedback of the discriminator network together constitute the loss function of the overall system. The joint loss function of the actor network and the critic network is shown in Equation (4):
[0031]
[0032] Where ρ t Represents the probability ratio of the old and new strategies under the current action; θ = {θ i ; i = 1, 2, ..., N; ∈ is a clip factor that controls the update magnitude of the actor network.
[0033] The loss function of each discriminator network corresponding to the actor network is shown in Equation (5):
[0034]
[0035] in Equivalent to the discriminator network F ij(t) represents a function of the current time step t; It's the sigmoid function. By using gradient descent, the parameters of the actor and critic networks are optimized, allowing the system to maximize cost optimization while minimizing the risk of privacy breaches.
[0036] Beneficial effects
[0037] By deploying multiple energy management units within the system and using modules that provide inherent privacy protection, multi-user energy management units can collaborate highly in complex community scenarios with shared energy storage devices, intelligently optimize battery charging and discharging strategies, and significantly reduce electricity costs at the community level.
[0038] Meanwhile, with the addition of a discriminator network, the system can effectively control the mutual information between meter readings and the user's actual electricity load, thereby ensuring optimized electricity costs while achieving user-level privacy protection without relying on a trusted third party.
[0039] Furthermore, the design of this system allows for flexible adjustment of the weight λ of the reward function in practical applications, ensuring both the adaptability of the technical solution and meeting the actual needs of users. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the multi-agent algorithm framework according to an embodiment of the present invention;
[0041] Figure 2 This is a schematic diagram of the system deployment according to an embodiment of the present invention;
[0042] Figure 3 This is a load comparison curve for the system simulation deployment in an embodiment of the present invention. Detailed Implementation
[0043] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0044] Figure 1 This is a schematic diagram of the improved multi-agent algorithm framework used in the system of the present invention, taking a user count of 2 as an example;
[0045] Figure 2 This is a schematic diagram of the system deployment of the present invention, taking a user count of 3 as an example;
[0046] Figure 3 The load comparison curve for the simulated deployment of the system of this invention is shown, taking a user count of 3 as an example.
[0047] In the diagram, I(·) represents mutual information. This represents the action performed by actor network numbered n in time slot t. Let X(·) represent the local observations of the actor network with ID n in time slot t, and let X(·) represent protected privacy data. Let s represent the advantage function estimated by the central commentator network. t Represents the global state of time slot t, r t This represents the joint instant reward for time slot t.
[0048] A community-level energy utilization and multi-user privacy protection system includes: multiple energy management units, multiple smart meters, a centralized energy storage device, and processing equipment that assists the energy management units in performing privacy protection.
[0049] This system optimizes the charging and discharging strategies of multiple energy management units within a community in a multi-user collaborative environment by deploying a novel multi-agent reinforcement learning algorithm.
[0050] The energy management unit corresponds to the smart meter and deploys agents in the multi-agent reinforcement learning algorithm.
[0051] The smart meter is used to monitor the user's electricity load, electricity price, and battery charging and discharging status in real time, and interacts with the system.
[0052] The centralized energy storage device is a battery, and the intelligent agent protects user privacy and saves electricity costs by controlling its charging and discharging behavior.
[0053] The processing device that assists the energy management unit in completing privacy protection is deployed with a centralized network of critics to evaluate the strategies of each energy management unit. The advantage of this is that it can improve the coordination between multiple energy management units through global information, thereby effectively improving the stability and efficiency of the entire system when dealing with the complex coordination of multiple energy management units to complete energy utilization and privacy protection.
[0054] The multi-agent reinforcement learning algorithm used in this invention is an improvement on the traditional multi-agent reinforcement learning algorithm. The improved algorithm can provide system-level intrinsic privacy protection in scenarios where multiple energy management units coordinate to control energy storage devices.
[0055] Furthermore, each of the energy management units deploys an independent actor network to process user meter data and formulate localized energy storage device control strategies to control meter readings.
[0056] Furthermore, the energy management unit also includes a discriminator network for measuring the mutual information between the smart meter reading and the user's actual load. In terms of algorithm structure, the discriminator network and the actor network correspond one-to-one and appear in pairs.
[0057] Furthermore, in addition to the decentralized actor and discriminator networks, a centralized commentator network is deployed in the system to receive meter data and charging / discharging decisions from all users, and optimize the overall battery scheduling strategy based on the global state. This network evaluates and provides feedback on the strategies of each energy management unit based on community-level global information, thereby achieving dual optimization of community-level energy utilization management and user-level privacy protection at the community level.
[0058] In practice, after receiving the current environmental state, each energy management unit first generates the next action to control the energy storage device through its local actor network. These actions, once aggregated, collectively control the charging and discharging of the energy storage device. Subsequently, the actions taken by these strategies are sent to a centralized commentator network, which comprehensively evaluates the strategies of all energy management units and generates an advantage function. The network of actors is then guided to update to a better strategy (π).
[0059] In terms of privacy protection, the discriminator network adopts a multilayer perceptron structure to reflect the mutual information between smart meter readings and the user's actual load. The mutual information is estimated using a KL (Kullback-Leibler) divergence metric. During the optimization process, the weights of the discriminator network are continuously updated through backpropagation, thereby accurately estimating the mutual information between the meter readings and the user's actual load, and thus feeding back the privacy risks of the actor network's actions.
[0060] To achieve multi-objective optimization of community-level energy utilization and user-level privacy protection, this system designs a weighted sum single-step reward function as shown in formula (1), which will affect the energy utilization effect r. c and privacy risks p Different weights λ and 1-λ are assigned respectively, as described in Example 1. The advantage of this is that λ can be dynamically adjusted according to actual application needs, thereby enabling more flexible deployment of the energy management proposed in this invention.
[0061] r t =λr c +(1-λ)r p (1)
[0062] In this system, the battery storage unit enables energy sharing among multiple users through intelligent scheduling. The reviewer network dynamically adjusts the battery charging and discharging strategy based on global power demand and electricity price fluctuations to maximize energy utilization efficiency and minimize user electricity costs.
[0063] The actor network and critic network are trained based on an improved multi-agent reinforcement learning algorithm. The critic network performs comprehensive analysis on the power data of multiple users to generate a globally optimal battery charging and discharging strategy, ensuring that the system saves costs while protecting user privacy.
[0064] The discriminator network reduces the correlation between meter data and user electricity consumption habits by calculating the mutual information between the user's smart meter reading and their actual electricity load, thus optimizing the balance between user privacy protection and system scheduling.
[0065] The battery storage unit communicates in real time with the smart meters of multiple users. The system can automatically adjust the charging and discharging strategy according to each user's electricity load and electricity price fluctuations to ensure that electricity costs are minimized during peak hours and that battery energy storage is fully utilized during off-peak hours.
[0066]
Example 1
[0067] Each household smart meter is equipped with an independent actor network and a corresponding discriminator network, such as Figure 1 The network shown is responsible for outputting battery charging and discharging decisions based on the current environmental conditions.
[0068] Each actor network consists of an input layer, a hidden layer, and an output layer. The input layer receives information including electricity price, current battery level, and household load. The hidden layer processes the input information using a non-linear activation function, and the output layer provides the final charging and discharging actions. By minimizing mutual information, the needs of privacy protection and energy utilization can be balanced when optimizing the actor network.
[0069] The discriminator network is used to estimate the mutual information between smart meter readings and the actual load. Its paired actor network, with the assistance of the discriminator network, effectively reduces sensitive information contained in the meter readings by decreasing mutual information, thereby protecting user privacy.
[0070] The critic network plays an evaluation role in the algorithm, such as... Figure 1 As shown, it receives global information, including the decisions of all energy management units and the overall environmental state, to calculate and feedback an index of the quality of each smart meter's actions. Its core is a deep neural network based on global information, which takes the status information of all meters as input and outputs the quality of each energy management unit's actions, guiding the network to update accordingly.
[0071] The multi-agent reinforcement learning algorithm used in this embodiment is as follows:
[0072] Step S1, Environment Initialization: Initialize the state of all smart meters, initial battery charge, power load, and electricity price information (s t Actor network V and its parameter φ, critic network πi and its parameter θ i and the discriminator network F i and its parameter ω i Initialize to random values.
[0073] Step S2, Strategy Execution:
[0074] Within each time step t, the actor network calculates based on the current partial observations. Select Action After interacting with the environment for T time slots, a joint reward sequence (r0, r1, ..., r) is obtained. T (The calculation method is shown in formula (1)).
[0075] The environment described is a community where two intelligent agents share a single battery. The process involves protecting individual user privacy by controlling the battery's charging and discharging behavior, receiving feedback on energy costs and privacy protection effectiveness during interaction, and continuously optimizing its own strategy.
[0076] The observed values include the actual power and battery energy percentage of the user corresponding to the agent at time t.
[0077] The action referred to is the charging (discharging) power of the battery;
[0078] The centralized network of critics evaluates the value of each action according to formula (2);
[0079]
[0080] Wherein, the global state is s t The actions of all agents except agent i are as follows: (N represents the number of actor networks, which is 2 in this example), discount factor γ = 1, generalized dominance parameter λ A =0.97, Based on the current critic network V φ Calculated TD-error This represents the value function estimated by the current critic network, which is output by the critic network.
[0081] Step S3, Privacy Feedback Calculation: Discriminator network F corresponding to the i-th actor network i The corresponding actor network (π) is calculated in real time according to formula (3). i Readings of smart meters under control Compared to real load Mutual information between them:
[0082]
[0083] This information is fed back into the reinforcement learning process, As the reward function r p The value is used to measure privacy risk, while Where g t Indicates the current electricity price, g t The settings can flexibly simulate electricity pricing models; one example is tiered pricing. t =λr c +(1-λ)r p This helps optimize the parameters of the critic network and its corresponding actor network.
[0084] Step S4, Loss Function Optimization: Integrating Actor Network The action selection, the evaluation results of the critic network, and the feedback of the discriminator network together constitute the loss function of the overall system. The joint loss function of the actor network and the critic network is shown in Equation (4):
[0085]
[0086] Where, ρ t Represents the probability ratio of the old and new strategies under the current action; θ = {θ i ;i = 1, 2, ..., N;∈ is a clip factor that controls the update magnitude of the actor network.
[0087] The loss function of each discriminator network corresponding to the actor network is shown in Equation (5):
[0088]
[0089] in Equivalent to the discriminator network F i j(t) represents a function of the current time step t; It's the sigmoid function. By using gradient descent, the parameters of the actor and critic networks are optimized, allowing the system to maximize cost optimization while minimizing the risk of privacy breaches.
[0090] Throughout the system, smart meters collaborate through multi-agent reinforcement learning, forming a highly efficient distributed energy management system. The introduction of a discriminator network enhances user privacy protection while maintaining high energy efficiency. Deploying a centralized critic network increases the algorithm's adaptability and robustness, and replacing decentralized critic networks with a centralized one reduces computational complexity.
[0091]
Example 2
[0092] Based on Embodiment 1, this embodiment further provides a deployment method for the system of the present invention in a building setting, applicable to scenarios where multiple users share the same energy storage device (community-level battery). The system achieves intelligent battery scheduling and user privacy protection through a multi-agent reinforcement learning algorithm.
[0093] like Figure 2 As shown, each user in the building (smart meter) is connected to a central shared battery via a local smart meter. Users' daily electricity consumption is monitored in real time by the smart meters and transmitted to the system's energy management unit for decision-making. All users share a community-level battery, and the control strategies of each user's energy management unit are aggregated at the battery to control battery behavior, achieving the dual goals of optimizing electricity costs and protecting privacy.
[0094] Specifically, each user's smart meter monitors key information such as actual household electricity consumption and the current status of the community-level battery in real time, transmitting this information to their respective energy management unit (EMU) for processing. Each EMU makes local charging or discharging suggestions based on the received local information. The management of the community-level battery is jointly handled by all users' EMUs. Simultaneously, the reviewer network receives decision-making and global status information from all smart meters, evaluates each user's charging and discharging requests, and comprehensively considers the electricity demand of all users in the building, electricity price fluctuations, and privacy protection goals, guiding the EMU optimization from a global perspective. The system ensures that the EMU meets users' electricity needs while intelligently charging the battery, taking into account complex electricity price changes, to maximize overall economic benefits. Furthermore, the EMU must ensure that meter readings do not excessively reveal users' true electricity consumption habits. In this way, the system optimizes energy utilization while ensuring that user privacy is not easily compromised. Over time, electricity prices and user needs may change. The EMU within the system can continuously learn and optimize its local battery charging and discharging strategies to ensure that it meets users' dynamic needs and effectively adapts to various electricity fluctuation scenarios.
[0095] Under this system deployment scheme, users can achieve intelligent and optimized power scheduling in building scenarios. By sharing a single battery, users can effectively share energy resources, reducing energy costs for the community and individuals. The system is not only economical but also provides users with a sense of security and privacy protection.
[0096]
Example 3
[0097] Based on Example 1, this example addresses the need for user privacy protection by shaping the user's load to prevent external parties from inferring the user's actual electricity consumption behavior through meter readings. Figure 3 As shown:
[0098] Figure 3 The left side shows the user's actual electricity load curve, which represents the user's load situation when actually using electricity.
[0099] like Figure 3 On the right, to protect user privacy, the system uses a discriminator network to process the actual load, causing the output load curve to exhibit irregular noise characteristics. This shaped load curve interferes with external analysis of the user's electricity usage habits.
[0100] During this process, the discriminator network compares the difference between the user's actual load and the shaped load curve to ensure that privacy is protected by perturbing the actual load. The battery achieves simultaneous privacy protection and power cost optimization through charge and discharge scheduling.
[0101] Through this implementation method, the system not only effectively improves the privacy of users' electricity consumption behavior, but also ensures the stability of system operation and the efficient use of power resources.
[0102] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.
Claims
1. A community-level energy utilization and multi-user privacy protection system, characterized in that, The system comprises: a plurality of energy management units, a plurality of smart meters, a centralized energy storage device, and a processing device assisting the energy management units in privacy protection; The system optimizes the charging and discharging strategies of multiple energy management units in a community by deploying a new multi-agent reinforcement learning algorithm in a multi-user collaborative environment. The energy management unit corresponds to the smart meter, and the agent in the multi-agent reinforcement learning algorithm is deployed. The smart meter is used to monitor the user's power load, electricity price and battery charging and discharging status in real time, and interacts with the system. The centralized energy storage device is a battery, and the agent controls its charging and discharging behavior to protect user privacy and save electricity cost. The processing device assisting the energy management unit to complete the privacy protection function deploys a centralized critic network to evaluate the strategy of each energy management unit.
2. The community-level energy utilization and multi-user privacy protection system of claim 1, wherein, The centralized critic network receives all user meter data and charging and discharging decisions, and optimizes the overall scheduling strategy of the battery based on the global state. Based on the global information of the community, the network evaluates and feeds back the strategy of each energy management unit, thereby achieving double optimization of community-level energy utilization management and user-level privacy protection at the community level.
3. The community-level energy utilization and multi-user privacy protection system of claim 1, wherein, The agent in each energy management unit deploys an independent actor network to process the user's meter data and develop a localized energy storage device control strategy to control the meter reading.
4. The community-level energy utilization and multi-user privacy protection system of claim 3, wherein, The energy management unit also includes a discriminator network for measuring the mutual information between the smart meter reading and the user's real load. From the algorithm structure, the discriminator network and the actor network correspond one-to-one and appear in pairs.
5. The community-level energy utilization and multi-user privacy protection system of claim 4, wherein, The multi-agent reinforcement learning algorithm process is as follows: Step S1, environment initialization: initialize the state of all smart meters, the initial battery capacity, power load and electricity price information ; critic network and its parameters ; actor network and its parameters ; and discriminator network and its parameters are initialized to random values; Step S2, strategy execution: At each time step , the actor network selects an action according to the current partial observation , and interacts with the environment for T time slots to obtain a joint reward sequence , and the calculation method is shown in formula (1); (1) The environment is a community where multiple agents share a battery. By controlling the battery charging and discharging behavior, the privacy of individual users is protected. During the interaction process, feedback on electricity cost and privacy protection effect is received, and the strategy is continuously optimized. The observation value includes the actual power of the agent's corresponding user at time t and the battery storage percentage; The action is the battery charging / discharging power; The centralized critic network evaluates the value of each action according to formula (2); (2) where the global state is , and the action of the critic is , N represents the number of actor networks, the discount factor , the generalized advantage parameter , is the TD-error calculated based on the current critic network , and represents the value function estimated by the critic network, which is output by the critic network; Step S3, privacy feedback computation: the discriminative network corresponding to the i-th actor network is computed According to formula (3), the reading of the smart meter under the control of its corresponding actor network is computed in real time The mutual information between the real load is computed (3) This information is fed back to the reinforcement learning process, which measures the privacy risk as the value of a reward function while where denotes the current electricity price, the setting can flexibly simulate the electricity price model; helps to optimize the critic network and its corresponding actor network parameters; Step S4, Loss Function Optimization: Integrating Actor Network The action selection, the evaluation results of the critic network, and the feedback of the discriminator network together constitute the loss function of the overall system; the joint loss function of the actor network and the critic network is shown in formula (4): (4) where denotes the probability ratio of the new and old policies under the current action; ; is a clipping factor that can control the update magnitude of the actor network; The loss function of each discriminator network corresponding to the actor network is shown in formula (5): (5) wherein is equivalent to the discriminator network ; denotes a function with respect to the current time step t; is a sigmoid function; the parameters of the actor and critic networks are optimized by gradient descent such that the system maximizes the cost optimization objective while minimizing the risk of privacy leakage.
6. The community-level energy utilization and multi-user privacy protection system of claim 4, wherein, The battery storage unit in the system realizes energy sharing among multiple users through intelligent scheduling. The critic network dynamically adjusts the battery charging and discharging strategy based on global power demand and price fluctuations to maximize energy utilization efficiency and minimize user electricity cost.
7. The community-level energy utilization and multi-user privacy protection system of claim 4, wherein, The actor network and critic network are trained based on the improved multi-agent reinforcement learning algorithm. The critic network analyzes the power data of multiple users to generate a globally optimal battery charging and discharging strategy, ensuring that the system saves cost while protecting user privacy.
8. The community-level energy utilization and multi-user privacy protection system of claim 4, wherein, The system introduces a privacy protection discriminator network. The discriminator network calculates the mutual information between the user's smart meter reading and the actual power load to reduce the correlation between the meter data and the user's power habits, optimizing the balance between user privacy protection and system scheduling.
9. The community-level energy utilization and multi-user privacy protection system of claim 4, wherein, The battery storage unit communicates with the smart meters of multiple users in real time. The system can automatically adjust the charging and discharging strategy according to the power load and price fluctuation of each user, to minimize the electricity bill during peak hours and fully utilize the battery storage during off-peak hours.
Citation Information
Patent Citations
Cooperative control system and method for comprehensive energy system of multiple industrial parks
CN117272842A
User side renewable energy utilization method giving consideration to cost effectiveness and privacy security
CN118802329A