A Community Digital Grid Autonomous Management Method Based on Multi-Agent System
By adopting the hierarchical architecture and related optimization technologies of multi-agent systems in the community digital grid management system, the limitations of existing systems in flexibility, adaptability and global optimization are solved, and efficient and flexible community management and resource scheduling are achieved.
Patent Information
- Application Number
- CN202411331570.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-09-24
AI Technical Summary
The existing community digital grid management system has the limitations of a centralized management architecture, lacks adaptive learning capabilities, and it is difficult to achieve collaborative optimization between multiple devices and multiple systems. The data processing capabilities are insufficient, so it is impossible to effectively use large-scale data for intelligent decision-making and risk prediction.
The hierarchical architecture design of multi-agent system is adopted, combined with conjugate gradient optimization technology, adaptive meta-strategy and multi-agent entropy regularization stochastic exploration algorithm, to realize the adaptive optimization management of community digital grids. Local agents, regional coordinated agents and global coordinated agents manage community environment, energy consumption and security status through real-time communication, data sharing and collaborative optimization.
It realizes the system's high flexibility and scalability, can make real-time responses in a dynamic and complex community environment, ensures overall optimal resource scheduling, improves overall management efficiency and emergency response capabilities, and enhances data-driven intelligent decision-making capabilities.
Smart Images

Figure CN119168229B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control and automation, and particularly to a community digital grid autonomous management method based on a multi-agent system. Background Art
[0002] With the rapid development of Internet of Things technology and artificial intelligence technology, the concept of intelligent community has gradually entered the public view. Community management has gradually transitioned from the traditional manual management method to the digital and automated management stage. In an intelligent community, collaborative management in multiple aspects such as environmental monitoring, security systems, and energy management is involved, forming the concept of "community digital grid". The community digital grid comprehensively monitors and manages various resources and devices in the community through a large number of sensor devices, intelligent devices, and management systems, aiming to improve the overall management efficiency of the community, reduce energy consumption, and enhance the quality of life of residents. However, although the existing technologies have achieved a certain degree of automation, there are still significant limitations in terms of flexibility, adaptability, and global optimization.
[0003] The existing community management systems mainly rely on a centralized control architecture, and all monitoring and decision-making functions are usually concentrated at a central node. Although this centralized management method can achieve a certain degree of automation, its management ability for large-scale communities is limited. Especially when facing a complex and dynamically changing community environment, the centralized system is prone to problems such as delays, overloading, and untimely responses. In particular, when the community scale is large, the number of sensors, monitoring devices, and types of resources to be managed involved are numerous, and the centralized architecture is prone to bottlenecks due to insufficient data processing capabilities. In addition, once the centralized architecture encounters a failure or network interruption, the management ability of the entire system will be severely limited, and it is impossible to achieve real-time monitoring and management of the community environment.
[0004] In the management of the community digital grid, various data sources such as environmental data, energy consumption, and security information need to be processed in a timely manner, and these data have significant dynamics and complexity. The existing community management systems in the art usually manage through preset rules, but this rule-based management method lacks sufficient adaptability and cannot make dynamic adjustments to the real-time changing community environment. For example, the energy demand in the community may fluctuate significantly at different times, and the existing system cannot effectively adjust the energy scheduling strategy in a short time, often resulting in energy waste or imbalance between supply and demand. Similarly, in terms of environmental monitoring and security management, the fixed-rule management method is difficult to cope with sudden abnormal situations such as extreme weather, sudden fires, or safety accidents, and the system can often only react afterwards, lacking the real-time dynamic response ability to handle emergencies.
[0005] In addition, the intelligent community management systems in the prior art usually adopt a single-level control method. Even when intelligent management means are used, most of them are based on the learning and control of a single intelligent agent, lacking the ability of collaborative management and global optimization. In a complex community environment, there are a large number of interdependent relationships among various devices and resources. Single-agent management cannot effectively coordinate the cooperation among multiple devices, thus achieving global optimality. For example, environmental control devices and energy management systems usually need to cooperate with each other to minimize energy consumption while adjusting the environment. However, a single intelligent agent control strategy is difficult to take into account these conflicting goals simultaneously, resulting in a locally optimal rather than globally optimal result.
[0006] To address the above problems, multi-agent systems have gradually received attention in recent years. Through the collaborative work of multiple agents, multi-agent systems can effectively solve problems in complex and dynamic environments, especially suitable for large-scale distributed management systems. Each agent can independently handle part of the tasks and, through information sharing and strategy coordination with other agents, achieve the optimization of the overall goal. However, in the existing technology, the application of multi-agent systems in community digital grid management is still relatively limited. In particular, the problems of global coordination and adaptive optimization in a large-scale community environment have not been fully solved. Specifically, there are still deficiencies in the collaborative optimization, real-time communication, and the ability to cope with dynamic changes among agents in multi-agent systems. Traditional multi-agent systems usually rely on fixed collaboration rules and cannot dynamically adjust the collaborative mode among agents according to the real-time changes of the environment. In addition, due to the complexity of the community environment itself and the unpredictability of emergencies, it is difficult for the existing technology to achieve simultaneous optimization at the regional and global levels, resulting in inflexible resource scheduling and the need to improve the response speed and efficiency in dealing with emergencies.
[0007] Another obvious defect in the prior art is the data processing ability. The sensors and monitoring devices involved in the community digital grid generate a large amount of real-time data every day, which need to be processed and responded to within a short time. Although the Internet of Things technology provides means for data collection, the existing systems usually limit the data processing to simple rule judgment and threshold triggering, and cannot fully utilize the potential value in the data for in-depth analysis and learning. In addition, when dealing with complex events, the existing technology is difficult to effectively use historical data for prediction and risk prevention and control, lacking data-driven intelligent decision-making ability.
[0008] Generally speaking, the existing community digital grid management systems have the following main defects: First, the limitations of the centralized management architecture lead to poor scalability, high load pressure and lack of flexibility of the system; second, the lack of adaptive learning ability makes it impossible to dynamically adjust management strategies to cope with the real-time changing community environment; third, the control mode of a single intelligent agent is difficult to achieve collaborative optimization among multiple devices and systems; fourth, the data processing ability is insufficient, making it difficult to effectively utilize large-scale data for intelligent decision-making and risk prediction.
[0009] Therefore, how to provide a community digital grid autonomous management method based on a multi-agent system is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0010] An object of the present invention is to propose a community digital grid autonomous management method based on a multi-agent system. By adopting a multi-agent hierarchical architecture design, combining conjugate gradient optimization technology, adaptive meta-strategy and multi-agent entropy regularization stochastic exploration algorithm, the present invention realizes the adaptive optimization management of the community digital grid. Through real-time communication, data sharing and collaborative optimization among local agents, regional coordination agents and global coordination agents, it can effectively manage the community environment, energy consumption and security status. The system has high flexibility and scalability, can make real-time responses in a dynamic and complex community environment, ensure the global optimality of resource scheduling, and improve the overall management efficiency and emergency response ability.
[0011] A community digital grid autonomous management method based on a multi-agent system according to an embodiment of the present invention includes the following steps:
[0012] S1. Design an agent hierarchical architecture, including local agents, regional coordination agents and global coordination agents;
[0013] S2. Establish a distributed communication network. Low-latency communication protocols are used between local agents, Wi-Fi communication is adopted between regional coordination agents and local agents, and global coordination agents communicate with regional coordination agents through a high-speed network;
[0014] S3. Deploy local agents at key points in the community, including environmental sensors, security devices and energy controllers, deploy regional coordination agent servers in community blocks, deploy global coordination agents in the community control center, install agent autonomous learning and control software, and configure regional coordination and global coordination software;
[0015] S4. Local agents operate independently, collect data in real time, and perform autonomous learning and strategy update through conjugate gradient optimization technology and adaptive meta-strategy, and execute operations according to the autonomous learning strategy;
[0016] S5. The regional coordination agent collects local agent data, constructs a regional environmental state model, collaborates and adjusts strategies with local agents through the multi-agent entropy-regularized stochastic exploration algorithm, and executes regional resource scheduling optimization;
[0017] S6. The global coordination agent obtains regional state information, executes global learning and optimization across the entire community, adjusts resource scheduling and collaboration, and issues optimization strategies to the regional coordination agent;
[0018] S7. Local agents upload status information to the regional coordination agent for local status monitoring; the global coordination agent monitors the overall community status; adjust strategies regularly through data analysis, and execute regional and global continuous optimization, including real-time monitoring, preventive operations, and system adaptive adjustment.
[0019] Optionally, the local agents include sensing agents, executive agents, and decision-making agents, which are configured with environmental monitoring, security monitoring, and energy management functions; the regional coordination agent is configured in the community block and is responsible for the collaborative work of the local agents; the global coordination agent is configured at the community level and is responsible for global optimization.
[0020] Optionally, the specific content of S4 includes:
[0021] S41. Local agents use the deployed sensors to collect environmental data, security data, and energy usage data in real time;
[0022] S42. Preprocess the collected raw data, including data filtering, denoising, and standardization processing, to obtain a standardized data set D t ={d 1,t ,d 2,t ,,d n,t}, where d i,t represents the i-th data sample collected by the local agent at time t;
[0023] S43. Construct a learning model for local agents through conjugate gradient optimization technology, define the objective function J(θ), where θ represents the parameters of the local agent policy, calculate the gradient of the objective function, and obtain the gradient Use the conjugate gradient method to iteratively optimize the policy parameters of local agents. The iteration formula is:
[0024]
[0025] where, θ k+1 represents the policy parameters of the (k + 1)-th iteration, θ k represents the policy parameters of the k-th iteration, α k represents the step size parameter of the k-th iteration, p kDenote the conjugate direction of the k-th iteration, γ k Denote the dynamic adjustment coefficient of the k-th iteration, Denote the gradient of the k-th iteration strategy parameter, p k-1 Denote the conjugate direction of the (k - 1)-th iteration;
[0026] S44. During the conjugate gradient optimization process, in combination with the adaptive meta-policy module, dynamically adjust the local agent policy. The adjustment formula of the adaptive meta-policy is:
[0027]
[0028] where, π ′ (θ) represents the policy after the update of the adaptive meta-policy, λ represents the adjustment coefficient, T represents the number of time steps, δ t represents the error term at time step t, v t represents the moving average of the square of the gradient, ∈ represents the smoothing term, represents the gradient of the policy parameter at time step t;
[0029] S45. Define the reward function R(s, a):
[0030]
[0031] where, s represents the current state, a represents the current action, β1 and β2 represent the policy sensitivity parameters, Q env (t) represents the environmental quality parameter at time t, Q env,ideal represents the ideal environmental quality value, Q env,max represents the maximum environmental quality value, Q env,min represents the minimum environmental quality value, E consumed (t) represents the energy consumption at time t, E threshold represents the preset energy consumption threshold, E max represents the maximum energy capacity, S sec (t) represents the safety factor at time t, η represents the safety factor decay rate, ω1, ω2 and ω3 represent the dynamic weight coefficients, used to balance the impacts of environmental quality, energy consumption and safety, and are defined as:
[0032]
[0033] where, i = 1, 2, 3, κ represents the adjustment rate parameter, t0 represents the time point when the agent focuses on each target;
[0034] S46. The local agent uses the exploration and exploitation mechanism in the reinforcement learning framework to execute the updated policy π ′(θ) to maximize the reward function R(s,a), obtain the reward feedback and adjust the policy parameters according to the reward feedback;
[0035] S47. The local agent will apply the self - learned policy π ′ (θ) to the actual operation, determine the control actions for the environment, security, and energy systems according to the policy, and store the experience data in the learning process in the experience pool, and conduct policy evaluation and retraining regularly.
[0036] Optionally, the S5 specifically includes:
[0037] S51. The regional coordination agent regularly collects the environmental state data from the local agents, constructs a regional environmental state model, aggregates and standardizes the collected environmental state data to form a regional state vector S region ={s1, s2, …, s n}, where s i represents the state information of the i - th local agent, including environmental quality, energy consumption, and safety status parameters;
[0038] S52. According to the regional state vector S region construct a regional environmental state model M region , the regional environmental state model M region is expressed as the state transition probability P(S ′ region |S region , A), where S ′ region represents the regional state at the next moment, A represents the set of joint actions of the local agents in the region, | represents the conditional symbol, and use the multi - agent reinforcement learning algorithm to learn the model, calculate the joint policy Π={π1, π2, …, π n} of each local agent in the region, where π i represents the policy of the i - th local agent;
[0039] S53. Introduce a multi - agent entropy - regularization random exploration algorithm, add an entropy - regularization term in the policy optimization process, and define the entropy - regularization term H(Π) of the joint policy as:
[0040]
[0041] where n represents the number of local agents in the region, a i represents the action of the i - th local agent, A i represents the action space of the i - th local agent, s i represents the state of the i - th local agent;
[0042] S54. Maximize the entropy regularization term H(Π) of the joint policy to achieve the entropy regularization of the joint policy;
[0043] S55. During the policy optimization process, the regional coordination agent randomly explores and evaluates the policies of local agents, and executes the random exploration policy Π ′ to obtain the joint reward feedback R region :
[0044]
[0045] where τ represents the entropy regularization coefficient, and R(s i , a i ) represents the reward function of the i-th local agent;
[0046] S56. The regional coordination agent optimizes the joint policy Π according to the joint reward feedback R region to achieve the optimal coordination of resource scheduling within the region. The regional coordination agent issues adjustment commands to local agents according to the optimized joint policy Π * to execute the optimization of regional resource scheduling, including environmental regulation, energy management, and security response.
[0047] Optionally, the specific content of S6 includes:
[0048] S61. The global coordination agent regularly obtains the regional state information from the regional coordination agent, constructs a global state model for the entire community, and aggregates all regional state vectors to form a global state vector where m represents the number of regions in the community, represents the state of the i-th local agent in region k;
[0049] S62. Based on the global state vector S global , the global coordination agent constructs a global environmental state model M global , where the model describes the dynamic changes of the states within the entire community, and defines the global state transition probability P(S ′ global |S global , A global ), where S ′ global represents the global state at the next moment, A global represents the set of joint actions for the entire community, and | represents the conditional symbol. The global coordination agent uses a hierarchical reinforcement learning algorithm to optimize the global policy and solve the optimal global policy where represents the optimal policy for region k;
[0050] S63. The global coordination agent introduces a multi-level entropy regularization method. During the global policy optimization process, an entropy regularization term H(Π global ) is added. The entropy regularization term is defined as:
[0051]
[0052] where m represents the number of regional coordination agents in the whole community, represents the j-th action in region k, represents the action space of region k;
[0053] S64. During the global policy optimization process, the global coordination agent executes a joint action across the whole community according to the current global state S global and the policy Π global , and calculates the global reward R global according to the operation situation of the whole community:
[0054]
[0055] where n k represents the number of local agents in region k, represents the reward of the i-th local agent in region k, ξ represents the entropy regularization coefficient, represents the entropy regularization term of the policy of region k;
[0056] S65. The global coordination agent uses the global reward R global to optimize the global policy Π global through a hierarchical reinforcement learning algorithm, and issues adjustment instructions to each regional coordination agent according to the optimized global policy . Each regional coordination agent executes the optimized policy to adjust the resource scheduling and cooperation within the region.
[0057] Optionally, the specific content of S7 includes:
[0058] S71. During the independent operation process, the local agent uses the deployed sensors to monitor the environmental data, energy usage situation, and security status in real time, and generates local state information after standardizing the collected data;
[0059] S72. The local agent uploads the current state and the actions taken to the regional coordination agent according to the autonomously learned policy π i to form a local monitoring data set, and uploads the data set to the regional coordination agent for real-time monitoring;
[0060] S73. The regional coordination agent analyzes the overall operation situation within the region according to the data uploaded by the local agent, and generates the regional state information S region, and in combination with the dynamic changes within the region, the strategies of local agents are evaluated and adjusted in real time to form a closed-loop feedback mechanism;
[0061] S74. The global coordination agent receives the regional state information S transmitted by the regional coordination agent region , and conducts a comprehensive analysis of the global state S of the entire community global . Based on the global monitoring data and predictive analysis, the global coordination agent issues a policy adjustment instruction π to the regional coordination agent global ;
[0062] S75. The regional coordination agent updates the strategies π of local agents within the region according to the policy adjustment instruction issued by the global coordination agent i . The local agents re-execute tasks according to the adjusted strategies and feedback the status and actions to the regional coordination agent again, forming a closed-loop management system from local to regional and then to global.
[0063] The beneficial effects of the present invention are as follows:
[0064] First of all, by designing a hierarchical architecture of agents and adopting a multi-level structure of local agents, regional coordination agents and global coordination agents, the community management system of the present invention can distributively process complex tasks within the community, avoiding the bottleneck problems of traditional centralized management architectures. This hierarchical architecture makes the system have stronger scalability and flexibility. Each layer of agents can independently process local tasks, and at the same time optimize resource scheduling globally, effectively reducing the load pressure of the system and improving the processing efficiency of the system.
[0065] Secondly, the present invention introduces multi-agent reinforcement learning algorithms, such as conjugate gradient optimization technology, adaptive meta-policy, entropy regularization stochastic exploration algorithm, etc., enabling local agents and regional coordination agents to autonomously learn and adjust strategies in a dynamic community environment, thus possessing self-adaptability. This self-adaptive ability enables the system to dynamically adjust management strategies according to the real-time changes in the environment, energy demand, and security situation, avoiding the limitations of fixed rules in existing systems that cannot cope with complex dynamic environments. Through the application of reinforcement learning and adaptive strategies, local agents can optimize their behaviors in different scenarios, thereby achieving reasonable scheduling of resources and efficient utilization of energy, reducing unnecessary energy consumption, and enhancing the overall efficiency of the community.
[0066] In addition, through the collaborative learning and global optimization mechanism in the multi-agent system, the present invention solves the problem that it is difficult to achieve global optimization in the single-agent management mode of the prior art. Through real-time communication and data sharing among local agents, regional coordination agents, and global coordination agents, multi-level collaborative work can be achieved, ensuring that while local goals are achieved, overall benefits are also taken into account. Through the collaborative optimization among multiple agents, the present invention can achieve global optimality in dealing with problems such as resource conflicts and scheduling balance, avoiding the phenomenon that local agents only pursue local optimality and ignore overall benefits. This optimization mechanism that takes into account both the global and local aspects enables the system to operate efficiently in a large-scale community environment, ensuring the reasonable allocation and dynamic scheduling of various resources within the community.
[0067] In view of the complexity of the community environment and the unpredictability of emergencies, the closed-loop management mechanism of the present invention provides great advantages. Local agents continuously adjust their strategies based on the real-time collected data and transmit feedback information to regional coordination agents and global coordination agents, forming a closed-loop management system from local to regional and then to global. This closed-loop feedback mechanism not only improves the response speed of the system but also enhances the flexibility and robustness of the system in dealing with emergencies. Through the global monitoring and predictive analysis of the global coordination agent, the system can predict potential risks in the community environment and make countermeasures in advance, thus greatly improving the community safety and the accuracy of resource management.
[0068] The present invention also enhances the intelligent level of community management through a data-driven intelligent decision-making mechanism. Local agents, regional coordination agents, and global coordination agents can effectively utilize historical data for prediction and optimization decisions through the collection, preprocessing, and analysis of real-time data. Especially through the designed reward function and entropy regularization strategy, the present invention can achieve a balance among environmental quality, energy consumption, and community safety, ensuring the coordinated operation of various agents in a dynamic environment. This real-time analysis and decision-making mechanism based on data effectively make up for the problem of insufficient data processing ability in the prior art, enabling the system to quickly make reasonable decisions in the face of complex environmental changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0070] Figure 1 is a flowchart of a community digital grid autonomous management method based on a multi-agent system proposed by the present invention;
[0071] Figure 2 is a schematic diagram of the system architecture of a community digital grid autonomous management method based on a multi-agent system proposed by the present invention;
[0072] Figure 3 This is a schematic diagram of the closed-loop feedback mechanism of a community digital grid autonomous management method based on a multi-agent system proposed by the present invention. Detailed implementation manners
[0073] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0074] Refer to Figures 1-3 , a community digital grid autonomous management method based on a multi-agent system, includes the following steps:
[0075] S1. Design an agent hierarchical architecture, including local agents, regional coordination agents, and global coordination agents;
[0076] S2. Establish a distributed communication network. Low-latency communication protocols are used between local agents, Wi-Fi communication is adopted between regional coordination agents and local agents, and global coordination agents communicate with regional coordination agents through a high-speed network;
[0077] S3. Deploy local agents at key points in the community, including environmental sensors, security devices, and energy controllers. Deploy regional coordination agent servers in community blocks, deploy global coordination agents in the community control center, install agent autonomous learning and control software, and configure regional coordination and global coordination software;
[0078] S4. Local agents operate independently, collect data in real time, and perform autonomous learning and policy update through conjugate gradient optimization technology and adaptive meta-policies, and execute operations according to the autonomous learning policies;
[0079] S5. Regional coordination agents collect data from local agents, build a regional environmental state model, and perform policy coordination and adjustment with local agents through the multi-agent entropy regularization random exploration algorithm, and execute regional resource scheduling optimization;
[0080] S6. Global coordination agents obtain regional state information, perform global learning and optimization across the entire community, and adjust resource scheduling and cooperation, and issue optimization policies to regional coordination agents;
[0081] S7. Local agents upload status information to regional coordination agents for local status monitoring; global coordination agents monitor the overall community status; regularly adjust policies through data analysis, and execute regional and global continuous optimization, including real-time monitoring, preventive operations, and system adaptive adjustment.
[0082] In this embodiment, the local agent includes a sensing agent, an executive agent, and a decision-making agent, and is configured with functions of environmental monitoring, security monitoring, and energy management; the regional coordination agent is configured in the community block and is responsible for the collaborative work of the local agents; the global coordination agent is configured at the community level and is responsible for global optimization.
[0083] In this embodiment, the specific steps of S4 are as follows:
[0084] S41. The local agent uses the deployed sensors to collect environmental data, security data, and energy usage data in real time;
[0085] S42. Preprocess the collected raw data, including data filtering, denoising, and standardization, to obtain a standardized data set D t ={d 1,t ,d 2,t ,...,d n,t}, where d i,t represents the i-th data sample collected by the local agent at time t;
[0086] S43. Construct a learning model for the local agent through conjugate gradient optimization technology, define the objective function J(θ), where θ represents the parameters of the local agent's policy, calculate the gradient of the objective function, and obtain the gradient Use the conjugate gradient method to iteratively optimize the policy parameters of the local agent. The iterative formula is:
[0087]
[0088] where, θ k+1 represents the policy parameters of the (k + 1)-th iteration, θ k represents the policy parameters of the k-th iteration, α k represents the step size parameter of the k-th iteration, p k represents the conjugate direction of the k-th iteration, γ k represents the dynamic adjustment coefficient of the k-th iteration, represents the gradient of the policy parameters of the k-th iteration, p k-1 represents the conjugate direction of the (k - 1)-th iteration;
[0089] S44. During the conjugate gradient optimization process, combine the adaptive meta-policy module to dynamically adjust the local agent's policy. The adjustment formula of the adaptive meta-policy is:
[0090]
[0091] where, π ′ (θ) represents the policy after the update of the adaptive meta-policy, λ represents the adjustment coefficient, T represents the number of time steps, δ trepresents the error term at time step t, v t represents the moving average of the square of the gradient, ∈ represents the smoothing term, represents the gradient of the policy parameters at time step t;
[0092] S45. Define the reward function R(s, a):
[0093]
[0094] where s represents the current state, a represents the current action, β1 and β2 represent the policy sensitivity parameters, Q env (t) represents the environmental quality parameter at time t, Q env,ideal represents the ideal environmental quality value, Q env,max represents the maximum environmental quality value, Q env,min represents the minimum environmental quality value, E consumed (t) represents the energy consumption at time t, E threshold represents the preset energy consumption threshold, E max represents the maximum energy capacity, S sec (t) represents the safety factor at time t, η represents the safety factor decay rate, ω1, ω2 and ω3 represent the dynamic weight coefficients used to balance the impacts of environmental quality, energy consumption and safety, and are defined as:
[0095]
[0096] where i = 1, 2, 3, κ represents the adjustment rate parameter, and t0 represents the time point when the agent focuses on each goal;
[0097] S46. The local agent uses the exploration and exploitation mechanism in the reinforcement learning framework to execute the updated policy π ′ (θ) to maximize the reward function R(s, a), obtain the reward feedback and adjust the policy parameters according to the reward feedback;
[0098] S47. The local agent applies the self - learned policy π ′ (θ) to actual operations, determines the control actions for the environment, security and energy systems according to the policy, and stores the experience data in the learning process in the experience pool, and conducts policy evaluation and retraining regularly.
[0099] In this embodiment, the S5 specifically includes:
[0100] S51. The regional coordination agent regularly collects the environmental state data from the local agents, constructs a regional environmental state model, aggregates and standardizes the collected environmental state data to form a regional state vector S region ={s1, s2, …, s n}, where s i represents the state information of the i-th local agent, including environmental quality, energy consumption, and safety status parameters;
[0101] S52. According to the regional state vector S region Construct a regional environmental state model M region , the regional environmental state model M region is expressed as the state transition probability P(S ′ region |S region , A), where S ′ region represents the regional state at the next moment, A represents the set of joint actions of local agents within the region, | represents the conditional symbol, and the multi-agent reinforcement learning algorithm is used to learn the model to calculate the joint policy Π = {π1, π2, …, π n} of local agents within the region, where π i represents the policy of the i-th local agent;
[0102] S53. Introduce the multi-agent entropy regularization random exploration algorithm, add an entropy regularization term during the policy optimization process, and define the entropy regularization term H(Π) of the joint policy as:
[0103]
[0104] where n represents the number of local agents within the region, a i represents the action of the i-th local agent, A i represents the action space of the i-th local agent, and s i represents the state of the i-th local agent;
[0105] S54. Maximize the entropy regularization term H(Π) of the joint policy to achieve the entropy regularization of the joint policy;
[0106] S55. During the policy optimization process, the regional coordination agent randomly explores and evaluates the policies of local agents, and executes the random exploration policy Π ′ to obtain the joint reward feedback R region :
[0107]
[0108] where τ represents the entropy regularization coefficient, and R(s i , a i ) represents the reward function of the i-th local agent;
[0109] S56. The regional coordination agent determines according to the joint reward feedback R regionOptimize the joint policy Π to achieve optimal coordination of resource scheduling within the region. The regional coordination agent issues adjustment commands to local agents according to the optimized joint policy Π * , and execute the optimization of regional resource scheduling, including environmental regulation, energy management, and security response.
[0110] In this embodiment, S6 specifically includes:
[0111] S61. The global coordination agent regularly obtains regional status information from the regional coordination agent, constructs a global status model for the entire community, and aggregates all regional status vectors to form a global status vector where m represents the number of regions in the community, represents the status of the i-th local agent in region k;
[0112] S62. Based on the global status vector S global , the global coordination agent constructs a global environmental status model M global , where the model describes the dynamic changes of the status within the entire community, defines the global state transition probability P(S ′ global |S global ,A global ), where S ′ global represents the global status at the next moment, A global represents the set of joint actions within the entire community, | represents the conditional symbol, and the global coordination agent uses a hierarchical reinforcement learning algorithm for global policy optimization to solve the optimal global policy where represents the optimal policy for region k;
[0113] S63. The global coordination agent introduces a multi-level entropy regularization method and adds an entropy regularization term H(Π global ) during the global policy optimization process. The entropy regularization term is defined as:
[0114]
[0115] where m represents the number of regional coordination agents within the entire community, represents the j-th action in region k, represents the action space of region k;
[0116] S64. During the global policy optimization process, the global coordination agent executes joint actions within the entire community according to the current global status S global and the policy Π global , and calculates the global reward R global based on the operation of the entire community:
[0117]
[0118] Among them, n k represents the number of local agents in area k, represents the reward of the i-th local agent in area k, ξ represents the entropy regularization coefficient, represents the entropy regularization term of the policy in area k;
[0119] S65. The global coordination agent uses the global reward R global to optimize the global policy Π global through a hierarchical reinforcement learning algorithm, and according to the optimized global policy send adjustment instructions to each area coordination agent. Each area coordination agent executes the optimized policy to adjust the resource scheduling and cooperation within the area.
[0120] In this embodiment, S7 specifically includes:
[0121] S71. During the independent operation of the local agent, it uses the deployed sensors to monitor the environmental data, energy usage, and security status in real time, and generates local status information after standardizing the collected data;
[0122] S72. The local agent uploads the current state and the actions taken to the area coordination agent according to the self-learned policy π i to form a local monitoring data set, and uploads the data set to the area coordination agent for real-time monitoring;
[0123] S73. The area coordination agent analyzes the overall operation situation within the area according to the data uploaded by the local agent, generates the area status information S region , and combines the dynamic changes within the area to evaluate and adjust the policies of the local agents in real time, forming a closed-loop feedback mechanism;
[0124] S74. The global coordination agent receives the area status information S region transmitted by the area coordination agent, and comprehensively analyzes the global status S global of the entire community. Based on the global monitoring data and predictive analysis, the global coordination agent issues a policy adjustment instruction π global to the area coordination agent;
[0125] S75. The area coordination agent updates the policies π i of the local agents within the area according to the policy adjustment instruction issued by the global coordination agent. The local agents re-execute the tasks according to the adjusted policies and feedback the status and actions to the area coordination agent again, forming a closed-loop management system from local to area to global.
[0126] Example 1:
[0127] To verify the feasibility of the present invention in implementation, the present invention is applied to a large city. A newly built intelligent community covers approximately 5,000 households, and the community covers an area of approximately 50 hectares. The community is equipped with various intelligent facilities, including an environmental monitoring system, a security system, a smart grid system, etc. These facilities need to achieve efficient resource scheduling, environmental quality monitoring, and security management within the community. However, the energy demand, environmental conditions, and security situation within the community vary dynamically with time and location. For example, the electricity demand fluctuates greatly between day and night, the air quality during morning and evening rush hours is affected by external traffic, and the security needs of residents increase significantly during certain emergencies (such as fires, intrusions).
[0128] Under the existing technology, the resource scheduling and management of the community usually adopt a centralized control scheme. This mode cannot quickly respond to the dynamic changes within the community, easily leading to energy waste, a decline in environmental quality, and security vulnerabilities. Therefore, the community introduces an autonomous management scheme based on a multi-agent system, aiming to achieve optimal scheduling and dynamic adjustment of resources within the community through the collaborative work of local agents, regional coordination agents, and global coordination agents.
[0129] In this intelligent community, local agents are deployed at key nodes of each building, such as environmental sensors on the rooftops, smart electricity meters in households, surveillance cameras, etc. Each local agent is responsible for collecting environmental data, energy usage, and security status. Regional coordination agents are respectively deployed in the management centers of each community block, used to receive real-time data from all local agents within the region, and are responsible for optimizing resource scheduling within the region. The global coordination agent is deployed in the community control center, used to coordinate the operating status of the entire community and adjust the strategies of regional agents when necessary.
[0130] In a typical summer daily scenario, the electricity demand of residents surges, especially during the period from 6 pm to 9 pm. Due to the extensive use of high-energy-consuming devices such as air conditioners, the energy load of the community significantly increases. The traditional centralized energy scheduling system cannot adjust in a timely manner according to demand, often resulting in energy waste or insufficient supply. In the multi-agent system of the present invention, local agents monitor the electricity consumption of each household in real time and feed it back to the regional coordination agent. Through conjugate gradient optimization technology and adaptive meta-strategies, local agents can adjust energy distribution according to real-time electricity demand, reducing ineffective energy consumption. The regional coordination agent optimizes energy scheduling across the entire region by collecting data from all local agents and using the multi-agent entropy regularization stochastic exploration algorithm, thus effectively balancing the energy consumption of each block. At the same time, the global coordination agent monitors the energy usage of the entire community for global optimization, ensuring reasonable and balanced electricity scheduling between different blocks.
[0131] During the implementation process, local agents were installed in each of the 5000 households in the community to monitor energy usage, environmental quality, and security status. The experimental data records on a certain day showed that during the peak electricity consumption period (18:00 - 21:00) in the community, the utilization rate of high-energy-consuming devices such as air conditioners detected by local agents reached 90%. Under the traditional centralized management system, the energy supply could not meet the demand, resulting in unstable electricity loads for some resident households and an energy waste rate of 15%. Under the autonomous management of the multi-agent system of the present invention, through the autonomous optimization of local agents and the collaborative scheduling of regional coordination agents, the energy consumption of the community was effectively controlled, the energy waste rate was reduced to 6%, and the power supply stability was ensured.
[0132] In addition, during the same period, the global coordination agent detected an imbalance in energy supply and demand in different blocks of the community by monitoring the operation data of all regional coordination agents. For example, the energy demand in some blocks had approached the upper limit, while there was still surplus energy usage in other blocks. The global coordination agent immediately issued an optimization strategy to mobilize the energy resources of the surplus blocks to support the high-demand blocks, ultimately achieving energy balance across the entire community and avoiding local power shortages.
[0133] Meanwhile, data from the environmental monitoring section indicate that during the morning rush hour (7:00 - 9:00), local agents detected that the Air Quality Index (AQI) within the community reached 150 (moderate pollution), indicating a relatively high demand for environmental regulation at this time. In traditional systems, the activation of air purification equipment is usually preset for specific time periods and lacks the ability to respond to sudden situations. After using the multi-agent system of the present invention, when local agents detect a deterioration in air quality, they quickly feedback the status information to the regional coordination agent. The regional coordination agent analyzes the environmental conditions across the entire region, combines with the entropy regularization algorithm, activates the necessary air purification equipment, and conducts optimization by time periods. Through data comparison, after the air purification equipment is activated, the AQI within the community drops from 150 to 80 (good) within 30 minutes, and the environmental recovery is significant.
[0134] In terms of security, during a simulated fire drill in the community, local agents detected a fire signal through smoke sensors at the first moment of the fire and immediately fed back the information to the regional coordination agent. The regional coordination agent activates the emergency response mechanism and automatically controls the security equipment within the region, including opening emergency evacuation channels, activating the automatic fire extinguishing system, etc. After receiving the fire alarm from the regional coordination agent, the global coordination agent immediately issues a warning to other regions and mobilizes resources from other regions, such as dispatching additional security forces and coordinating the support of the fire systems in other blocks. Data show that during the simulated drill, the response time of the regional coordination agent was 5 seconds, and the coordination time of the global coordination agent was 10 seconds. The entire community completed global warning and resource scheduling within 15 seconds, with a response speed improvement of more than 50% compared to the traditional system. Table 1 shows the specific performance comparison between the traditional system and the multi-agent system in community management.
[0135] Table 1 Performance comparison between the traditional system and the multi-agent system in community management
[0136]
[0137] After introducing the community digital grid autonomous management method with a multi-agent system, the community management system has achieved significant improvements in various key aspects, specifically manifested in the following aspects:
[0138] First of all, the energy consumption waste rate has been significantly reduced. The energy waste rate under the traditional system was 15%, which reflects that in the centralized management mode, the waste phenomenon caused by untimely or unreasonable energy scheduling is relatively serious. By introducing a multi-agent system, local agents can monitor and adjust the energy demand of each household in real time, and the regional coordination agent further reduces the ineffective energy consumption through optimizing energy scheduling, reducing the waste rate to 6%, a reduction of 9 percentage points. This shows a significant improvement in the intelligent regulation ability of the system in energy management.
[0139] Secondly, the effect of shortening the air quality recovery time is particularly significant. In the traditional system, the recovery of air quality usually depends on preset rules and lacks the ability to flexibly respond to sudden air pollution events, resulting in a relatively long air quality recovery time of about 60 minutes. After introducing the multi-agent system, local agents can quickly respond to the deterioration of air quality and immediately activate air purification equipment. At the same time, the regional coordination agent optimizes the operation strategy of the purification equipment, shortening the air quality recovery time to 30 minutes, a reduction of 50%. This result shows that the adaptive ability of the multi-agent system in environmental regulation has been greatly improved.
[0140] In terms of the security emergency response time, the multi-agent system also shows significant improvement. Under the traditional system, the security emergency response time is 30 seconds, while in the multi-agent system, through the real-time monitoring of local agents and the rapid response mechanism of regional coordination agents, the security emergency response time is shortened to 15 seconds, reducing the time by half. This improvement is crucial for enhancing the security guarantee ability of the community. Especially when dealing with emergencies, it can take emergency measures faster to ensure the safety of residents.
[0141] In terms of the overall system response time, the overall response speed of the traditional system is slow, usually taking 137 seconds to coordinate and complete various tasks. The multi-agent system, through hierarchical and regional coordination optimization, has significantly improved the response efficiency, shortening the overall response time to 72 seconds, a reduction of 53%. This improvement means that the multi-agent system can more efficiently allocate resources and execute tasks when facing complex community management tasks, greatly enhancing the operation efficiency of the system.
[0142] The regional energy balance efficiency has also been significantly improved in the new system. The energy balance efficiency of the traditional system is 82%, indicating that when the traditional system allocates energy between different regions, it cannot fully consider real-time load changes and demand differences. After introducing the multi-agent system, local agents monitor the energy demands of each region in real time, and the regional coordination agent makes dynamic allocations according to the load situation, increasing the energy balance efficiency to 93%, an increase of 11%. This result reflects the significant advantages of the multi-agent system in optimizing energy scheduling, being able to more reasonably allocate the community's energy resources and avoid excessive or insufficient energy distribution.
[0143] Finally, the efficiency of handling emergencies has also been significantly improved. In the traditional system, the efficiency of handling emergencies was 74%, indicating that the system was relatively slow to respond and had poor speed and flexibility in allocating resources when faced with emergencies. Through the collaborative work of local and regional coordination agents, the multi-agent system can quickly identify and respond at the early stage of an event, and the handling efficiency has been increased to 92%, an increase of 18%. This improvement is particularly crucial for enhancing the community's emergency management capabilities. Especially when faced with sudden disasters such as fires and earthquakes, the system can more quickly coordinate relevant resources for emergency response to ensure the safety of residents' lives and property.
[0144] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A community digital grid autonomous management method based on a multi-agent system, characterized in that: The steps include: S1. Design a hierarchical architecture of intelligent agents, including local agents, regional coordination agents, and global coordination agents; S2. Establish a distributed communication network; S3, deploy local intelligent agents at key points in the community, including environmental sensors, security equipment and energy controllers; S4, local agents operate independently, collect data in real time, and autonomously learn and update strategies through conjugate gradient optimization technology and adaptive meta-strategy, and perform operations according to autonomous learning strategies; S5, the regional coordination agent collects local agent data, builds a regional environmental state model, coordinates and adjusts strategies with local agents through a multi-agent entropy regularized random exploration algorithm, and performs regional resource scheduling optimization; S6, the global coordination agent obtains regional status information, performs global learning and optimization across the entire community, adjusts resource scheduling and collaboration, and sends optimization strategies to regional coordination agents; S7, local agents upload status information to regional coordination agents to monitor local status; global coordination agents monitor the overall status of the community; regularly adjust strategies through data analysis, and perform regional and global continuous optimization, including real-time monitoring, preventive operations and system adaptive adjustments; The S4 specifically includes: S41, local intelligent agents use deployed sensors to collect environmental data, security data and energy usage data in real time; S42, preprocessing the collected raw data, including data filtering, denoising and standardization, to obtain a standardized data set D t ={d 1,t ,d 2,t ,…,d n,t }, where d i,t represents the i-th data sample collected by the local agent at time t; S43. Construct a learning model for the local agent through conjugate gradient optimization technology, define the objective function J(θ), where θ represents the parameters of the local agent strategy, calculate the gradient of the objective function, and obtain the gradient Iteratively optimize the policy parameters of the local agent using the conjugate gradient method; S44, in the conjugate gradient optimization process, the adaptive meta-strategy module is combined to dynamically adjust the local agent strategy; S45. Define the reward function R(s,a): Among them, s represents the current state, a represents the current action, β1 and β2 represent the policy sensitivity parameters, and Q env (t) represents the environmental quality parameter at time t, Q env,ideal Indicates the ideal environmental quality value, Q env,max Indicates the maximum environmental quality, Q env,min Indicates the minimum environmental quality, E consumed (t) represents the energy consumption at time t, E threshold Indicates the preset energy consumption threshold, E max Indicates the maximum energy capacity, S sec (t) represents the safety factor at time t, η represents the safety factor decay rate, ω1, ω2 and ω3 represent dynamic weight coefficients, which are used to balance the impact of environmental quality, energy consumption and safety; S46. The local agent uses the exploration and exploitation mechanism in the reinforcement learning framework to execute the updated strategy π′(θ) to maximize the reward function R(s,a), obtain reward feedback and adjust the strategy parameters according to the reward feedback; S47. The local intelligent agent applies the autonomously learned strategy π′(θ) to actual operations, determines the control behavior of the environment, security and energy systems based on the strategy, stores the experience data of the learning process in the experience pool, and conducts strategy evaluation and retraining regularly.
2. According to claim 1, a community digital grid autonomous management method based on a multi-agent system is characterized in that: The local intelligent agents include perceptual intelligent agents, executive intelligent agents and decision-making intelligent agents, which are equipped with environmental monitoring, security monitoring and energy management functions; the regional coordination intelligent agent is configured in the community block and is responsible for the collaborative work of the local intelligent agents; The global coordination agent is configured at the community level and is responsible for global optimization; the regional coordination agent server is deployed in the community block, the global coordination agent is deployed in the community control center, the agent autonomous learning and control software is installed, and the regional coordination and global coordination software are configured; a low-latency communication protocol is used between local agents, Wi-Fi communication is used between the regional coordination agent and the local agents, and the global coordination agent communicates with the regional coordination agent through a high-speed network.
3. According to the method of autonomous management of community digital grid based on multi-agent system in claim 1, it is characterized in that: The iteration formula in S43 is: Among them, θ k+1 represents the policy parameter of the k+1th iteration, θ k represents the strategy parameter of the kth iteration, α k represents the step size parameter of the kth iteration, p k represents the conjugate direction of the kth iteration, γ k represents the dynamic adjustment coefficient of the kth iteration, represents the gradient of the policy parameter at the kth iteration, p k-1 represents the conjugate direction of the k-1th iteration; The adjustment formula of the adaptive meta-strategy in S44 is: Among them, π′(θ) represents the updated strategy of the adaptive meta-strategy, λ represents the adjustment coefficient, T represents the number of time steps, and δ t represents the error term at time step t, v t represents the square moving average of the gradient, ∈ represents the smoothing term, represents the gradient of the policy parameters at time step t; In S45, ω1, ω2 and ω3 represent dynamic weight coefficients, which are used to balance the impact of environmental quality, energy consumption and safety, and are defined as: Among them, i = 1, 2, 3, κ represents the adjustment rate parameter, and t0 represents the time point when the agent pays attention to each target.
4. According to claim 1, a community digital grid autonomous management method based on a multi-agent system is characterized in that: The S5 specifically includes: S51, the regional coordination agent regularly collects environmental status data from local agents, builds a regional environmental status model, aggregates and standardizes the collected environmental status data, and forms a regional state vector S region ={s1,s2,…,s n }, where s i Represents the state information of the i-th local agent, including environmental quality, energy consumption and safety status parameters; S52, according to the regional state vector S region Constructing regional environmental status model M region , regional environmental status model M region Expressed as state transition probability P(S′ region |S region ,A), where S′ region represents the regional state at the next moment, A represents the joint action set of local agents in the region, | represents the conditional symbol, and the multi-agent reinforcement learning algorithm is used to learn the model and calculate the joint strategy of each local agent in the region Π={π1,π2,…,π n }, where π i represents the strategy of the i-th local agent; S53. Introduce the multi-agent entropy regularized random exploration algorithm, add the entropy regularization term in the strategy optimization process, and define the entropy regularization term H(Π) of the joint strategy as: Where n represents the number of local agents in the region, a i represents the action of the i-th local agent, A i represents the action space of the ith local agent, s i Represents the state of the i-th local agent; S54, maximizing the entropy regularization term H(Π) of the joint strategy to achieve entropy regularization of the joint strategy; S55. During the strategy optimization process, the regional coordination agent randomly explores and evaluates the strategies of the local agents, executes the random exploration strategy Π′, and obtains the joint reward feedback R region : Among them, τ represents the entropy regularization coefficient, R(s i ,a i ) represents the reward function of the i-th local agent; S56, the regional coordination agent gives feedback R based on the joint reward region The joint strategy Π is optimized to achieve the optimal coordination of resource scheduling within the region. The regional coordination agent uses the optimized joint strategy Π * , issue adjustment commands to local intelligent agents and perform regional resource scheduling optimization, including environmental control, energy management and security response.
5. The community digital grid autonomous management method based on a multi-agent system according to claim 1 is characterized in that: The S6 specifically includes: S61, the global coordination agent regularly obtains regional state information from the regional coordination agent, builds a global state model for the entire community, and converts all regional state vectors Aggregate to form a global state vector Where m represents the number of regions in the community, represents the state of the i-th local agent in region k; S62, based on the global state vector S global , the global coordination agent builds the global environment state model M global , where the model describes the dynamic changes of the state within the entire community and defines the global state transition probability P(S′ global |S global ,A global ), where S′ global Represents the global state at the next moment, A global represents the set of joint actions in the entire community, | represents the conditional symbol, and the global coordination agent uses a hierarchical reinforcement learning algorithm to optimize the global strategy and solve the optimal global strategy in represents the optimal strategy for region k; S63, the global coordination agent introduces a multi-level entropy regularization method. In the process of global strategy optimization, the entropy regularization term H(Π global ), the entropy regularization term is defined as: Among them, m represents the number of regional coordination agents in the whole community, represents the jth action in region k, represents the action space of region k; S64, in the process of global strategy optimization, the global coordination agent according to the current global state S global and strategy global Perform community-wide joint actions and calculate the global reward R based on the operation of the entire community global : Among them, n k represents the number of local agents in region k, represents the reward of the i-th local agent in region k, ξ represents the entropy regularization coefficient, represents the entropy regularization term of the region k strategy; S65. Global coordination agent uses global reward R global The global strategy Π is trained by a hierarchical reinforcement learning algorithm global Optimize and use the optimized global strategy Adjustment instructions are issued to each regional coordination agent, and each regional coordination agent executes the optimized strategy to adjust resource scheduling and collaboration within the region.
6. The community digital grid autonomous management method based on a multi-agent system according to claim 1 is characterized in that: The S7 specifically includes: S71. During independent operation, the local agent monitors environmental data, energy usage, and security status in real time through deployed sensors, and generates local status information after standardizing the collected data; S72, local agent based on autonomous learning strategy π i , upload the current status and actions taken to the regional coordination agent to form a local monitoring data set, and upload the data set to the regional coordination agent for real-time monitoring; S73, the regional coordination agent analyzes the overall operation status in the region based on the data uploaded by the local agent and generates regional status information S region , and combined with the dynamic changes in the area, the strategies of local agents are evaluated and adjusted in real time to form a closed-loop feedback mechanism; S74, the global coordination agent receives the regional status information S transmitted by the regional coordination agent region , and the global state S of the entire community global Perform comprehensive analysis, based on global monitoring data and predictive analysis, the global coordination agent issues strategy adjustment instructions to the regional coordination agent global ; S75, the regional coordination agent updates the strategy π of the local agent in the region according to the strategy adjustment instruction issued by the global coordination agent i , the local agent re-executes the task according to the adjusted strategy, and again feeds back the state and action to the regional coordination agent, forming a closed-loop management system from local to regional to global.
Citation Information
Patent Citations
Multi-agent federated cooperation method based on deep reinforcement learning
CN112465151A