Autonomous Decision-Making System and Operation Method of Multi-Level Intelligent Agents in Urban Areas

By constructing a multi-level intelligent agent autonomous decision-making system in the urban integrated energy system, the problems of response delay and decision reliability in the traditional cloud-based centralized processing mode are solved. This enables on-site data analysis and local strategy generation, improving the system's real-time response capability and stability.

CN121477981BActive Publication Date: 2026-03-10STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Traditional cloud-based centralized processing models are insufficient to meet the needs of urban integrated energy systems for real-time response, agile regulation, and efficient bandwidth utilization. In particular, they suffer from insufficient global coordination capabilities, response timeliness, and decision reliability in multi-energy coordinated scheduling and rapid handling of extreme events.

Method used

Construct an autonomous decision-making system for multi-level intelligent agents in urban areas, including edge, cloud, and terminal decision-making units. By decentralizing decision-making to the edge and terminal sides, data analysis and policy generation can be achieved locally. Combined with lightweight models and distributed algorithms, collaborative optimization and regulation of multi-level intelligent agents can be realized.

Benefits of technology

It significantly reduces system latency, improves decision-making efficiency, achieves second-level regional scheduling, and possesses rapid self-healing and elastic control capabilities, ensuring continuous and stable operation of the system in complex network environments and enhancing the robustness of the decision-making process and the accuracy of regional scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121477981B_ABST
    Figure CN121477981B_ABST
Patent Text Reader

Abstract

An autonomous decision-making system and its operation method for multi-level intelligent agents in urban areas are proposed. Each end-side decision-making unit broadcasts equipment state variables to the edge-side decision-making units. The edge-side decision-making units determine the global baseline state and generate control commands for each end-side decision-making unit. After the equipment executes the control commands, each end-side decision-making unit updates the equipment state variables based on real-time equipment operation data. When an abnormal event is detected, the edge-side decision-making units predict the decision variables of each equipment based on a collaborative optimization model of the equipment's objectives and the urban area's objectives. The cloud-based control platform predicts globally consistent variables based on the collaborative optimization model of all edge-side decision-making units. Under system operation constraints, the control commands of each edge-side decision-making unit are predicted based on the global dynamic optimization objective. The edge-side decision-making units decompose the control commands into control commands for each end-side decision-making unit. This system solves the problems of weak global coordination capabilities and untimely collaborative scheduling under extreme events in urban integrated energy systems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of multi-level agent control of urban regional integrated energy systems, and particularly relates to a general-purpose computing unit for autonomous decision-making of lightweight urban regional integrated energy multi-level agents and a running method thereof. BACKGROUND

[0002] The rapid development of Internet of Things technology has promoted the digital transformation of energy systems. The number and types of terminal devices connected to urban regional integrated energy systems are increasing, forming a complex sensing and regulation system with multi-source heterogeneity and multi-layer distribution. In the face of the challenges of massive device access and high-frequency data concurrency, the traditional cloud-centric processing mode has been difficult to meet the actual needs of real-time response, agile regulation and efficient bandwidth utilization. Especially in the scenes of multi-energy collaborative scheduling and rapid disposal of extreme events, the system still has obvious deficiencies in global coordination capability, response timeliness and decision reliability.

[0003] Under this background, edge computing provides a new technical path for building a hierarchical autonomous and multi-level collaborative urban regional integrated energy system. Its distributed architecture is highly consistent with the structure, function and regulation target of the regional integrated energy system, and can sink part of the decision-making function to the edge side and the terminal side, realize data processing nearby, strategy generation and dynamic optimization on site, effectively reduce the cloud load, reduce communication delay, and improve the overall response capability of the system. With the continuous improvement of communication protocols and embedded processor performance, the potential of edge computing in collaborative optimization, distributed reasoning and multi-agent decision-making is constantly released, laying a technical foundation for the realization of "cloud-edge-terminal" collaborative regulation of urban regional integrated energy systems.

[0004] Currently, edge computing units have been gradually applied in industrial Internet of Things, smart grids and other fields, but their application in integrated energy systems still faces many challenges: including how to realize lightweight and standardized hardware design of edge and terminal side agents, how to build a general decision-making unit supporting multi-level collaborative optimization, how to improve the processing efficiency and transmission reliability of massive multi-source data on the edge side, and how to realize efficient collaboration and control across devices, regions and levels through distributed algorithms, which have become key problems restricting the large-scale deployment and intelligent level improvement of the system. SUMMARY

[0005] To solve the problems in the prior art, the application provides an autonomous decision system of a multi-level agent of a city area and a running method thereof, light autonomous decision of multi-level agent collaborative optimization and regulation of a regional integrated energy system, and deployment at the edge side and the terminal side of the multi-level agent collaborative regulation system of the city area integrated energy system; by sinking the decision to the edge side, on-site data analysis, strategy generation and issuance are realized, and the system decision efficiency is improved; meanwhile, multi-level and multi-node agents in the system form a layered controllable resource group, through dynamic collaboration of the multi-level agents, the multi-level agent regulation strategies are complementary and mutually beneficial, the group regulation strategy is optimized, and thus the technical problems of weak global collaboration ability of the city area integrated energy system, and untimely collaborative scheduling under extreme events are solved.

[0006] The application adopts the following technical scheme.

[0007] The application provides an autonomous decision system of a multi-level agent of a city area, comprising: a terminal-side decision unit, an edge-side decision unit and a cloud regulation platform.

[0008] Each terminal-side decision unit broadcasts a device state variable to the edge-side decision unit; the edge-side decision unit determines a global reference state according to all the device state variables, and generates a control instruction of each terminal-side decision unit by using the deviation of each device state variable from the global reference state;

[0009] After the device executes the control instruction, each terminal-side decision unit updates the device state variable according to real-time operation data of the device, and based on the updated device state variable, the edge-side decision unit predicts a decision variable of each device based on a collaborative optimization model of the device target and the city area target;

[0010] The cloud regulation platform fuses the collaborative optimization model of all the edge-side decision units according to the decision variable of each device, predicts a global consistency variable by using the fused collaborative optimization model, establishes a reward function based on the global consistency variable, and predicts a regulation instruction of each edge-side decision unit according to a global dynamic optimization target established by using the reward function under system operation constraints;

[0011] The edge-side decision unit decomposes the regulation instruction into a control instruction of each terminal-side decision unit.

[0012] Preferably, one city area in an integrated energy system corresponds to one edge-side decision unit, and each device in the city area corresponds to each terminal-side decision unit.

[0013] Preferably, the terminal-side decision unit comprises: a multi-source sensor array, a microprocessor and a communication module.

[0014] The multi-source sensor array is connected with the master control unit of each device through a standard digital interface to collect real-time operation data of the device.

[0015] The microprocessor integrates a neural network processing unit, the neural network processing unit is built-in a lightweight model, device state variables are extracted from real-time collected device operation data based on the lightweight model, and it is judged whether an abnormal event exists based on the device state variables;

[0016] The end-side decision unit uploads the device state variables and the abnormal event determined based on the device state variables to the edge-side decision unit through the communication module.

[0017] Preferably, the core of the hardware architecture of the edge-side decision unit comprises: a processing module, a storage module and a communication module.

[0018] The storage module stores deviation detection rules and lightweight decision rules.

[0019] The processing module performs weighted aggregation on the device state variables to obtain a global reference state, reads the deviation detection rules in the storage module, detects the deviation of each device state variable from the global reference state, reads the lightweight decision rules in the storage module when it is detected that there is no abnormality in the deviation of each device state variable from the global reference state, generates control instructions for each end-side decision unit based on the lightweight decision rules, and generates a request for the cloud regulation platform to update the lightweight model when it is detected that there is an abnormality in the deviation of each device state variable from the global reference state.

[0020] The edge-side decision unit broadcasts the control instructions to the end-side decision unit through the communication module, and sends a request for updating the strategy network parameters to the cloud regulation platform through the communication module.

[0021] Preferably, the cloud regulation platform comprises: a heterogeneous computing cluster and a communication module; the heterogeneous computing cluster learns to obtain strategy network parameters in the state space and the action space based on a cooperative stability judgment mechanism with the maximum reward function as the target; wherein, based on the global consistency variable, the city area energy efficiency improvement amount, the network communication overhead and the control instruction fluctuation amount are used to establish a reward function; the learned strategy network parameters are used to prune and reconstruct the lightweight model; and the cloud regulation platform feeds back the pruned and reconstructed lightweight model to each end-side decision unit through the communication module.

[0022] The application also provides a running method of the autonomous decision system of the city area multi-level intelligent agent.

[0023] A unified lightweight model is integrated in each end-side decision unit; based on the lightweight model, the end-side decision unit broadcasts the device state variables extracted from the real-time collected device operation data to the edge-side decision unit.

[0024] The edge decision unit performs weighted aggregation of device state variables to obtain a global baseline state; based on deviation detection rules, it detects the deviation between each device state variable and the global baseline state; when no abnormality is detected in the deviation between each device state variable and the global baseline state, it generates control commands for each edge decision unit based on lightweight decision rules; when an abnormality is detected in the deviation between each device state variable and the global baseline state, it generates a request for the cloud control platform to update the lightweight model.

[0025] Each end-side decision unit uses a lightweight model to extract the state variables of each device from the operating data after each device executes control commands; each end-side decision unit iteratively updates the state variables of each device based on the comprehensive weight factor issued by the side-side decision unit; when the end-side decision unit sends an abnormal event detected based on the updated device state variables to the side-side decision unit, the side-side decision unit uses the subgradient method to perform distributed solution of the collaborative optimization model under joint constraints to obtain the decision variables of each device;

[0026] The cloud-based control platform uses the alternating direction multiplier method to fuse the collaborative optimization models of the side decision units in different regions based on the decision variables of each device, and then performs distributed solution on the fused collaborative optimization model to obtain globally consistent variables.

[0027] The cloud-based control platform establishes a global dynamic optimization objective using a reward function based on globally consistent variables. It then employs a rolling optimization model predictive control framework to iteratively solve the global dynamic optimization objective under system operation constraints, obtaining predicted control commands. These predicted control commands are then distributed to edge decision units via an encrypted channel. The edge decision units use a distributed model predictive control framework to decompose the control commands into control commands for each edge decision unit.

[0028] Preferably, the side decision-making unit loads the topology weight matrix from the cloud-based control platform. The resulting global baseline state is as follows:

[0029]

[0030] In the formula, For period The global baseline state, For the first The period of broadcast from each end-side decision unit to the side-side decision unit State variables, Topological weight matrix The first in Each weight, This refers to the number of decision-making units on the edge.

[0031] Preferably, the deviation detection rules include:

[0032] 1) If all deviations show a consistent deviation, then the system is determined to have been subjected to a common, global disturbance;

[0033] 2) If any deviation is different from the other deviations, the equipment corresponding to any deviation is determined to be faulty.

[0034] Preferably, the control command for each end-side decision unit is as follows:

[0035]

[0036] In the formula, For the first Each end-side decision unit in the cycle +1 control command; These are built-in functions for a lightweight decision rule engine. This is the adjustment coefficient.

[0037] Preferably, the edge decision-making unit generates a request for the cloud control platform to update the lightweight model; the cloud control platform establishes a state space and an action space, and establishes a reward function using the energy efficiency improvement of urban areas, network communication overhead, and control command fluctuations; with the goal of maximizing the reward function, the platform learns policy network parameters in the state space and action space based on a collaborative stability judgment mechanism; the cloud control platform uses the learned policy network parameters to prune and reconstruct the lightweight model, and feeds back the pruned and reconstructed lightweight model to each edge decision-making unit;

[0038] Among them, time reward function as follows:

[0039]

[0040] In the formula, , , All are weighting coefficients. , They are time points The equipment status variables and control commands, Based on and The amount of energy efficiency improvement in urban areas Based on and network communication overhead, Based on and The fluctuation amount of the control command.

[0041] Preferably, the edge decision unit calculates a comprehensive weight factor according to the electrical coupling factor and the communication quality factor and distributes it to each end-side decision unit, and the calculation formula is as follows:

[0042]

[0043] In the formula, is a comprehensive weight factor between the device and , is an equivalent electrical impedance between the device and , is the maximum adjustable power of the device , is the historical communication success rate between the device and , is a set of neighboring devices of the device in the communication topology.

[0044] Each end-side decision unit extracts the state variables of each device from the running data after each device executes the control instruction using a lightweight model ; each end-side decision unit iteratively updates the state variables of each device based on the comprehensive weight factor, as shown in the following formula:

[0045]

[0046] In the formula, is the state variable of the device in the cycle +1.

[0047] Preferably, the collaborative optimization model based on the device target and the urban area target is as follows:

[0048]

[0049] In the formula, is the target of the device ; is a collaborative cost function between the device and its neighboring devices , representing the urban area target; is a collaborative weight coefficient, is the number of end-side decision units, is a set of neighboring devices of the device .

[0050] The obtained decision variables are as follows:

[0051]

[0052] wherein, is the state variable of the device in the period is the decision variable of the device is the projection operation to the feasible set ; is the state variable of the device in the period , is the subgradient of the Lagrangian function, is the step sequence of the period .

[0053] Preferably, the fused collaborative optimization model is as follows:

[0054]

[0055] The global consistency variable obtained by solving is:

[0056]

[0057] wherein, is the global consistency variable of the period , is the dual variable of the device in the period , is the decision variable of the device in the period , is the local optimization objective of the nth end-side decision unit, is the collaborative weight coefficient; When

[0058] , , is the set error threshold, and the system converges.

[0059] Preferably, the global dynamic optimization objective is as follows:

[0060]

[0061] wherein, is the strategy network parameter learned by the cloud regulation platform, is the discount factor, , is the optimization period length, is the expected function.

[0062] The application also relates to a terminal, which comprises a processor and a storage medium; the storage medium is used for storing instructions; the processor is used for operating according to the instructions to execute the steps of the method.

[0063] The present invention is also a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method.

[0064] The beneficial effects of this invention, compared with the prior art, include at least the following: This invention constructs a cloud-edge-device multi-level intelligent agent collaborative technology architecture for urban regional integrated energy systems. This architecture physically consists of three hardware entities: a cloud-based control platform, edge-side decision-making units, and device-side decision-making units. By decentralizing lightweight intelligent decision-making capabilities to the edge and device sides, it achieves second-level regional scheduling, significantly reducing overall system latency and cloud load. It exhibits rapid self-healing and elastic control capabilities in extreme event scenarios with extremely high real-time requirements. The device side can maintain basic operation based on a rule base even in the event of a network outage, and the edge side can achieve regional autonomy, thus maintaining continuous and stable system operation even in complex network environments. The strategy update mechanism with stability judgment introduced in the cloud further enhances the robustness of the decision-making process. By designing a lightweight device-side model process of "structured pruning—low-rank approximation—nonlinear quantization," the model and hardware are highly matched. The topology weight matrix used in edge-side collaboration is dynamically generated by the cloud based on device coupling relationships, improving the accuracy of regional scheduling. This design significantly reduces end-side resource consumption and extends device battery life while ensuring decision-making accuracy. Furthermore, by using simulation-optimized weighting coefficients in the reward function, the system balances multiple conflicting objectives such as energy efficiency, communication overhead, and control stability. Attached Figure Description

[0065] Figure 1 This is a schematic diagram of the autonomous decision-making system of a multi-level intelligent agent in an urban area proposed in this invention.

[0066] Figure 2 This is a flowchart of the operation method of the autonomous decision-making system of multi-level intelligent agents in urban areas proposed in this invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.

[0068] like Figure 1As shown, this invention constructs a cloud-edge-device multi-level intelligent agent collaborative technology architecture for urban regional integrated energy systems. Physically, this architecture consists of three hardware entities: a cloud-based control platform, edge-side decision-making units, and device-side decision-making units. Functionally, these correspond to the cloud-side, edge-side, and device-side intelligent agents. The core of this architecture lies in the efficient collaboration of these three hardware levels: the device-side decision-making unit, relying on its multi-source sensors and dedicated AI chip, achieves real-time perception and rapid control at the device level; the edge-side decision-making unit aggregates regional data and performs regional collaborative scheduling through multi-core processors and hardware acceleration engines; and the cloud-based control platform, based on a heterogeneous computing power cluster, completes global strategy optimization and model training. This forms an intelligent closed loop where perceived data converges from the bottom up and decision commands are executed from the top down. By constructing a three-level collaborative closed-loop decision-making system of global optimization, regional collaboration, and local execution, this invention achieves deep integration of hardware architecture, decision-making mechanisms, and algorithm models, fundamentally solving the problems of high response latency, high risk of single-point failure, low resource utilization efficiency, and poor system scalability inherent in traditional cloud-based centralized processing.

[0069] The cloud-based control platform, serving as the global decision-making hub of the entire system, plays a crucial role in providing the core computing power and policy support environment for the collaborative control of multi-level intelligent agents across the cloud, edge, and terminal layers. To achieve real-time processing of massive amounts of multi-source heterogeneous data and centralized training of complex strategies, the platform preferably employs a GPU / TPU heterogeneous computing cluster and a large-scale distributed storage system as its hardware foundation. Based on this infrastructure, the cloud-based control platform realizes core functions such as multi-dimensional situational awareness, multi-level collaborative optimization scheduling, and multi-scale grid-friendly interaction.

[0070] The autonomous decision-making system of multi-level intelligent agents in urban areas proposed in this invention includes: edge decision-making units, side decision-making units, and cloud control platform; in the integrated energy system, one urban area corresponds to one side decision-making unit, and each device in one urban area corresponds to one edge decision-making unit.

[0071] In this embodiment, the edge decision unit is used to perceive the operating data of various devices in the urban area in real time and make local lightweight decisions.

[0072] Various types of equipment in urban areas include, but are not limited to: intelligent circuit breakers, distributed photovoltaic controllers, energy storage converters, charging piles, and environmental control devices;

[0073] As the underlying sensing and execution terminal of the urban area integrated energy system, the edge decision unit has an integrated hardware architecture designed to meet the needs of real-time sensing, local intelligent decision-making and reliable collaborative control in resource-constrained environments. The architecture integrates a multi-source sensor array, a low-power microprocessor with a dedicated AI acceleration chip, a local storage unit, a heterogeneous communication module and a high-efficiency power supply module, which together support the real-time sensing and collaborative control of more than ten types of intelligent agents such as intelligent circuit breakers, distributed photovoltaic controllers, energy storage converters, charging piles and environmental regulation devices.

[0074] The core hardware architecture of the edge decision unit includes: a multi-source sensor array, a microprocessor, a local storage unit, a communication module, and a power supply module; among which,

[0075] 1) The multi-source sensor array connects to the main control unit of each device through a standard digital interface to collect equipment operation data in real time, thus building a data foundation for accurate equipment status assessment and real-time anomaly diagnosis; the multi-source sensor array includes sensors for electrical parameters, equipment status, environmental parameters, etc.

[0076] 2) The microprocessor integrates a neural network processing unit, which has a built-in lightweight model. Based on the lightweight model, it extracts device state variables from real-time collected device operation data and determines whether there are abnormal events based on the device state variables. The low-power microprocessor provides ultra-low standby power consumption and real-time control capabilities. At the same time, the integrated neural network processing unit (NPU) significantly improves the inference efficiency of the lightweight model locally. It is the computing power core for realizing device-level autonomous decision-making and rapid response to anomalies.

[0077] 3) The local storage unit adopts a hierarchical storage design. In addition to storing device firmware, compressed lightweight model parameters and code, it has built-in high-speed SRAM to cache intermediate calculation results of the model to improve inference speed, and uses FRAM or EEPROM to reliably store key configuration parameters and short-term data, ensuring data persistence and fast recovery capability in abnormal situations such as power outages.

[0078] 4) The end-side decision unit uploads device status variables and abnormal events determined based on the device status variables to the edge-side decision unit through the communication module; the heterogeneous communication module integrates a low-power wide area network module, which is optimized for uploading device status variables and alarm information to the edge-side decision unit with low bandwidth, long distance and low power consumption; at the same time, the optional Wi-Fi module facilitates device configuration and local debugging, and supports establishing direct communication links with adjacent end-side devices when necessary to form a local collaborative group;

[0079] 5) The high-efficiency power module supports wide voltage input and provides multiple turn-off power domains to achieve fine dynamic power consumption management, enabling it to adapt to complex field environments without mains power or with intermittent power supply, and ensuring long-term continuous and stable operation.

[0080] The collaborative work of the aforementioned hardware modules enables the edge decision unit to become an autonomous decision-making node with local intelligence and edge collaboration capabilities.

[0081] like Figure 1 As shown, the edge decision unit serves as a regional-level collaborative processing hub, with its hardware architecture specifically configured to undertake multi-agent collaborative scheduling and rapid optimization decision-making tasks. The edge decision unit possesses the capability for unified access and collaborative scheduling of more than ten types of heterogeneous intelligent agents, including photovoltaic inverters, energy storage systems, charging piles, and building energy management systems. This aims to achieve collaborative optimization and flexible control of the urban regional integrated energy system under extreme events across multiple time scales.

[0082] The core hardware architecture of the edge decision unit includes: a processing module, a storage module, a communication module, and a power supply module; among which,

[0083] 1) The processing module performs weighted aggregation of device state variables to obtain a global baseline state; it reads the deviation detection rules in the storage module to detect the deviation between each device state variable and the global baseline state; when no abnormality is detected in the deviation between each device state variable and the global baseline state, it reads the lightweight decision rules in the storage module and generates control instructions for each end-side decision unit based on the lightweight decision rules; when an abnormality is detected in the deviation between each device state variable and the global baseline state, it generates a request for the cloud control platform to update the strategy network parameters; the processing module adopts a multi-core embedded processor with an integrated Dynamically Reconfigurable Processor (DRP) hardware acceleration engine. This design is optimized through instruction-level parallelism and data pipeline optimization to improve the throughput and real-time performance of multi-node collaborative computing tasks, thereby supporting the real-time generation and release of hierarchical group control strategies.

[0084] 2) The storage module stores deviation detection rules and lightweight decision rules; the storage module adopts a three-level lightweight architecture for edge collaboration scenarios: it uses a high-speed SRAM ring buffer to realize low-latency caching of real-time data streams on the edge; it is equipped with ≥4GB LRDDR4x running memory to match the running requirements of the pruned and quantized lightweight model; it uses ≥64GB eMMC storage to solidify pre-trained model parameters and scheduling policy library, providing support for fast policy switching and system autonomy in extreme scenarios.

[0085] 3) The edge decision unit broadcasts control commands to the end decision unit through the communication module, and sends a request to update the policy network parameters to the cloud control platform through the communication module. The communication module adopts a heterogeneous communication design, integrates a LoRaWAN gateway to support low-cost access of a large number of end decision units in the area, and is equipped with an NB-IoT uplink module, which is dedicated to reliably transmitting high-value feature data and abnormal alarms to the cloud control platform.

[0086] 4) The power module supports wide voltage input and is configured with redundant power supplies to ensure continuous and stable operation under voltage fluctuations or extreme environments.

[0087] Under this architecture, the edge decision-making unit and the end decision-making unit complement each other: the end decision-making unit focuses on low-power perception and local lightweight decision-making at the single device level; while the edge decision-making unit, with its stronger heterogeneous computing capabilities, more complex multi-source data processing capabilities and higher communication and storage capacity, focuses on regional-level collaboration and can form a multi-agent collaborative rapid optimization scheduling strategy for urban regional integrated energy systems to cope with extreme events at multiple time scales.

[0088] Both the edge-side decision-making units and the cloud-based control platform communicate with each other. The cloud-based control platform includes a heterogeneous computing power cluster and a communication module. The heterogeneous computing power cluster learns policy network parameters in the state space and action space, aiming to maximize the reward function and based on a collaborative stability judgment mechanism. Among them, the reward function is established based on the global consistency variable, using the energy efficiency improvement of urban areas, network communication overhead, and control command fluctuations. The lightweight model is pruned and reconstructed using the learned policy network parameters. The cloud-based control platform feeds back the pruned and reconstructed lightweight model to each edge-side decision-making unit through the communication module.

[0089] The cloud-based control platform proposed in this invention is designed around the collaborative needs of lower-level edge and terminal intelligent agents. It aggregates multi-source heterogeneous data extracted and compressed from the edge through a distributed storage system, and utilizes heterogeneous computing power clusters for deep feature extraction and global state analysis oriented towards multi-agent collaboration. Simultaneously, its core function lies in using algorithms such as multi-agent deep reinforcement learning, based on global state information, to centrally train cloud-edge-terminal collaborative strategies and model high-fidelity virtual power plants. The cloud-based control platform dynamically allocates resources for collaborative training tasks of different priorities through elastic resource scheduling and containerized deployment, and achieves seamless and rapid distribution and version management of the trained strategies to the edge. Ultimately, through the aforementioned customized hardware architecture and operating mechanism, the cloud-based control platform, together with the lower-level edge and terminal decision-making units, constitutes a closed-loop, efficient collaborative intelligent system, realizing comprehensive perception, centralized optimization, and collaborative control of the urban regional integrated energy system.

[0090] The present invention constructs a cloud-edge-device communication architecture for hierarchical collaboration among multiple agents, and relies on a low-latency, high-reliability transmission mechanism to achieve orderly data interaction and intelligent decision-making closed loop across the entire domain.

[0091] In this embodiment, the edge-side decision unit extracts key features of the operational data and generates alarm information, which is then asynchronously uploaded to the edge-side decision unit via a low-power wide-area network (LPWAN). The system flexibly selects different LPWAN technologies based on the specific application scenario: for scenarios with long transmission distances, LoRaWAN technology is used, achieving a transmission distance of 2-5km in urban environments and exceeding 15km in line-of-sight environments; for scenarios with medium transmission distances and low latency, Zigbee technology is selected, with a transmission distance of 10-100 meters and end-to-end latency typically below 100ms, enabling effective data extraction and low-power transmission at the source.

[0092] After key features and alarm information reach the edge decision unit, the system aggregates and compresses the data from multiple nodes, and then uploads it to the cloud control platform via the uplink. The cloud control platform generates control commands and model parameters based on the global optimization strategy, and distributes them to the end decision unit via the edge decision unit, forming an efficient closed loop of "perception-aggregation-decision-control".

[0093] In terms of communication protocols, the system adopts a lightweight application protocol to achieve cross-level command and data transmission.

[0094] For communication between the endpoint and the edge, the CoAP (Constrained Application Protocol) based on UDP is selected. It uses only a 4-byte binary message header, supports asynchronous communication and acknowledgment retransmission mechanism, and its RESTful architecture style facilitates resource operation on resource-constrained devices. In scenarios requiring higher reliability (such as communication between the edge and the cloud), the MQTT protocol based on TCP can be selected.

[0095] Furthermore, the system effectively reduces communication load by relying on data compression and differential update technologies. All transmission processes are based on two-way authentication and link encryption mechanisms (such as CoAP combined with DTLS to achieve secure transmission), ensuring the integrity, confidentiality, and availability of critical data and control signaling, thereby supporting stable collaboration and autonomous decision-making among multiple agents in complex environments.

[0096] This invention also proposes an operational method for an autonomous decision-making system of multi-level intelligent agents in urban areas, such as... Figure 2 As shown, it includes:

[0097] Step 1: Establish a multi-level intelligent agent collaborative technology architecture for the urban area integrated energy system, including: cloud control platform, edge decision-making unit, and edge decision-making unit; one urban area corresponds to one edge decision-making unit in the integrated energy system, and each device in one urban area corresponds to one edge decision-making unit.

[0098] Specifically, the multi-level intelligent agent of the urban regional integrated energy system, based on the aforementioned hardware architecture, can achieve hierarchical autonomous decision-making and control. Among them, the cloud performs long-term strategy optimization and model training based on global information; the edge relies on its regional coordination capabilities to perform collaborative scheduling and decision-making based on regional multi-node state aggregation and rule base; and the end relies on its local intelligence to perform rapid response and control based on local real-time data and lightweight models. The three form a differentiated decision-making system with close collaboration between "global optimization - regional coordination - local execution".

[0099] Step 2: Each end-side decision unit broadcasts the device state variables to the side-side decision unit.

[0100] Specifically, each edge decision unit integrates a unified lightweight model; based on the lightweight model, the edge decision unit extracts the equipment state variables from the real-time collected equipment operation data and broadcasts the equipment state variables to the edge decision unit.

[0101] During operation, the first Each edge decision unit broadcasts its respective cycle via LoRaWAN. state variables Meanwhile, the side decision-making units receive data periodically to update the memory matrix. , The number of end-side decision-making units is used to enable the end-side autonomous decision-making process, in which each end-side decision-making unit broadcasts the corresponding equipment state variables to the side-side decision-making unit.

[0102] This invention proposes an end-side autonomous decision-making process that includes rule-based direct control and model-based intelligent decision-making. A pre-built rule base supports real-time matching of abnormal states such as overcurrent, and directly outputs control signals to the actuator through the GPIO / PWM interface, supporting on / off switching of digital quantities, analog quantity adjustment, and PID parameter write-back. A locally integrated lightweight model performs real-time preprocessing and feature extraction on the collected data, generating high-information-density feature vectors, which are used as inference inputs to support intelligent responses of the device under normal or abnormal conditions.

[0103] Step 3: The side-side decision unit determines the global reference state based on all device state variables, and uses the deviation between each device state variable and the global reference state to generate control commands for each end-side decision unit.

[0104] Specifically, the edge decision unit performs weighted aggregation of device state variables to obtain a global baseline state; based on deviation detection rules, it detects the deviation between each device state variable and the global baseline state; when no abnormality is detected in the deviation between each device state variable and the global baseline state, it generates control commands for each edge decision unit based on lightweight decision rules; when an abnormality is detected in the deviation between each device state variable and the global baseline state, it generates a request for the cloud control platform to update the lightweight model.

[0105] Specifically, step 3 includes:

[0106] Step 3.1: The side-side decision unit performs weighted aggregation of the device state variables to obtain the global baseline state;

[0107] The side decision-making unit utilizes a topology weight matrix pre-loaded from the cloud-based control platform. The first in the matrix Weights Reflects the first The importance and reliability of each end-side decision-making unit in regional coordination are determined by the weighted state of each end-side decision-making unit calculated in parallel by the DRP accelerator. The processor then performs aggregation operations to generate a global baseline state that characterizes the overall operational status of the region, as shown in the following equation:

[0108]

[0109] In the formula, For period The global baseline state, The number of edge decision units;

[0110] Step 3.2: Obtain the deviation between each device's state variables and the global baseline state, and perform detection based on the deviation detection rules; when no abnormality is detected in the deviation between each device's state variables and the global baseline state, proceed to step 3.3; when an abnormality is detected in the deviation between each device's state variables and the global baseline state, proceed to step 3.4.

[0111] The global baseline state calculated by the method proposed in this invention is no longer a fixed value, but a predicted value maintained by the side decision-making unit. Therefore, the detection rules for abnormal deviations between the state variables of each device and the global baseline state include:

[0112] 1) If all deviations show a consistent deviation, then the system is determined to have been subjected to a common, global disturbance;

[0113] 2) If any deviation is different from the other deviations, the equipment corresponding to any deviation is determined to be faulty;

[0114] Step 3.3: The side-side decision unit generates control commands for each end-side decision unit based on lightweight decision rules; after receiving the control commands, each end-side decision unit controls the corresponding device to execute the control commands.

[0115] Steps 2 and 3 of this invention implement a complete edge-side rapid collaborative decision-making process. Its core lies in combining the global strategy of the cloud-based control platform with the real-time situation at the edge, achieving intelligent autonomy at the city / region level. The edge-side decision-making unit achieves distributed autonomous decision-making based on weighted aggregation of state variables from multiple edge devices and a lightweight rule engine. Its autonomous decision-making system includes model loading, state awareness, collaborative computation, instruction generation, and anomaly handling. During system deployment, the edge-side decision-making unit receives the initial rule base, the topology weight matrix set representing the coupling relationship between devices, and lightweight aggregation model parameters from the cloud-based control platform via NB-IoT, storing them in the eMMC for dynamic loading at runtime. The edge-side decision-making unit loads lightweight decision rules customized for the current scenario from the eMMC into memory and executes a decision-making step that integrates the global baseline state and individual differences, generating the next cycle control instruction for each edge-side decision-making unit, as shown in the following formula:

[0116]

[0117] In the formula, For the first Each end-side decision unit in the cycle +1 control command; This refers to the built-in function of the lightweight decision rule engine, specifically the if-then rule base stored in eMMC in the example. The adjustment factor is typically set to 0.2 to 0.5 in the examples;

[0118] In this embodiment, to capture local anomalies in real time, the first anomaly is calculated synchronously during state aggregation. Deviation of each end-side decision unit The system performs anomaly detection on deviations and uploads anomaly summaries and key data to the cloud control platform based on lightweight decision rules, requesting further updates to the strategy network parameters. The architecture proposed in this invention effectively distributes the cloud computing pressure through the above closed-loop process, supports regional self-consistent operation in the event of network outages, and ensures the real-time performance and reliability of the system's decision-making in complex environments.

[0119] The side decision unit will use LoRaWAN to... Once the command is sent to the corresponding node, the device on the endpoint will immediately execute the control command.

[0120] Step 3.4: The edge decision-making unit generates a request for the cloud control platform to update the lightweight model; the cloud control platform establishes a state space and an action space, and establishes a reward function using the urban area energy efficiency improvement, network communication overhead, and control command fluctuations; with the goal of maximizing the reward function, the platform learns policy network parameters in the state space and action space based on a collaborative stability judgment mechanism; the cloud control platform uses the learned policy network parameters to prune and reconstruct the lightweight model, and feeds back the pruned and reconstructed lightweight model to each edge decision-making unit.

[0121] To overcome the bottlenecks of traditional centralized optimization in terms of response speed and computational load, this invention designs a hierarchical decision-making mechanism. In this mechanism, a cloud-based control platform performs long-term policy learning and model training across regions. The cloud-based control platform obtains the state variables of each device from each edge decision unit to establish a state space; it establishes a reward function based on the regional energy efficiency improvement, network communication overhead, and control command fluctuations at each time point; and then, it employs an adapted multi-agent deep reinforcement learning mechanism as the core of the autonomous decision-making and model training mechanism to obtain the action space.

[0122] Specifically, a state space is established to characterize the global operating state of the system. Defined as:

[0123]

[0124] In the formula, This is the device state vector. For the first The device state variables corresponding to each end-side decision unit , The number of edge decision units, for The set of real numbers;

[0125] The action space of the established representation system's coordinated regulation capability Defined as:

[0126]

[0127] In the formula, To control the command vector, For the first Regulatory commands, , To control the number of commands, for The set of real numbers;

[0128] The reward function designed in this invention Taking into account the three conflicting objectives of energy efficiency, communication costs, and control stability in integrated energy system coordination, its expression is as follows:

[0129]

[0130] In the formula, For a moment The reward function, , , These are all weighting coefficients, and after extensive simulation verification, their optimal values ​​are typically as follows: , , This range of values ​​can effectively balance the relationship between system energy efficiency, communication overhead and control stability. , They are time points The equipment status variables and control commands, Based on and The amount of energy efficiency improvement in urban areas Based on and network communication overhead, Based on and The fluctuation amount of the control command;

[0131] Within the above framework, the cloud-based control platform serves as the global decision-making center, relying on a high-performance computing cluster to achieve autonomous optimization and centralized decision-making. Based on a collaborative technology architecture, a collaborative stability judgment mechanism is introduced during the update process of the policy network parameters, as shown in the following equation:

[0132]

[0133] In the formula, To adjust the policy network parameters Find the gradient. For loss function, For the target Q value, For policy-based network parameters The current Q value, It is a norm 2. This is the error threshold;

[0134] The square of the L2 norm of the difference between the current Q value and the target Q value (i.e., the mean square error) is used as the loss function because it is smooth and differentiable and penalizes large errors, thereby improving cooperative stability.

[0135] To ensure reliable execution and closed-loop control of cloud-based strategies in complex network environments, the generated macro-control strategies are distributed to edge and end-side devices via the cloud-edge channel. Addressing potential latency, packet loss, and heterogeneity issues in the multi-layered cloud-edge-end network, the distribution process employs and improves multiple fault-tolerance and security mechanisms: strategy commands are transmitted after lightweight serialization encoding (such as CBOR format) and bidirectional authentication based on digital certificates. In case of communication anomalies, differentiated retransmission and redundancy checks are automatically triggered based on predefined service priorities. Simultaneously, as a crucial link in closed-loop control, the cloud continuously monitors the status feedback of each unit. By comparing the deviation between the expected and feedback states, once a deviation in command execution or operational anomaly is detected, a preset dynamic compensation strategy is immediately activated—including real-time adjustment of control parameters, switching to a backup strategy set, or rolling back to the most recent stable state. This constructs a reliability assurance system spanning the entire "decision-distribution-execution-feedback" chain, ensuring the system maintains stable, autonomous, and efficient operation across the entire domain under disturbances and abnormal conditions.

[0136] To achieve autonomous decision-making capabilities and overcome the strict limitations of edge decision-making units in terms of computing power, memory, and power consumption, this invention designs a lightweight model collaborative generation method of "perception-pruning-reconstruction". The core innovation of this process lies in the targeted optimization process executed before the cloud model is deployed to the edge, based on the characteristics of strong time sequence and diverse operating conditions of comprehensive energy data.

[0137] The cloud-based control platform uses the learned policy network parameters to prune and reconstruct the lightweight model. The specific process is as follows:

[0138] 1) During the process of training a lightweight model using the learned strategy network parameters, the cloud-based control platform identifies the sensitivity of neurons in each layer to various working conditions.

[0139] Specifically, calculate the first The KL divergence of the average activation values ​​of layer neurons under abnormal and steady-state conditions is used as the first... The sensitivity of layer neurons to various operating conditions, including but not limited to start-stop, steady state, and overload.

[0140] 2) Based on the sensitivity of each layer of neurons to key operating conditions, adjust the pruning threshold of each layer as shown in the following formula:

[0141]

[0142] In the formula, For the first The pruning threshold of layer neurons Based on the threshold, The adjustment factor is typically set to 0.2~0.5 in the examples. For the first The sensitivity of layer neurons to critical operating conditions.

[0143] For the basic feature extraction layer, a smaller adjustment coefficient is used to achieve low sensitivity, thereby using a higher threshold for coarse pruning to quickly compress the model size. For decision layer neurons that are strongly correlated with key abnormal states (such as overcurrent and overvoltage), a larger adjustment coefficient is used to achieve high sensitivity, thereby using a lower threshold for fine pruning to maximize the retention of their decision sensitivity. This step ensures the high reliability of the compressed model in actual operation.

[0144] 3) Determine the rank of each layer of neurons based on the energy contribution gradient of the singular values ​​of the weight matrix of neurons in each layer before and after pruning;

[0145] Energy-preserving guided low-rank reconstruction: In the low-rank decomposition stage, this invention proposes a dynamic rank determination method based on the energy contribution gradient, that is, for the pruned model... Weight matrix of layer neurons Singular Value Decomposition (SVD) yields a sequence of singular values ​​{ }, its reserved first Rank of layer neurons Determined by the following formula:

[0146]

[0147] In the formula, Let be the rank of the model after pruning. Let be the rank of the original model. The energy retention threshold is set to 0.95 in this embodiment. This method prioritizes the retention of the main components of the dominant energy data features, enabling the reconstructed model to accurately capture the core operating rules of the system while significantly reducing computational complexity.

[0148] Hardware-adapted mixed-precision quantization fully considers the hardware optimization characteristics of the edge NPU for low-bit integer operations during the quantization stage. Based on the numerical distribution range of parameters in the model after pruning and low-rank decomposition, mixed-precision quantization combining 8-bit and 4-bit methods is performed on the weights and activation values; for the weight matrix... Its quantization bit width The selection is based on the kurtosis of its numerical distribution. Decide:

[0149]

[0150] In the formula, The set kurtosis threshold is used. This hybrid strategy ensures that weights with sharp distributions and high accuracy impact receive higher bit widths, while weights with gentle distributions are aggressively compressed, achieving an optimal balance between inference speed and accuracy on the edge hardware.

[0151] The lightweight model generated through the aforementioned collaborative process has computational characteristics that are highly compatible with the hardware architecture of the edge NPU, effectively balancing the relationship between model accuracy, inference speed, and energy consumption. This provides core algorithmic support for the edge to complete highly reliable local decisions within milliseconds. Furthermore, in abnormal situations such as communication interruptions, the unit can maintain basic operation based on the last valid parameters stored in FRAM, and combined with the power module's dynamic shutdown mechanism for non-core circuits, achieve autonomous power reduction and continuous operation of the system under extreme conditions. Moreover, this invention only calls the cloud control platform to reconstruct the lightweight model when there are abnormal deviations between the state variables of each device and the global baseline state, thereby improving the reliability of the lightweight model in identifying device state variables. This allows the edge decision unit to accurately identify device state variables using the reconstructed lightweight model in subsequent steps.

[0152] Step 4: After the equipment executes the control command, each end-side decision unit updates the equipment status variables based on the real-time operating data of the equipment. When the end-side decision unit determines that an abnormal event has occurred based on the updated equipment status variables, it predicts the decision variables of each equipment based on the collaborative optimization model of the equipment target and the urban area target.

[0153] Specifically, each end-side decision unit uses a lightweight model to extract the state variables of each device from the operating data after each device executes control commands; each end-side decision unit iteratively updates the state variables of each device based on the comprehensive weight factor issued by the side-side decision unit; when the end-side decision unit sends an abnormal event detected based on the updated device state variables to the side-side decision unit, the side-side decision unit uses the subgradient method to perform distributed solution of the collaborative optimization model under joint constraints to obtain the decision variables of each device.

[0154] Specifically, step 4 includes:

[0155] Step 4.1: The side-side decision unit calculates the comprehensive weighting factor based on the electrical coupling factor and the communication quality factor, and distributes it to each end-side decision unit. The calculation formula is as follows:

[0156]

[0157] In the formula, For equipment and The combined weighting factor between them For equipment and The equivalent electrical impedance between them is determined by the system topology parameters, reflecting the tightness of the physical connection; For equipment The maximum adjustable power gives higher weight to devices with strong adjustment capabilities; For equipment and Historical communication success rate between them For equipment The set of neighboring devices within the communication topology;

[0158] The mechanism for calculating the comprehensive weighting factor proposed in this invention enables devices with short electrical distances, high adjustment potential, and reliable communication to dominate in the collaboration, thereby guiding the system state to converge toward the optimal equilibrium point that conforms to the physical laws of the power grid and has high reliability, thus significantly improving the efficiency and quality of distributed collaboration.

[0159] Step 4.2: Each end-side decision unit uses a lightweight model to extract the state variables of each device from the operational data after each device executes control commands; each end-side decision unit iteratively updates the state variables of each device based on a comprehensive weighting factor, as shown in the following formula:

[0160]

[0161] in, For equipment In the cycle The +1 state variable is the output power setting value used in this embodiment; The end-side decision unit uses a lightweight model to extract equipment data from the operational data after each device executes control commands. In the cycle State variables;

[0162] This invention addresses the diverse equipment types and volatile topologies inherent in urban integrated energy systems. To achieve rapid self-organization and collaborative operation of equipment clusters, it proposes a dynamic consensus and collaboration method based on electrical and communication dual-factor weighting. This method significantly improves upon classic distributed consensus algorithms by deeply binding their state variables and weighting mechanisms to the physical characteristics of the power system. First, the method constructs a communication topology between equipment based on their electrical connections and communication relationships. Each end-side device periodically broadcasts its own key state information and receives state data from neighboring devices, enabling plug-and-play and rapid self-collaborative optimization for equipment clusters. The state variable is explicitly defined as its "equivalent adjustable power", which integrates real-time output and adjustable variables. Through multiple iterations, the state of all devices can gradually become consistent within seconds, ultimately enabling the device group to autonomously achieve multi-objective coordination such as cluster power balance and voltage stability without relying on real-time intervention from upper-level collaborative units, significantly improving the system's distributed autonomy and response speed.

[0163] Step 4.3: When the end-side decision unit detects an abnormal event based on the updated state variables, it sends an abnormal event alarm to the side-side decision unit. The side-side decision unit uses the subgradient method to perform distributed solution of the collaborative optimization model under joint constraints to obtain the decision variables of each device.

[0164] Edge-end collaboration achieves efficient distributed closed-loop control through an event-triggered mechanism. The edge-side decision unit issues lightweight control strategies based on a rule engine to the end-side decision unit. The end-side decision unit continuously monitors the operating status based on local multi-source sensor data. When events such as exceeding limits, faults, or policy update requirements are detected, abnormal events are automatically reported. To address the system state abrupt changes and source load uncertainties introduced by event triggering, the edge-side decision unit adopts an event-driven distributed robust optimization method based on multi-end event information and regional collaboration requirements. The collaborative optimization model of this method can be used to balance individual device optimization and regional collaborative stability. Therefore, this invention proposes a collaborative optimization model for the edge-side decision unit based on device objectives and urban area objectives to predict the decision variables of each device. The collaborative optimization model is as follows:

[0165]

[0166] In the formula, For equipment The objective, as described in this embodiment, is to minimize operating costs or operating losses as the equipment objective. For equipment Adjacent equipment The collaborative cost function between them represents the urban area target. In the embodiment, the collaborative cost function in terms of power balance, voltage coordination, etc. is adopted. The collaborative weighting coefficient is initialized in the cloud and updated through edge learning, and is used to dynamically balance local targets and system targets. The number of edge decision units, For equipment The set of adjacent devices.

[0167] Joint constraints include:

[0168] 1) Local operating constraints of the equipment, including physical constraints such as upper and lower power limits and ramp rate;

[0169] 2) Inter-device coupling constraints characterize the operational correlation between adjacent devices caused by electrical connections;

[0170] The optimization model is solved in a distributed manner using the subgradient method. Each edge decision unit updates its decision variables based on local information and limited information exchanged with neighboring units, according to the following iterative format:

[0171]

[0172] In the formula, For equipment In the cycle The decision variable is +1, and in this embodiment, the output power setting value is used; To reach the feasible set Projection operation; For equipment In the cycle State variables, The subgradient of the Lagrange function, For period The step size sequence;

[0173] In this embodiment, the set of decision variables is defined as follows: , Indicates device The method controls variables such as output setpoints; through local information interaction, it effectively improves the real-time performance and reliability of system control while reducing communication overhead, providing a stable and reliable execution foundation for cloud-based global optimization.

[0174] Step 5: The cloud-based control platform integrates the collaborative optimization models of all side decision units based on the decision variables of each device, and uses the integrated collaborative optimization model to predict the globally consistent variables.

[0175] Specifically, the cloud-based control platform uses the alternating direction multiplier method to fuse the collaborative optimization models of the side decision units in different regions based on the decision variables of each device, and then performs distributed solution on the fused collaborative optimization model to obtain globally consistent variables.

[0176] Multi-edge collaborative scheduling achieves resource coordination and power allocation between regions through a distributed optimization algorithm. Each edge decision unit receives the collaborative objective from the cloud and uses the Alternating Directional Multiplier Method (ADMM) for distributed solution: each edge decision unit first initializes its local decision variables (such as adjustable load power and energy storage output), and exchanges boundary coupling information (such as planned values ​​for tie-line transmission power) with adjacent edge decision units through a communication network; the fused collaborative optimization model is as follows:

[0177]

[0178] In the formula, For period Globally consistent variables, For equipment In the cycle dual variables, For equipment In the cycle Decision variables, For equipment Target, For collaborative weighting coefficients;

[0179] The globally consistent variables obtained by solving are:

[0180]

[0181] The dual variable is then updated as follows:

[0182]

[0183] The iterative process continues until the difference between the decision variables of all devices and the globally consistent variable is less than the set tolerance. , The error threshold is set according to the control accuracy requirements; at this point, the system converges to the global optimal or near-optimal solution; after multiple iterations of convergence, each side decision unit achieves the collaborative goal without centralized optimization, realizing cross-regional economic scheduling and power mutual assistance.

[0184] Step 6: Establish a reward function based on the globally consistent variables. Under the constraints of system operation, predict the control instructions of each side decision unit according to the global dynamic optimization objective established by the reward function. The side decision unit decomposes the control instructions into control instructions of each end decision unit.

[0185] Specifically, the cloud-based control platform uses a reward function based on globally consistent variables to establish a global dynamic optimization objective; it adopts a rolling optimization model predictive control framework to iteratively solve the global dynamic optimization objective under system operation constraints to obtain predicted control instructions; and it sends the predicted control instructions to the edge decision units through an encrypted channel; the edge decision units use a distributed model predictive control framework to decompose the control instructions into control instructions for each edge decision unit.

[0186] Specifically, step 6 includes:

[0187] Step 6.1: The cloud-based control platform uses a reward function based on a globally consistent variable and a multi-agent deep reinforcement learning method to establish a global dynamic optimization objective;

[0188] Specifically, based on the globally consistent variables obtained in step 5, the cloud-based control platform further employs a multi-agent deep reinforcement learning method to establish a global dynamic optimization objective to achieve long-term optimization; based on the timing of the globally consistent variables... reward function , , Based on globally consistent variables in the state space and action space The determined time No. Each terminal-side decision unit corresponds to the equipment state variables and control commands; reward function. Characterizing time Based on globally consistent variables, the urban area's energy efficiency improvement, network communication overhead, and control command fluctuations; utilizing a reward function. Construct a global dynamic optimization objective representing future cumulative rewards as follows:

[0189]

[0190] In the formula, The policy network parameters are learned by the cloud-based control platform. The policy network parameters are introduced here to establish a correspondence between the lightweight model used in practice and the global dynamic optimization objective, which facilitates improved computational efficiency when calling the lightweight model in the future. As a discount factor, This is used to balance immediate rewards with long-term benefits; To optimize cycle length; Let it be the expected function;

[0191] Based on the time corresponding to the globally consistent variables The global dynamic optimization objective is established based on the reward function, forming a hierarchical closed-loop optimization system.

[0192] Step 6.2: The cloud-based control platform adopts a rolling optimization model predictive control framework to iteratively solve the global dynamic optimization objective under system operation constraints, and obtain the predicted control instructions.

[0193] The rolling optimization model predictive control framework aims to minimize system operating costs and maximize energy efficiency. It generates predictive control commands through centralized training. System operating constraints include local device operating constraints (such as power upper and lower limits, ramp rate, etc.) and inter-device coupling constraints.

[0194] Step 6.3: The cloud-based control platform sends the predicted control instructions to the edge decision-making units through an encrypted channel; the edge decision-making units adopt a distributed model predictive control framework to decompose the control instructions into control instructions for each edge decision-making unit.

[0195] The cloud-based control platform generates macro-level instructions such as demand response strategies and unit combination schemes, and distributes them to each edge-side decision-making unit through a cloud-edge encrypted channel. After receiving the instructions, the edge-side decision-making unit decomposes the global instructions into executable equipment group control sequences based on the real-time regional operating status, and distributes them to terminal devices through the NB-IoT network. The terminal devices convert the received control codes into specific actuator drive signals to achieve precise energy regulation and ensure the step-by-step implementation of the control intentions.

[0196] Furthermore, the autonomous decision-making system proposed in this invention possesses a three-level fault-tolerance mechanism to cope with network anomalies: edge devices can autonomously activate rule-based emergency control strategies to maintain basic operation; edge nodes utilize cached historical strategies and regional state estimates to maintain regional coordination; the cloud generates compensation strategies through historical data analysis, which are incrementally synchronized to the terminal after communication is restored; through the above multi-layered optimization architecture and closed-loop control mechanism, a complete optimization chain is realized from global decision-making in the cloud to precise execution on the edge. This system not only ensures the overall efficiency of global optimization but also maintains the real-time performance and reliability of local operations through layered decoupling and distributed control, ultimately achieving efficient, stable, and intelligent operation of the urban integrated energy system across multiple time scales.

[0197] The urban integrated energy system adopts a multi-layered collaborative control architecture. It achieves power and voltage coordination of equipment groups through a multi-terminal distributed consensus algorithm. The edge-end side achieves real-time closed-loop control based on event triggering and distributed robust optimization. The multi-edge side completes inter-regional resource coordination through the ADMM algorithm. The cloud-edge-end side relies on multi-agent reinforcement learning and model predictive control to achieve global optimization decision-making. Combined with federated learning and fault tolerance mechanisms, it ensures the stable and efficient operation of the system at multiple time scales.

[0198] In terms of decision-making efficiency, traditional systems rely on centralized cloud processing, which is difficult to meet the needs of high-frequency data concurrency and real-time control. This invention sinks lightweight intelligent decision-making capabilities to the edge and end sides, combines the edge-side NPU and dedicated AI chip to achieve millisecond-level local response, and uses the edge-side DRP accelerator to complete the parallel aggregation and collaborative computing of regional data, achieving second-level regional scheduling. This significantly reduces the overall system latency and cloud load, and demonstrates rapid self-healing and elastic control capabilities in extreme event scenarios with extremely high real-time requirements.

[0199] Regarding system reliability, existing centralized architectures suffer from single points of failure, with communication interruptions easily leading to system paralysis. This invention integrates end-side FRAM persistent storage, edge-side eMMC policy library, and multi-level power management mechanisms at the hardware level, and constructs a three-level closed-loop fault-tolerant strategy at the software level. This enables the end-side to maintain basic operation based on the rule base even in the event of a network outage, and allows the edge-side to achieve regional autonomy, thus maintaining continuous and stable system operation even in complex network environments. The policy update mechanism with stability judgment introduced in the cloud further enhances the robustness of the decision-making process.

[0200] In terms of resource coordination, traditional methods decouple algorithms from hardware, leading to low coordination efficiency. This invention designs a lightweight edge-side model process of "structured pruning—low-rank approximation—nonlinear quantization," ensuring a high degree of model-hardware matching. The topology weight matrix used in edge-side coordination is dynamically generated by the cloud based on device coupling relationships, improving the accuracy of regional scheduling. This design significantly reduces edge-side resource consumption and extends device endurance while ensuring decision-making accuracy. Furthermore, through simulation-optimized weight coefficients in the reward function, the system balances multiple conflicting objectives such as energy efficiency, communication overhead, and control stability.

[0201] Regarding system scalability and adaptability, existing technologies lack efficient plug-and-play and self-organizing mechanisms to address the challenges of heterogeneous devices and dynamic topology changes. This invention achieves rapid self-coordination of device groups through a distributed consensus algorithm and employs event triggering and distributed robust optimization to address system uncertainties. This allows new devices to automatically integrate into the collaborative network and achieve autonomous power and voltage balancing without upper-layer intervention, significantly improving the system's deployment flexibility and environmental adaptability in large-scale, dynamic scenarios.

[0202] In summary, this invention has made systematic innovations at multiple levels, from architecture design and hardware implementation to algorithm collaboration, effectively improving the overall performance of urban integrated energy systems in terms of real-time response, reliable operation, resource optimization, and scale expansion.

[0203] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0204] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0205] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0206] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. An autonomous decision system for urban area multi-level agents, characterized in that, Comprise: End-side decision unit, edge-side decision unit, cloud-side regulation platform; Each end-side decision unit broadcasts device state variables to the edge-side decision unit; The edge-side decision unit determines the global reference state according to all device state variables, and generates control instructions for each end-side decision unit using the deviation of each device state variable from the global reference state; After the device executes the control instruction, each end-side decision unit updates the device state variable according to the real-time running data of the device and determines whether an abnormal event occurs based on the updated device state variable, and the edge-side decision unit predicts the decision variable of each device based on the collaborative optimization model of the device target and the city area target; The cloud-side regulation platform fuses the collaborative optimization model of all edge-side decision units according to the decision variable of each device, and predicts the global consistency variable using the fused collaborative optimization model; Based on the global consistency variable, a reward function is established, and based on the global dynamic optimization target established by the reward function under the system operation constraint, the regulation instruction of each edge-side decision unit is predicted; The edge-side decision unit decomposes the regulation instruction into control instructions for each end-side decision unit.

2. The autonomous decision system of the multi-level agent of the city area according to claim 1, wherein One city area in a comprehensive energy system corresponds to one edge-side decision unit, and each device in a city area corresponds to each end-side decision unit.

3. The autonomous decision system of the multi-level agent of the city area according to claim 2, wherein The end-side decision unit comprises: a multi-source sensor array, a microprocessor, and a communication module; The multi-source sensor array is connected to the main control unit of each device through a standard digital interface to collect device running data in real time; The microprocessor has a neural network processing unit integrated therein, the neural network processing unit has a lightweight model built-in, the device state variable is extracted from the real-time collected device running data based on the lightweight model, and it is judged whether there is an abnormal event based on the device state variable; The end-side decision unit uploads the device state variable and the abnormal event determined based on the device state variable to the edge-side decision unit through the communication module.

4. The autonomous decision system of the multi-level agent of the city area according to claim 3, wherein The cloud-side regulation platform comprises: a heterogeneous computing cluster and a communication module; the heterogeneous computing cluster learns to obtain policy network parameters in the state space and the action space with the maximum reward function as the target based on the collaborative stability judgment mechanism; wherein, based on the global consistency variable, the city area energy efficiency improvement, the network communication overhead and the control instruction fluctuation are used to establish the reward function; the policy network parameters learned are used to prune and reconstruct the lightweight model; the cloud-side regulation platform feeds back the pruned and reconstructed lightweight model to each end-side decision unit through the communication module.

5. The autonomous decision system of the multi-level agent of the city area according to claim 4, wherein The core of the hardware architecture of the edge-side decision unit comprises: a processing module, a storage module, and a communication module; The storage module stores the deviation detection rule and the lightweight decision rule; The processing module aggregates the device state variables to obtain a global reference state; the deviation detection rule in the storage module is read to detect the deviation of each device state variable from the global reference state; when it is detected that the deviation of each device state variable from the global reference state is not abnormal, the lightweight decision rule in the storage module is read, and the control instruction of each end-side decision unit is generated based on the lightweight decision rule; when it is detected that the deviation of each device state variable from the global reference state is abnormal, a request for the cloud regulation platform to update the lightweight model is generated; The edge-side decision unit broadcasts the control instruction to the end-side decision unit through the communication module, and sends a request for updating the strategy network parameters to the cloud regulation platform through the communication module.

6. A method for operating an autonomous decision system suitable for use in a multi-level agent for an urban area according to any one of claims 1 to 5, characterized in that, It comprises: A unified lightweight model is integrated in each end-side decision unit; Based on the lightweight model, the end-side decision unit extracts the device state variables from the real-time collected device operation data, and then broadcasts the device state variables to the edge-side decision unit; The edge-side decision unit aggregates the device state variables to obtain a global reference state; based on the deviation detection rule, the deviation of each device state variable from the global reference state is detected; when it is detected that the deviation of each device state variable from the global reference state is not abnormal, the control instruction of each end-side decision unit is generated based on the lightweight decision rule; when it is detected that the deviation of each device state variable from the global reference state is abnormal, a request for the cloud regulation platform to update the lightweight model is generated; Each end-side decision unit extracts the device state variables from the operation data of each device after executing the control instruction using the lightweight model; each end-side decision unit iteratively updates each device state variable based on the comprehensive weight factor issued by the edge-side decision unit; When the end-side decision unit sends an abnormal event detected based on the updated device state variable to the edge-side decision unit, the edge-side decision unit uses the sub-gradient method to solve the collaborative optimization model in a distributed manner under joint constraints to obtain the decision variables of each device; The cloud regulation platform fuses the collaborative optimization models of the edge-side decision units in different regions using the alternating direction multiplier method based on the decision variables of each device, and solves the fused collaborative optimization model in a distributed manner to obtain global consistency variables; The cloud regulation platform establishes a global dynamic optimization objective using a reward function based on the global consistency variables; a rolling optimization model predictive control framework is used to iteratively solve the global dynamic optimization objective under system operation constraints to obtain predicted regulation instructions; the predicted regulation instructions are issued to the edge-side decision unit through an encrypted channel; the edge-side decision unit uses a distributed model predictive control framework to decompose the regulation instructions into control instructions for each end-side decision unit.

7. The operation method of the autonomous decision system of the urban area multi-level agent according to claim 6, wherein The edge side decision unit loads the topology weight matrix from the cloud regulation platform The global reference state obtained is as follows: wherein is a global reference state of period , is a state variable broadcasted by the th end-side decision unit to the side-side decision units of period , is the th weight in the topology weight matrix , is the number of end-side decision units.

8. The operation method of the autonomous decision system of the urban area multi-level agent according to claim 6, wherein The deviation detection rule comprises: 1) If all deviations have consistent deviations, it is determined that the system is subject to a common, global disturbance; 2) If there is a deviation between any deviation and the remaining deviations, it is determined that the device corresponding to any deviation of the deviation exists a fault.

9. The operation method of the urban area multi-level agent autonomous decision system according to claim 7, characterized in that, The control instruction of each end-side decision unit is as follows: wherein is the first end-side decision unit in the control command of the cycle +1; is a built-in function of the lightweight decision rule engine, is a regulation coefficient.

10. The operation method of the urban area multi-level agent autonomous decision system according to claim 7, characterized in that, The edge-side decision unit generates a request for the cloud regulation platform to update the lightweight model; the cloud regulation platform establishes a state space and an action space, and establishes a reward function based on global consistency variables, urban area energy efficiency improvement, network communication overhead and control instruction fluctuation; based on the collaborative stability judgment mechanism, the strategy network parameters are learned in the state space and the action space with the maximum reward function as the target; The cloud regulation platform prunes and reconstructs the lightweight model using the learned strategy network parameters, and feeds back the pruned and reconstructed lightweight model to each end-side decision unit; wherein the time instant The reward function is as follows: In the formula, , , are weight coefficients, , are the device state variables and control instructions at time , is the urban area energy efficiency improvement amount based on and , is the network communication overhead based on and , is the control instruction fluctuation amount based on and .

11. The operation method of the urban area multi-level agent autonomous decision system according to claim 10, characterized in that, The edge-side decision unit calculates a comprehensive weight factor according to the electrical coupling factor and the communication quality factor and distributes it to each end-side decision unit, and the calculation formula is as follows: wherein, is a device with a combined weight factor between, is a device with an equivalent electrical impedance between, is a device a maximum adjustable power of, is a device with a historical communication success rate between, is a device a set of neighbor devices within a communication topology of, Each end-side decision unit extracts state variables of each device from running data after each device executes a control instruction by using a lightweight model Each end-side decision unit iteratively updates the state variables of each device based on a comprehensive weight factor, as shown in the following formula: wherein for the device in the cycle +1 state variable.

12. The operation method of the urban area multi-level agent autonomous decision system according to claim 11, characterized in that, The collaborative optimization model based on the device target and the urban area target is as follows: In the formula, is the device target; is the device adjacent device cooperative cost function between the city area target; is the cooperative weight coefficient, is the number of end-side decision units, is the device adjacent device set of the device The decision variables obtained by solving are as follows: wherein is the device is the state variable of the device is the decision variable of the device is the device is the state variable of the device is the subgradient is the step size sequence of the period 13. The operation method of the urban area multi-level agent autonomous decision system according to claim 12, characterized in that, The fused collaborative optimization model is as follows: The global consistency variables obtained by solving are as follows: wherein is a global consistent variable of period , is a device variable of period , is a device variable of period , is a local optimization objective of the th end-side decision unit, is a coordination weight coefficient; When , is a set error threshold, the system converges.

14. The operation method of the urban area multi-level agent autonomous decision system according to claim 13, characterized in that, Global dynamic optimization objective As follows: In the formula, The strategy network parameters are learned by the cloud-based control platform. As a discount factor, , To optimize cycle length, Let be the expected function. , Based on globally consistent variables in the state space and action space The determined time No. Each terminal-side decision unit corresponds to the equipment state variables and control commands. For moments based on globally consistent variables The reward function.

15. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to perform the steps of the method of any one of claims 6-14.

16. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the method of any one of claims 6-14.

Citation Information

Patent Citations

  • Virtual power plant global optimization scheduling method, system and device based on cloud edge collaboration and storage medium

    CN120955686A

  • Coordinated control method and system for comprehensive energy multi-agent coordinated group control and autonomous decision

    CN121076819A