Primary and secondary fusion complete ring main unit energy efficiency optimization system based on cloud edge collaboration

By deploying energy efficiency optimization service packages in edge computing units and combining them with cloud service orchestration, and using deep reinforcement learning models for real-time decision-making, the problems of limited computing resources at edge nodes and communication latency in wide area networks are solved, enabling real-time and accurate optimization of the energy efficiency of ring network equipment and efficient system operation.

CN121567702AActive Publication Date: 2026-02-24NANJING GREEN POWER INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202610077026.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-02-24
Estimated Expiration
2046-01-21

AI Technical Summary

Technical Problem

Existing technologies, under large-scale real-time data stream processing, suffer from limited computing resources at edge nodes and the impact of wide area network communication latency, making it difficult to deploy and execute data-driven distributed network service applications in real time. Furthermore, traditional local services lack the ability to adapt to complex operating conditions, resulting in reduced system operating efficiency and reliability.

Method used

A cloud-edge collaborative primary and secondary integrated ring network box energy efficiency optimization system is adopted. By deploying energy efficiency optimization service packages in the edge computing unit and combining them with cloud service orchestration, a deep reinforcement learning model is used for real-time decision-making, thereby achieving efficient data processing and rapid response to control commands.

Benefits of technology

While ensuring real-time performance and reliability, the energy efficiency of the equipment was optimized, the computational complexity of the edge computing unit was reduced, and the problem of limited edge resources was solved by offloading high-computing tasks to the cloud. At the same time, the safety of the power grid voltage and the accuracy of energy efficiency evaluation were ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567702A_ABST
    Figure CN121567702A_ABST
Patent Text Reader

Abstract

The invention relates to the field of edge computing, in particular to a primary and secondary fusion complete ring main unit energy efficiency optimization system based on cloud edge collaboration. The system comprises an edge device deployment and data acquisition module used for collecting operation electric parameters and device states and generating a service calling request; the service calling module is used for screening target control parameters under security constraints based on the energy efficiency optimization service package; the data synchronization module unloads the data to a cloud platform through a message queue protocol; the cloud service arrangement module aggregates the data updating service package and issues the data updating service package to the edge; and the real-time control module converts the parameters into instructions to drive equipment to act. A cloud training and edge reasoning collaborative architecture is adopted, the problems that local control lacks adaptivity, edge resources are limited and full-cloud control communication time delay exists are solved, and real-time energy efficiency optimization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of edge computing, specifically to an energy efficiency optimization system for a complete ring network box based on cloud-edge collaboration, integrating primary and secondary components. Background Technology

[0002] The integrated primary and secondary ring network enclosure serves as a crucial edge node in the energy Internet of Things (IoT), continuously collecting and generating high-frequency operational data streams. Traditional control logic is deployed on local controllers, whose decisions are based on preset, simple data thresholds, executing basic equipment control operations. This type of local service application, based on simple thresholds, lacks flexibility in data processing and command exchange.

[0003] However, the large volume of real-time data streams collected by nodes exhibits significant time-varying and nonlinear characteristics. This local service application based on simple thresholds lacks adaptability to complex operating conditions. When data fluctuations are large and the amount of data to be collected is substantial, this decision-making approach can easily lead to frequent command transmissions or failure to converge to the globally optimal system operating efficiency, thereby reducing system performance and equipment reliability.

[0004] While complex distributed application services can improve the accuracy of data-driven decision-making, they face contradictions in network architecture deployment: on the one hand, edge nodes have limited computing power and storage space, making it difficult to bear the computing load required by such service applications; on the other hand, if a centralized service architecture is adopted in the cloud, the uploading of a large number of real-time data streams will cause excessive consumption of communication bandwidth, and the latency and packet loss risks of wide area network transmission will seriously affect the real-time performance and reliability of distributed network services.

[0005] In summary, the technical problem that existing technologies need to solve is: how to overcome the limitations of edge node computing resources and wide area network communication latency under the requirements of large-scale real-time data stream processing, and realize the deployment and real-time execution of data-driven low-latency distributed network service applications.

[0006] To address this, a cloud-edge collaborative primary and secondary integrated ring network box energy efficiency optimization system is proposed. Summary of the Invention

[0007] The purpose of this invention is to provide a cloud-edge collaborative primary and secondary integrated ring network box energy efficiency optimization system. Through a hierarchical collaborative architecture of cloud model training and edge real-time inference, it solves the contradiction between insufficient edge computing power and high network transmission latency in the context of large-scale real-time data stream processing, and achieves real-time and accurate optimization of the energy efficiency of ring network box equipment.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A cloud-edge collaborative primary and secondary integrated ring network enclosure energy efficiency optimization system specifically includes the following steps: Edge device deployment and data acquisition module: Deploys the edge computing unit in the smart power distribution terminal; receives real-time power parameters of the power distribution network and status data of ring network box equipment, integrates them into status data packets; and generates corresponding service call requests based on the status data packets. Service Invocation Module: Based on the service invocation request, invokes the energy efficiency optimization service package embedded in the edge computing unit, and under the condition of meeting the preset grid voltage and load safety constraints, selects a set of data as target control parameters from the equipment control parameter space that characterizes the equipment operating status; Data synchronization module: acquires the status data packet and formats it into a synchronization payload, and offloads the synchronization payload to the energy IoT cloud platform via the MQTT message transmission protocol; Cloud service orchestration module: It collects the synchronous payloads, updates the energy efficiency optimization service package as feedback data, and distributes the updated energy efficiency optimization service package to the edge computing unit through the network according to the preset strategy; Real-time control module: Encapsulates the target control parameters into remote device control commands and sends the remote device control commands to the controlled device in real time.

[0009] Preferably, the edge computing unit is an intelligent computing terminal integrating a central processing unit, memory, and network communication interface, and has a container runtime environment deployed inside, providing isolated computing resources and a running platform for the energy efficiency optimization service package; the specific method of integrating into the status data packet includes: acquiring power distribution network operating parameters including three-phase voltage, three-phase current, active power, and reactive power, as well as ring network box equipment status data including switch on / off position, energy storage status, and ambient temperature and humidity, and encapsulating the operating parameters and the ring network box equipment status data into a data frame of structured data object, and using the data frame as the status data packet; the specific method of generating the corresponding service call request includes: generating an interface call instruction containing a power distribution terminal identifier and the status data packet, and using the interface call instruction as the service call request.

[0010] Preferably, the energy efficiency optimization service package is an independent running instance encapsulated as a containerized microservice, which integrates an energy efficiency decision model based on deep reinforcement learning and an inference environment to support the model's operation. The energy efficiency decision model adopts a deep Q-network architecture and constructs a reward function to characterize the energy efficiency evaluation index and the grid voltage and load safety constraints. The grid voltage and load safety constraints are configured as boundary penalty terms in the reward function. When the monitored bus voltage deviation rate exceeds a preset threshold, a negative penalty is applied to the energy efficiency decision model. The energy efficiency evaluation index is configured as the target reward term in the reward function, and its specific value is the weighted sum of the line loss reduction rate and the power factor improvement rate within a unit power supply cycle.

[0011] Preferably, the specific method for selecting a set of data as the target control parameters includes: obtaining the allowable adjustment range of the on-load tap changer associated with the ring main unit and the number of available switching groups of the reactive power compensation switching device; constructing their respective discrete state sets with the minimum adjustment action unit as the step size; performing Cartesian product operation on each of the discrete state sets to generate a candidate action library containing all feasible combinations as the equipment control parameter space; constructing the current state data packet into a state vector and inputting it into the energy efficiency decision model based on deep reinforcement learning; evaluating the value score of each set of parameters in the candidate action library through the energy efficiency decision model, and outputting the target control parameters according to the preset action selection strategy, wherein the target control parameters include reactive power compensation switching commands and transformer tap position commands.

[0012] Preferably, the specific method of formatting the synchronous payload includes: extracting data from the status data packet and encoding it using the protobuf serialization data format to generate the synchronous payload; the specific method of offloading to the energy IoT cloud platform includes: establishing a message subscription and publishing channel with the energy IoT cloud platform based on the MQTT message transmission protocol, and pushing the synchronous payload through the channel.

[0013] Preferably, the specific method for aggregating the synchronization payload includes: subscribing to the MQTT message transmission protocol channel on the energy Internet of Things cloud platform, receiving and parsing the synchronization payload, and storing the parsed data into the power time series database according to the power distribution terminal identifier.

[0014] Preferably, the specific method for updating the energy efficiency optimization service package includes: monitoring the data aggregation scale of the power time series database, and when the data aggregation scale meets the preset retraining threshold, extracting historical operational interaction data containing state-action pairs, and performing offline retraining on the energy efficiency decision model based on deep reinforcement learning; the historical operational interaction data includes the corresponding operational electrical parameters, ring network equipment status, historically executed target control parameters, and calculated energy efficiency evaluation indicators at the same time segment; the specific method for distributing to the edge computing unit includes: verifying the version number of the energy efficiency decision model after offline retraining, and distributing the incremental update package of the energy efficiency decision model based on deep reinforcement learning to the edge computing unit through a cloud-edge collaborative container orchestration architecture.

[0015] Preferably, the specific method of encapsulating the remote device control command includes: converting the reactive power compensation device switching command and the transformer tap position command into a command frame conforming to the Modbus-TCP protocol format; the specific method of sending the command frame to the controlled device includes: sending the command frame to the programmable logic controller in the ring network box through the industrial Ethernet interface to remotely control the terminal device via the network.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention deploys edge computing units within intelligent power distribution terminals and utilizes a cloud service orchestration module to aggregate and synchronize payloads and update energy efficiency optimization service packages, thus constructing a cloud-edge collaborative architecture. This solution offloads high-computing-power training tasks to the cloud, solving the problem of limited edge resources; simultaneously, it uses service call requests to invoke the energy efficiency optimization service package embedded in the edge computing unit for local inference, avoiding the response lag caused by network fluctuations in pure cloud control and ensuring the real-time performance of power distribution network control.

[0017] 2. The energy efficiency decision-making model described in this invention adopts a deep Q-network architecture, where the grid voltage and load safety constraints are configured as boundary penalty terms in the reward function, and the energy efficiency evaluation index is configured as the target reward term in the reward function. This scheme uses boundary penalties to force the model to avoid exceeding limits, ensuring voltage safety; simultaneously, by clearly defining the optimization direction through the weighted sum of the line loss reduction rate and power factor improvement rate within a unit power supply cycle, it achieves coordinated optimization of line loss and power factor under dynamic load conditions.

[0018] 3. This invention constructs a discrete state set with the smallest adjustment action unit as the step size, and performs a Cartesian product operation on each discrete state set to generate a candidate action library containing all feasible combinations as the device control parameter space. This scheme strictly limits the algorithm search space to the physical range, thus preventing mechanical damage to the device from a mechanism perspective; at the same time, the pre-constructed finite action library narrows the optimization range and reduces the computational complexity of the edge computing unit.

[0019] 4. This invention uses protobuf serialized data format for encoding to generate the synchronization payload, and establishes a message subscription and publish channel with the energy IoT cloud platform based on the MQTT message transmission protocol. Compared with traditional text formats, this scheme significantly compresses data volume and reduces protocol overhead, ensuring the transmission stability of data uplink and model update packet distribution in environments with limited wide-area communication bandwidth in power distribution networks. Attached Figure Description

[0020] Figure 1 This is a structural diagram of a primary and secondary integrated ring network box energy efficiency optimization system based on cloud-edge collaboration, provided in Embodiment 1 of the present invention. Figure 2 This is a flowchart of a cloud-edge collaborative primary and secondary integrated ring network box energy efficiency optimization system provided in Embodiment 2 of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figures 1 to 2 This invention provides a cloud-edge collaborative primary and secondary integrated ring network box energy efficiency optimization system, comprising: Edge device deployment and data acquisition module: Deploys the edge computing unit in the smart power distribution terminal; receives real-time power parameters of the power distribution network and status data of ring network box equipment, integrates them into status data packets; and generates corresponding service call requests based on the status data packets. Service Invocation Module: Based on the service invocation request, invokes the energy efficiency optimization service package embedded in the edge computing unit, and under the condition of meeting the preset grid voltage and load safety constraints, selects a set of data as target control parameters from the equipment control parameter space that characterizes the equipment operating status; Data synchronization module: acquires the status data packet and formats it into a synchronization payload, and offloads the synchronization payload to the energy IoT cloud platform via the MQTT message transmission protocol; Cloud service orchestration module: It collects the synchronous payloads, updates the energy efficiency optimization service package as feedback data, and distributes the updated energy efficiency optimization service package to the edge computing unit through the network according to the preset strategy; Real-time control module: Encapsulates the target control parameters into remote device control commands and sends the remote device control commands to the controlled device in real time.

[0023] Example 1 As one embodiment of the present invention, refer to Figure 1 A structural diagram of a primary and secondary integrated ring network box energy efficiency optimization system based on cloud-edge collaboration.

[0024] Furthermore, the edge computing unit is an intelligent computing terminal integrating a central processing unit, memory, and network communication interface. It internally deploys a container runtime environment, providing isolated computing resources and a running platform for the energy efficiency optimization service package. The specific method of integrating the data into a status data packet includes: acquiring power distribution network operating parameters including three-phase voltage, three-phase current, active power, and reactive power, as well as ring network box equipment status data including switch on / off positions, energy storage status, and ambient temperature and humidity; encapsulating the operating parameters and the ring network box equipment status data into a JSON format data frame, and using the data frame as the status data packet; the specific method of generating the corresponding service call request includes: generating an interface call instruction containing a power distribution terminal identifier and the status data packet, and using the interface call instruction as the service call request.

[0025] Specifically, the edge computing unit, serving as the edge-side computing power carrier of the system, adopts an industrial-grade embedded architecture at the hardware level. It is equipped with a central processing unit supporting floating-point operations, such as an industrial-grade CPU based on ARM or x86 architecture, and memory for caching real-time data and model weights, including RAM (Random Access Memory) and Flash memory. At the software level, the unit runs an embedded operating system and comes pre-installed with a container runtime environment, such as Docker Engine or Containerd, enabling the loading and execution of energy efficiency optimization service packages as independent container instances, thus decoupling application logic from the underlying hardware. The edge device deployment and data acquisition module establishes a data communication channel with the underlying sensing devices, reading large-scale real-time data from the measurement unit. This large-scale real-time data is divided into distribution network operating parameters characterizing power load characteristics and ring network equipment status data characterizing physical operating conditions.

[0026] Specifically, the module uses a pre-built serialization program to map the heterogeneous data read above into a standard key-value pair structure and encapsulate it into a JSON-formatted data frame. For example, it generates a key-value pair structure containing voltage and switch status, thus forming a status data packet. Subsequently, the module reads the unique power distribution terminal identifier stored locally, generates an interface call instruction containing the identifier and the status data packet, and explicitly points to the energy efficiency optimization service container interface deployed in the edge computing unit, using it as a service call request to trigger subsequent processes.

[0027] This invention decouples the underlying data acquisition logic from the upper-level energy efficiency optimization algorithm by encapsulating operating electrical parameters and device status data into a unified JSON format data frame and generating interface call instructions containing terminal identifiers. This technical solution establishes a standardized data interaction format, eliminating the dependence of service calls within the edge computing unit on specific underlying register addresses, thus improving the maintainability of software modules and the standardization of data parsing.

[0028] Furthermore, the energy efficiency optimization service package is an independent running instance encapsulated as a containerized microservice. Internally, it integrates a deep reinforcement learning-based energy efficiency decision-making model and an inference environment to support its operation. The energy efficiency decision-making model employs a deep Q-network architecture and constructs a reward function to characterize the energy efficiency evaluation index and the grid voltage and load safety constraints. The grid voltage and load safety constraints are configured as boundary penalty terms in the reward function; when the monitored bus voltage deviation rate exceeds a preset threshold, a negative penalty is applied to the energy efficiency decision-making model. The energy efficiency evaluation index is configured as the target reward term in the reward function, specifically a weighted sum of the line loss reduction rate and power factor improvement rate within a unit power supply cycle.

[0029] Specifically, the energy efficiency optimization service package is explicitly defined as an independently running containerized microservice instance encapsulated within the storage space of an edge computing unit, such as a Docker container, which integrates a lightweight inference engine, such as TensorFlow Lite, to support the operation of the energy efficiency decision model. The energy efficiency decision model employs a deep Q-network architecture, comprising an input layer, fully connected hidden layers using the ReLU activation function, and an output layer. The model constructs a reward function to represent energy efficiency evaluation indicators and safety constraints, which consists of a boundary penalty term and a target reward term.

[0030] Specifically, the energy efficiency evaluation index refers to the numerical feedback used to quantitatively assess the degree to which a single control action improves the system's operational efficiency. Its calculation logic is as follows: The system first calculates the reduction rate of line loss per unit power supply cycle based on the electrical parameter data before and after the action is executed, denoted as... And the power factor improvement rate, denoted as The system introduces an adaptive dynamic weighting mechanism, dynamically adjusting the first weighting coefficient based on the deviation between the real-time power factor and the assessment threshold. Second weighting coefficient The aforementioned and It is an adaptive dynamic weighting coefficient. Its value is dynamically calculated by the system based on the deviation between the real-time power factor and the preset assessment threshold. The specific calculation process is as follows: When the power factor is not lower than the assessment threshold, an economic priority strategy is implemented, assigning a corresponding line loss reduction rate through a linear function. A higher percentage of power factor is used to guide the model to prioritize loss reduction; when the power factor falls below the assessment threshold, a compliance-first strategy is adopted. This compliance-first strategy refers to the system's approach of forcing the model to prioritize power factor correction in reward feedback to ensure that operational indicators return to the target range. Under this strategy, the system utilizes a non-linear growth function to significantly increase the corresponding power factor improvement rate as the negative deviation increases. And reduce accordingly This forces the model to prioritize correcting the power factor during reward feedback until the metric returns to the target range. The reward function is calculated as follows: In this formula The parameter refers to the energy efficiency evaluation index (i.e., reward value) obtained by the energy efficiency decision model during reinforcement learning training. It represents the merits of the current control action in numerical form and is directly used as the reward value during reinforcement learning training.

[0031] Specifically, the system reads the power grid operation standards. When the monitored bus voltage deviation rate exceeds a preset threshold, such as exceeding ±7%, a boundary penalty term is triggered, imposing a negative penalty on the energy efficiency decision model, for example, assigning a negative score of -100, forcing the algorithm to avoid exceeding the limit. Simultaneously, to prevent transformer tap changers and reactive power compensation switching devices from losing mechanical life due to frequent operation, the system also includes an action cost penalty mechanism: when the control command output by the model causes a change in the transformer tap position or capacitor bank state, the system introduces a negative cost related to the action type and frequency into the reward value of the current time step. This negative cost value will be multiplied by a preset action penalty coefficient. An additional penalty term is applied to the model. The purpose is to quantify the mechanical wear cost of the equipment as negative feedback to the energy efficiency decision-making model, guiding it to learn a strategy that minimizes the frequency of control actions while meeting energy efficiency targets. Simultaneously, the energy efficiency evaluation index is constructed using a tiered penalty mechanism. The reduction in line loss is configured as the main reward term, reflecting operational economic benefits; the power factor is configured as a soft constraint term, imposing a non-linear deduction penalty when it falls below the assessment standard, and not imposing an additional reward when it exceeds the standard, thereby avoiding optimization redundancy caused by physical quantity coupling. Furthermore, the system uses a Min-Max normalization method based on historical extreme values ​​to map the reward value to a fixed interval, eliminating the interference of the moving target effect on Q-value estimation.

[0032] As a preferred implementation, the system introduces a dynamic normalization mechanism when constructing the reward function. This mechanism maintains the moving average and variance of the line loss reduction rate and power factor improvement rate in real time, and uses the Z-Score normalization method to map the two indicators with different dimensions to the same numerical range, followed by weighted summation. This technical solution solves the gradient update bias problem caused by the difference in the magnitude of the indicators in multi-objective optimization, improves the stability of model training convergence, and ensures the balance of the two optimization objectives in weight allocation.

[0033] This invention employs containerized microservices to deploy the energy efficiency decision-making model, providing an independent operating environment and isolated computing resources, thus avoiding resource contention during concurrent application deployments. Simultaneously, by constructing a reward function that includes boundary penalty terms and target reward terms, it achieves a numerical representation of physical safety constraints and energy efficiency optimization goals. The boundary penalty term ensures that the model learns voltage safety boundaries during training, preventing outputs from exceeding limits; the target reward term, through weighted calculation, clarifies the optimization direction, ensuring that the system performs line loss reduction and power factor improvement while meeting safety constraints.

[0034] Furthermore, the specific method for selecting a set of data as the target control parameters is based on the construction of a discretized action space and the inference mechanism of a deep neural network. First, during the initialization phase, the system defines the operating boundaries of the physical equipment and reads the adjustment range of the mechanical tap changer of the on-load tap changer, for example, an integer sequence from -8 to +8, totaling 17 tap positions; simultaneously, it reads the number of physical groups of the reactive power compensation switching device, for example, an integer sequence from 0 to 4, totaling 5 switching states. The system discretizes the above physical boundaries based on the minimum adjustable accuracy as the preset sampling step size, generating a set of tap position parameters and a set of switching levels. Subsequently, the system performs a Cartesian product operation, performing full permutations and combinations of the set of tap position parameters and the set of switching levels to construct a two-dimensional matrix containing all physically feasible combinations. This matrix is ​​defined as the preset equipment control parameter space, i.e., the candidate action library, and its dimension is determined by the cardinality product of the two sets.

[0035] Specifically, during the real-time inference phase, the system collects current operating electrical parameters and equipment status data through sensors. The state vector is constructed as follows: the system selects the effective values ​​of the three-phase voltage, active power, and reactive power at the current and past historical time points as time-series features, and the current tap position code and the number of capacitor switching groups as discrete features. All the above features are concatenated to form a one-dimensional state feature vector. The data processing module performs max-min normalization on the above data to eliminate dimensional differences. This state feature vector is input into a pre-trained energy efficiency decision model, which adopts a deep Q-network architecture. The deep neural network layer inside the model performs feature extraction and nonlinear mapping on the input vector, outputting a value score vector with the same dimension as the candidate action library. Each element in this vector represents the expected cumulative energy efficiency return that can be obtained by performing the corresponding action in the current state. The system applies a preset action selection strategy, such as a greedy strategy or a maximization strategy, to lock the action index corresponding to the maximum value score from the value score vector, and maps this index back to the specific combination in the candidate action library, thereby determining the target control parameters including reactive power compensation device switching commands and transformer tap position commands.

[0036] Specifically, the deep Q-network architecture adopts an improved Dueling DQN architecture. It includes an input layer, a shared feature extraction layer, a value function branch, and a dominance function branch. The number of neurons in the input layer is consistent with the dimension of the state feature vector; the shared feature extraction layer consists of three fully connected layers, each containing 128 neurons, with ReLU activation functions used between layers; the value function branch outputs the value score of the current state, and the dominance function branch outputs the dominance value of each action relative to the average action. The outputs of the two branches are aggregated in the output layer to calculate the Q-value of each action. The number of nodes in the output layer is consistent with the cardinality of the candidate action library, and it is used to output the Q-value of each action.

[0037] This invention constructs a discrete state set with the smallest adjustment action unit as the step size and generates a candidate action library using Cartesian product operations, thus defining the physical boundaries and completeness of the algorithm's search space. Compared to continuous space search, the pre-constructed finite discrete action library effectively reduces the computational complexity of the edge-side inference process and, through a mechanism, excludes the output of physically infeasible solutions, ensuring that the control commands meet the adjustment step size requirements of the hardware device.

[0038] Furthermore, the specific method for selecting a set of data as the target control parameters includes: obtaining the allowable adjustment range of the on-load tap changer associated with the ring main unit and the number of available switching groups of the reactive power compensation switching device; constructing their respective discrete state sets with the minimum adjustment action unit as the step size; performing Cartesian product operation on each of the discrete state sets to generate a candidate action library containing all feasible combinations as the equipment control parameter space; constructing the current state data packet into a state vector and inputting it into the energy efficiency decision model based on deep reinforcement learning; evaluating the value score of each set of parameters in the candidate action library through the energy efficiency decision model, and outputting the target control parameters according to the preset action selection strategy, wherein the target control parameters include reactive power compensation switching commands and transformer tap position commands.

[0039] Specifically, during the initialization phase, the system first obtains the allowable adjustment range of the on-load tap changer, such as -8 to +8, and the number of available switching groups for the reactive power compensation switching device, such as 0 to 4 groups. The system constructs a discrete state set for each device using its smallest adjustment unit as the step size. The Cartesian product operation is used to construct a complete candidate action library, and its calculation formula is expressed as follows: .in, This represents the discrete set of states of the voltage regulator switch. Represents the discrete state set of the reactive power compensation device, ordered pair This represents a set of joint control commands. The system performs the Cartesian product operation on the voltage regulator switch state set and the reactive power compensation switching set to generate a candidate action library containing all physically feasible combinations. This action library constitutes the equipment control parameter space.

[0040] Specifically, during the real-time inference phase, the system constructs a normalized state vector from the current state data packet and inputs it into a deep Q-network architecture. The model outputs a value score vector with dimensions consistent with the candidate action library, representing the expected reward for each action. Based on a preset action selection strategy, such as a greedy strategy, the system locks the index with the largest value from the score vector and maps this index back to a specific combination in the candidate action library, thereby determining the target control parameters that include reactive power compensation switching commands and transformer tap position commands.

[0041] As a preferred implementation, the offline retraining process employs a priority experience replay mechanism. The system calculates sample priority based on the temporal difference error in historical interaction data, assigning higher sampling probabilities to samples with high errors. Furthermore, before generating the incremental update package, the cloud performs INT8 post-training quantization on the new model. This technical solution improves model iteration efficiency by focusing on high-error samples, while simultaneously compressing the model size through quantization, reducing the transmission load of the cloud-edge collaborative channel and the storage footprint on the edge side.

[0042] This invention employs the Protobuf binary serialization format combined with the MQTT message transmission protocol for data transmission, reducing the data payload size and bandwidth consumption of network communication, and adapting to the unstable network environment at the edge of the distribution network. The cloud parses the synchronized payload and stores it in a power time-series database, establishing a structured historical operating condition data index. This provides a data foundation for subsequent model training and state backtracking, supporting efficient writing and querying of massive amounts of time-series data.

[0043] Furthermore, the specific method of formatting into a synchronous payload includes: extracting data from the status data packet and encoding it using the protobuf serialization data format to generate the synchronous payload; the specific method of offloading to the energy IoT cloud platform includes: establishing a message subscription and publishing channel with the energy IoT cloud platform based on the MQTT message transmission protocol, and pushing the synchronous payload through the channel.

[0044] Specifically, the data synchronization module performs efficient data serialization and transmission. The system defines a data structure according to the Protocol Buffers standard, extracts data from the status data packet, encodes it using this format, and generates a compact binary byte stream as the synchronization payload to reduce data size.

[0045] Specifically, the module constructs a lightweight communication link based on the MQTT messaging protocol, acting as a client to connect to the cloud platform and establish a message subscription and publishing channel. For uplink data, the module encapsulates the binary synchronization payload into an MQTT publish message and pushes it to the cloud platform through the aforementioned channel; simultaneously, it subscribes to downlink control topics to receive reverse control or model update instructions from the cloud.

[0046] This invention aggregates synchronous payloads by subscribing to MQTT message transmission protocol channels in the cloud and uses a power time series database for persistent storage, thus constructing a structured data foundation indexed by distribution terminal identifiers. This storage method is adapted to the writing characteristics of high-frequency sampled data in the power system, supports efficient retrieval and backtracking of massive historical operating conditions, and provides a complete data foundation for the data-driven training of subsequent energy efficiency decision models.

[0047] Furthermore, the specific method for updating the energy efficiency optimization service package includes: monitoring the data aggregation scale of the power time series database, and when the data aggregation scale meets the preset retraining threshold, extracting historical operational interaction data containing state-action pairs, and performing offline retraining on the energy efficiency decision model based on deep reinforcement learning; the historical operational interaction data includes the corresponding operational electrical parameters, ring network equipment status, historically executed target control parameters, and calculated energy efficiency evaluation indicators on the same time segment; the specific method for distributing to the edge computing unit includes: verifying the version number of the energy efficiency decision model after offline retraining, and distributing the incremental update package of the energy efficiency decision model based on deep reinforcement learning to the edge computing unit through a cloud-edge collaborative container orchestration architecture.

[0048] Specifically, a data-driven model self-evolution closed loop is constructed in the cloud. The system monitors the data aggregation scale of the power time-series database, triggering iteration when the number of new samples meets a preset retraining threshold, such as 100,000 samples. The system extracts historical operational interaction data containing state-action pairs, which is formatted into experience replay tuples, explicitly including operating electrical parameters, equipment status, historical actions, and energy efficiency evaluation indicators. The cloud uses this data to perform offline retraining on the deep Q-network architecture to update the weights. After training, the system verifies the model version number. If the version number has been updated, the cloud first locks the latest model weight parameters generated in this training and reads the old model weight parameters from the previous version sent to the edge. Using a parameter differencing strategy, element-wise subtraction is performed on the weight values ​​at corresponding network layer positions in the new and old versions of the model to obtain the weight change. The weight change is encapsulated into an incremental update package and sent to the edge computing unit through a cloud-edge collaborative container orchestration architecture, such as the KubeEdge architecture. After receiving the data, the edge computing unit parses the weight changes. The energy efficiency optimization service container within the edge computing unit has a pre-installed model hot update agent process. This agent process monitors the local update directory. When it receives the weight incremental update package (.diff file) from the cloud, it reads the weight matrix of the currently running model, performs matrix addition, and reloads the updated weights into the inference engine memory, thus completing the online synchronization of model parameters without restarting the container.

[0049] As a preferred implementation, the formatting and push process incorporates a dead-zone compression strategy based on the rotating door algorithm. The edge computing unit maintains a dynamic dead-zone threshold locally and calculates the rate of change of operating electrical parameters in real time; Protobuf serialization and MQTT push are triggered only when the data change exceeds this dead-zone threshold; otherwise, the data is cached locally. This technical solution achieves an on-demand transmission mode, preserving key transient characteristics while eliminating steady-state redundant data, saving communication traffic and cloud database storage space.

[0050] This invention utilizes historical interaction data containing state-action pairs to perform offline retraining, constructing a data-driven model iteration closed loop that enables the algorithm to adapt to the time-varying characteristics of distribution network loads. Simultaneously, by distributing incremental update packages to the model through a cloud-edge collaborative architecture, the amount of data transmission during model deployment is reduced, achieving version updates and maintenance of the edge-side inference model with minimal network resource consumption.

[0051] Furthermore, the specific method of encapsulating the remote device control commands includes: converting the reactive power compensation device switching commands and transformer tap position commands into command frames conforming to the Modbus-TCP protocol format; the specific method of sending the commands to the controlled device includes: sending the command frames to the programmable logic controller in the ring network box through the industrial Ethernet interface to remotely control the terminal device via the network.

[0052] Specifically, the real-time control module adopts a hierarchical asynchronous control strategy. The edge computing unit, based on short-term load forecasting, issues optimization commands one control cycle in advance to offset the accumulated latency of network transmission and equipment operation. Simultaneously, millisecond-level safety interlocking logic is embedded within the programmable logic controller (PLC). When receiving commands from the edge side, the controller first verifies the instantaneous value of the current bus voltage. If the action might cause a momentary over-limit, it refuses to execute and triggers local protection, thereby ensuring the physical safety of the system under non-real-time communication links. The real-time control module performs protocol conversion and physical drive according to a preset mapping table. The real-time control module converts the target control parameters into command frames conforming to the Modbus-TCP protocol format, which contain the target register address and function code. Then, the module sends the command frames to the PLC in the ring network enclosure via an industrial Ethernet interface. After parsing the commands, the PLC drives the contactor to engage or the servo motor to rotate, thereby physically adjusting the capacitor bank and transformer taps, achieving remote network control and energy efficiency closed-loop optimization of the terminal equipment.

[0053] This invention converts target control parameters into instruction frames conforming to the Modbus-TCP protocol standard, achieving interoperability between algorithm decision results and industrial control equipment. By driving a programmable logic controller through an industrial Ethernet interface, it ensures that the digital control strategy can be correctly parsed and executed by existing underlying hardware, enabling remote network control of reactive power compensation devices and transformer tap changers.

[0054] Example 2 This embodiment demonstrates the application method of the cloud-edge collaborative primary and secondary integrated ring network box energy efficiency optimization system provided by the present invention in the distribution network of a high-tech industrial park with distributed photovoltaic access; see reference Figure 2 The specific steps are as follows: Furthermore, the system accesses the photovoltaic grid connection point and the low-voltage busbar of the ring network box in real time via the smart meter interface, and reads real-time data with an average sampling period of 100 milliseconds. The real-time data is divided into two categories: one category consists of operating electrical parameters characterizing the reverse power flow and load characteristics of the photovoltaic system, specifically including four-quadrant active and reactive power that reflect bidirectional energy flow, the three-phase voltage amplitude at the grid connection point, and the total harmonic distortion rate of the current introduced by the photovoltaic inverter; the other category consists of ring network box equipment status data characterizing the physical operating conditions, specifically including the current tap position of the on-load tap changer, the switching code of the reactive power compensation capacitor bank, and the temperature and humidity environmental values ​​inside the cabinet. Using a pre-set energy internet protocol template, the above multi-source heterogeneous data is mapped into a unified key-value pair structure and encapsulated into a JSON format data frame to form a status data packet containing photovoltaic output characteristics. Subsequently, the system reads the unique identifier of the high-tech zone distribution terminal stored locally and generates an interface call instruction containing the identifier and the status data packet. This instruction explicitly points to the "source-grid-load-storage collaborative optimization" business interface in the local containerized microservice and triggers subsequent processes as a service call request.

[0055] Furthermore, a deep reinforcement learning energy efficiency decision-making model based on physical information constraints, deployed within containerized microservices, is employed. Addressing the characteristics of strong randomness in photovoltaic output and high risk of voltage exceeding limits, the model uses bus voltage deviation rate, instantaneous photovoltaic penetration rate, and line load rate as state inputs, and reactive power compensation device switching combinations and transformer tap change commands as action outputs. Logically, a reward function is constructed to represent constraints and objectives: the allowable voltage deviation range specified by the national power grid standard is configured as a boundary penalty term in the reward function; when photovoltaic backfeeding during high-sunlight periods causes the predicted voltage to potentially exceed this limit, a high negative penalty is applied to the model, forcing the strategy to prioritize adjusting transformer taps or disconnecting capacitors. Additionally, energy efficiency evaluation indicators are configured as the target reward term in the reward function, their values ​​obtained by calculating the weighted sum of line loss reduction and power factor improvement per unit of power supply. The weighting coefficients are set to reduce line losses caused by backfeeding while also considering power factor assessment. The model also integrates a physical information neural network architecture, constructing a composite loss function that includes a reinforcement learning main loss term and a physical constraint regularization term. The main loss term is calculated based on the temporal difference error between the value score predicted by the deep Q-network and the target value score. The physical constraint regularization term is determined by measuring the deviation between the voltage prediction implicit in the network output and the theoretical voltage value derived from Kirchhoff's voltage law. Both terms are weighted and summed by introducing a pre-defined penalty coefficient, thus embedding the physical conservation law into the gradient descent process. This allows the neural network to automatically converge its output trajectory to a feasible region that conforms to the principles of circuit physics while optimizing energy efficiency strategies.

[0056] Furthermore, during the initialization phase, the adjustable physical boundaries of the ring main unit are defined, the adjustment range of the on-load tap changer and the number of capacitor groups of the reactive power compensation switching device are read, and their respective discrete state sets are constructed with the minimum adjustment action unit as the step size. A candidate action library containing all physically feasible combinations, i.e., the equipment control parameter space, is constructed through Cartesian product operation. During the real-time inference phase, the current grid operation parameters with a high proportion of photovoltaic access are collected, and after normalization, a state vector is constructed and input into the energy efficiency decision model. The value score (Q value) of each group of parameters in the candidate action library is calculated through a deep Q-network architecture. This score represents the expected benefit of performing the action at the current photovoltaic output level in terms of smoothing voltage fluctuations and reducing line losses. A greedy strategy or a maximization strategy is applied to lock the data with the largest score and determine it as the target control parameter. For example, during periods of high photovoltaic power generation, the transformer tap position is lowered and some capacitors are disconnected to preemptively offset the voltage rise effect.

[0057] Furthermore, based on the Protocol Buffers standard, a description file was written to define the structural schema of the power data in the high-tech park. Floating-point data such as bidirectional power and voltage harmonics, as well as enumerated data such as switch states, were assigned unique tag numbers. Status data packets were mapped and compressed into a synchronous payload in binary byte stream form. During the data offloading phase, a communication link with the energy IoT cloud platform was established based on the MQTT message transmission protocol. The system, acting as a client, published data and adopted a dead-zone compression transmission strategy based on the rotating door algorithm: the system maintains a dynamic dead-zone threshold associated with photovoltaic volatility. Serialization and message push are only triggered when sudden changes in illumination (such as cloud cover) cause the rate of change of the collected data to exceed this threshold; otherwise, stable data is cached locally, and the cloud platform subsequently reconstructs the full curve through interpolation. This mechanism captures the transient characteristics of photovoltaics while significantly reducing the network overhead caused by continuous uploading.

[0058] Furthermore, a cloud service orchestration module is deployed based on the KubeEdge open-source edge computing framework. The cloud serves as the control plane, and the edge-side ring network terminal acts as the computing node. The energy efficiency optimization service package is packaged into a Docker container image for unified orchestration. In the data aggregation stage, the cloud subscribes to the upload channels of all campus terminals via MQTT wildcards, receives and parses the synchronous payload, extracts the terminal identifier and time-series data, and stores it in the power time-series database. This accumulates massive amounts of photovoltaic output curves and load characteristic data, providing a data foundation for analyzing the impact characteristics of photovoltaics on the power grid.

[0059] Furthermore, by monitoring the power time-series database, when newly added photovoltaic fluctuation data (such as new samples generated by seasonal changes in sunlight) reaches a preset retraining threshold, a data-driven self-evolutionary closed-loop mechanism is constructed to automatically trigger the model iteration pipeline, extract historical operational interaction data containing state-action pairs, and perform offline retraining on the energy efficiency decision model. The historical data explicitly includes operational electrical parameters, equipment status, historically executed target control parameters, and energy efficiency evaluation indicators at the same time point. The update process employs a knowledge distillation strategy based on deep Q-networks. In this embodiment, the core component of the DQN architecture—the deep neural network used to fit the action value function—is designed as an asymmetric teacher-student structure to address the problem of limited inference resources on the edge side. Specifically, a teacher model based on an improved deep residual network (1D-ResNet) is maintained in the cloud. This model modifies the first-layer convolutional kernel from a 7×7 two-dimensional convolution to a one-dimensional temporal convolution with a kernel size of 7 to adapt to power time-series data with an input tensor shape of (time step, feature dimension). The student model runs at the edge using a lightweight fully connected network (3 hidden layers). The teacher model is trained on global data in the cloud until convergence. A KL divergence term is then introduced into the loss function between the soft labels output by the teacher model and the predicted probability distributions output by the student model, and a distillation temperature coefficient is set. The parameters of the cloud-based student model replica are updated using gradient descent until a converged new student model is obtained. When generating the update package, the weight deviation values ​​between the new and old student models in the isomorphic network layers are extracted, and these deviation values ​​are encapsulated into an incremental update package. After verifying the version number using the KubeEdge framework, the module uses container image layering technology to distribute the incremental update package to the edge computing unit. This allows the edge terminal to obtain the latest voltage control strategy adapted to the current seasonal light characteristics through rolling updates without interrupting control.

[0060] Furthermore, based on a pre-set PLC address mapping table, the target control parameters output by the model, including reactive power compensation switching instructions and transformer tapping instructions, are encapsulated into Modbus-TCP instruction frames and sent to the programmable logic controller (PLC) in the ring network enclosure via an industrial Ethernet interface. After PLC parsing, the PLC drives contactor actuation or servo motor rotation via DO output, physically executing transformer tap adjustment or capacitor switching. This process achieves a closed loop from cloud-based AI decision-making to physical device execution, ensuring that the power grid is in optimal operating condition through preventative regulation before drastic changes in photovoltaic output cause voltage over-limits, thus completing the value assessment and execution of the discrete action space by the deep Q-network architecture.

[0061] This invention constructs a cloud-edge collaborative architecture that decouples training and inference by deploying edge computing units within intelligent power distribution terminals and combining them with cloud service orchestration modules to iteratively update energy efficiency optimization service packages. Compared to centralized control in the cloud, local inference based on local service packages effectively reduces the impact of wide area network communication latency and link fluctuations on control real-time performance. Furthermore, it overcomes the physical bottleneck of edge-side embedded devices being unable to handle complex model training tasks due to limited computing and storage resources, ensuring the adaptive updating and accurate execution of energy efficiency optimization strategies.

[0062] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A cloud-edge collaborative primary and secondary integrated ring network box energy efficiency optimization system, characterized in that, include: Edge device deployment and data acquisition module: Deploys edge computing units within the intelligent power distribution terminal; Real-time reception of power parameters and status data of ring network equipment in the distribution network is integrated into status data packets; and corresponding service call requests are generated based on the status data packets. Service Invocation Module: Based on the service invocation request, invokes the energy efficiency optimization service package embedded in the edge computing unit, and under the condition of meeting the preset grid voltage and load safety constraints, selects a set of data as target control parameters from the equipment control parameter space that characterizes the equipment operating status; Data synchronization module: acquires the status data packet and formats it into a synchronization payload, and offloads the synchronization payload to the energy IoT cloud platform via the MQTT message transmission protocol; Cloud service orchestration module: It collects the synchronous payloads, updates the energy efficiency optimization service package as feedback data, and distributes the updated energy efficiency optimization service package to the edge computing unit through the network according to the preset strategy; Real-time control module: Encapsulates the target control parameters into remote device control commands and sends the remote device control commands to the controlled device in real time.

2. The energy efficiency optimization system for a primary and secondary integrated ring network enclosure based on cloud-edge collaboration according to claim 1, characterized in that, The edge computing unit is an intelligent computing terminal integrating a central processing unit, memory, and network communication interface. It internally deploys a container runtime environment, providing isolated computing resources and a running platform for the energy efficiency optimization service package. The specific method of integrating these into a status data packet includes: acquiring power distribution network operating parameters including three-phase voltage, three-phase current, active power, and reactive power, as well as ring network equipment status data including switch on / off positions, energy storage status, and ambient temperature and humidity; encapsulating the operating parameters and the ring network equipment status data into a data frame of a structured data object; and using this data frame as the status data packet. The specific method of generating the corresponding service call request includes: generating an interface call instruction containing a power distribution terminal identifier and the status data packet; and using this interface call instruction as the service call request.

3. The energy efficiency optimization system for a primary and secondary integrated ring network enclosure based on cloud-edge collaboration according to claim 1, characterized in that, The energy efficiency optimization service package is an independent running instance encapsulated as a containerized microservice. It integrates a deep reinforcement learning-based energy efficiency decision-making model and an inference environment to support its operation. The energy efficiency decision-making model employs a deep Q-network architecture and constructs a reward function to characterize the energy efficiency evaluation index and the grid voltage and load safety constraints. The grid voltage and load safety constraints are configured as boundary penalty terms in the reward function; when the monitored bus voltage deviation rate exceeds a preset threshold, a negative penalty is applied to the energy efficiency decision-making model. The energy efficiency evaluation index is configured as the target reward term in the reward function, specifically a weighted sum of the line loss reduction rate and power factor improvement rate within a unit power supply cycle.

4. The energy efficiency optimization system for a complete ring network enclosure based on cloud-edge collaboration as described in claim 1, characterized in that, The specific method for selecting a set of data as target control parameters includes: obtaining the allowable adjustment range of the on-load tap changer associated with the ring main unit and the number of available switching groups of the reactive power compensation switching device; constructing their respective discrete state sets with the minimum adjustment action unit as the step size; performing Cartesian product operation on each discrete state set to generate a candidate action library containing all feasible combinations as the equipment control parameter space; constructing the current state data packet into a state vector and inputting it into an energy efficiency decision model based on deep reinforcement learning; evaluating the value score of each set of parameters in the candidate action library through the energy efficiency decision model, and outputting the target control parameters according to a preset action selection strategy. The target control parameters include reactive power compensation switching commands and transformer tap position commands.

5. The energy efficiency optimization system for a primary and secondary integrated ring network enclosure based on cloud-edge collaboration according to claim 1, characterized in that, The specific method of formatting the synchronous payload includes: extracting data from the status data packet and encoding it using the protobuf serialization data format to generate the synchronous payload; the specific method of offloading to the energy IoT cloud platform includes: establishing a message subscription and publishing channel with the energy IoT cloud platform based on the MQTT message transmission protocol, and pushing the synchronous payload through the channel.

6. The energy efficiency optimization system for a primary and secondary integrated ring network enclosure based on cloud-edge collaboration according to claim 1, characterized in that, The specific method for aggregating the synchronization payload includes: subscribing to the MQTT message transmission protocol channel on the energy Internet of Things cloud platform, receiving and parsing the synchronization payload, and storing the parsed data into the power time series database according to the power distribution terminal identifier.

7. The energy efficiency optimization system for a primary and secondary integrated ring network enclosure based on cloud-edge collaboration according to claim 1, characterized in that, The specific method for updating the energy efficiency optimization service package includes: monitoring the data aggregation scale of the power time series database, and when the data aggregation scale meets the preset retraining threshold, extracting historical operation interaction data containing state-action pairs, and performing offline retraining on the energy efficiency decision model based on deep reinforcement learning; the historical operation interaction data includes the corresponding operating electrical parameters, ring network box equipment status, historically executed target control parameters, and calculated energy efficiency evaluation indicators on the same time segment; the specific method for distributing to the edge computing unit includes: verifying the version number of the energy efficiency decision model after offline retraining, and distributing the incremental update package of the energy efficiency decision model based on deep reinforcement learning to the edge computing unit through a cloud-edge collaborative container orchestration architecture.

8. The energy efficiency optimization system for a primary and secondary integrated ring network enclosure based on cloud-edge collaboration according to claim 1, characterized in that, The specific method of encapsulating remote device control commands includes: converting reactive power compensation device switching commands and transformer tap position commands into command frames conforming to the Modbus-TCP protocol format; the specific method of sending the commands to the controlled device includes: sending the command frames to the programmable logic controller in the ring network box through an industrial Ethernet interface to remotely control the terminal device via the network.

Citation Information

Patent Citations

  • Optimal energy consumption task unloading method, device and system based on cloud edge collaboration

    CN115051999A

  • Self-adaptive power grid state transformer area intelligent fusion terminal and control method thereof

    CN119275991A

  • Station area intelligent fusion terminal data processing system based on edge calculation

    CN119440800A

  • Intelligent management and control method for flexible power distribution network based on cloud side-end cooperation

    CN119921466A

  • Transformer energy efficiency optimization cloud platform based on edge computing

    CN121116575A

Cited By

  • Operation and maintenance data processing method and system based on cloud network fusion technology

    CN122120150A

  • Distributed new energy power generation regulation capability identification method and system based on cloud side end

    CN122196592A