Protocol Conversion and Self-Learning Method and System for Photovoltaic-Storage-Charging Collaboration in Power Grid Areas

By constructing a digital twin protocol agent and a protocol agent in the photovoltaic storage and charging system of the distribution area and conducting reinforcement learning collaborative training, dynamic optimization strategies are generated, which solves the problem of disconnect between protocol conversion and protocol control in the existing technology, realizes the dynamic adaptation and cross-domain collaborative optimization of the system, and improves communication efficiency and control effect.

CN121584881BActive Publication Date: 2026-04-21SICHUAN SIJI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN SIJI TECHNOLOGY CO LTD
Filing Date
2026-01-23
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing photovoltaic-storage-charging systems, there is a disconnect between protocol conversion and protocol control, which cannot dynamically adapt to the system's operating status, resulting in low communication efficiency and reduced control effectiveness. The lack of cross-domain collaboration mechanisms also affects the overall system performance.

Method used

A digital twin of the photovoltaic-storage-charging system in the distribution area is constructed. Through collaborative training of protocol agents and specification agents in a reinforcement learning framework, dynamic optimization strategies are generated to achieve protocol conversion and specification self-learning. By combining weighted indicators such as communication latency, packet loss rate, system network loss, and voltage deviation, collaborative optimization strategies are generated and deployed to edge control devices.

Benefits of technology

It achieves dynamic adaptation of protocol conversion and specification control, improves the system's adaptability and cross-domain collaborative optimization, ensures timely transmission of key control commands, enhances the system's robustness and reliability, can cope with photovoltaic power output fluctuations and load changes, and improves the operational efficiency of the photovoltaic-storage-charging system in the distribution area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121584881B_ABST
    Figure CN121584881B_ABST
Patent Text Reader

Abstract

This invention discloses a protocol conversion and self-learning method and system for photovoltaic-storage-charging coordination in power distribution areas. The method constructs a digital twin of the power distribution area's photovoltaic-storage-charging system, running a protocol agent and a protocol agent in parallel within a virtual environment. A hierarchical reinforcement learning architecture is used to achieve collaborative training between the two. The protocol agent learns dynamic priority scheduling and compression strategies for heterogeneous protocol messages such as Modbus, CAN, and IEC 104, while the protocol agent learns power balance and voltage stability control strategies based on photovoltaic output, energy storage SOC, and charging load. The two agents achieve cross-domain collaborative optimization through a reward function coupling mechanism. Finally, the trained collaborative strategies are securely deployed to the physical power distribution area edge control equipment. This invention solves the problems of disconnect between protocol conversion and collaborative control, rigid protocol strategies, and insufficient cross-domain collaboration in existing technologies, improving the operating efficiency and adaptability of the power distribution area's photovoltaic-storage-charging system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart grid technology, specifically to a protocol conversion and self-learning method and system for photovoltaic-storage-charging coordination in distribution substations. Background Technology

[0002] With the acceleration of the energy transition, the energy structure of distribution substations is undergoing profound changes. The large-scale integration of distributed photovoltaic, energy storage systems and electric vehicle charging facilities has transformed traditional distribution substations from a single power supply role into a complex energy system integrating power generation, energy storage and power consumption. While this structural change brings flexibility to system operation, it also poses new challenges to the coordinated control of substations.

[0003] The current state of technology in the integration of photovoltaic, energy storage, and charging systems in power distribution areas mainly includes the following:

[0004] In terms of system architecture, existing technologies mostly adopt hierarchical control or centralized control architecture. Some industrial control systems achieve equipment status monitoring and simulation by building digital twins and achieve optimized scheduling of production processes by adopting hierarchical control architecture. Although these solutions provide a basic framework for system integration, they fail to fully consider the unique communication protocol differences and protocol adaptation requirements of the distribution area.

[0005] In terms of communication protocol conversion, existing technologies typically rely on fixed protocol conversion gateways. These gateways use predefined mapping rules to convert between heterogeneous communication protocols such as Modbus, CAN, and IEC 104. However, this conversion mechanism lacks dynamic adaptability and cannot adjust communication strategies according to the real-time operating status of the system, resulting in low communication efficiency when system operating conditions change.

[0006] In terms of control strategy optimization, some existing solutions adopt traditional optimization algorithms or basic reinforcement learning methods. These methods often treat communication network optimization and power system control as two independent optimization problems, lacking an effective cross-domain coordination mechanism. When the system faces sudden load fluctuations or communication congestion, this fragmented optimization approach cannot guarantee the overall operating performance of the system.

[0007] The existing technical solutions mainly have the following drawbacks:

[0008] There is a disconnect between protocol conversion and system collaborative control. Existing protocol conversion methods use static mapping rules, which cannot dynamically adjust communication strategies based on the real-time operating status of the distribution area. During peak electricity consumption periods, critical control commands may experience significant transmission delays due to communication congestion, directly impacting system control performance. Furthermore, fixed protocol conversion rules are ill-suited to adapting to changes in communication requirements resulting from equipment upgrades or expansions.

[0009] The control protocol lacks necessary adaptability. The control protocol parameters of the existing system are usually preset based on typical operating scenarios. When the system operating conditions change significantly, it cannot automatically adjust key parameters such as energy storage charging and discharging thresholds and power allocation weights. This rigid control strategy leads to a significant decrease in control performance when the system faces atypical operating conditions such as photovoltaic output fluctuations and sudden load changes.

[0010] Finally, the lack of a cross-domain collaboration mechanism is a significant shortcoming of existing technologies. The optimization objectives at the communication layer and the optimization objectives at the control layer are independent of each other, lacking a unified optimization framework. The protocol conversion process only focuses on communication quality indicators, while the protocol control process only considers electrical operation indicators. This separate optimization mode makes it difficult to achieve the optimal overall system performance. Especially when communication resources are limited, the reliability of control command transmission cannot be effectively guaranteed, directly affecting the stable operation of the system.

[0011] Therefore, there is an urgent need for a technical solution that can achieve protocol conversion and self-learning collaborative optimization of specifications in order to solve the above problems and improve the overall operational efficiency of the photovoltaic-storage-charging system in the distribution area. Summary of the Invention

[0012] In order to overcome the problems of the prior art, the present invention discloses a protocol conversion and self-learning method and system for photovoltaic-storage-charging coordination in power distribution areas.

[0013] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0014] A protocol conversion and self-learning method for photovoltaic-storage-charging coordination in power distribution areas includes the following steps:

[0015] A digital twin of the photovoltaic, energy storage and charging system in the distribution area is constructed. The digital twin is used to perform high-fidelity mapping of the dynamic operation data of the photovoltaic, energy storage and charging equipment, the message sequence of the multi-source heterogeneous communication network and the topology of the distribution area.

[0016] In a digital twin, protocol agents and specification agents operate in parallel;

[0017] Among them, the protocol agent is used to learn dynamic priority scheduling and message compression strategies by parsing heterogeneous protocol messages such as Modbus, CAN and IEC104, and the specification agent is used to learn power balance and voltage stability control strategies for the distribution area by analyzing photovoltaic power output fluctuations, energy storage SOC and charging pile demand power.

[0018] By using a reinforcement learning framework, the protocol agent and the specification agent can conduct collaborative training and decision-making. The weighted indicators of communication delay, packet loss rate, system network loss, and voltage deviation are used as joint optimization objectives to generate collaborative optimization strategies.

[0019] The collaborative optimization strategy is deployed to the edge control equipment of the physical distribution area. Based on the real-time collected operation status of the distribution area, the protocol command conversion and collaborative control of photovoltaic inverters, energy storage converters and charging piles are executed.

[0020] Preferably, the protocol agent and the specification agent collaborate using a hierarchical reinforcement learning architecture, specifically including:

[0021] The protocol-based intelligent agent and the specification-based intelligent agent share some neural network layers;

[0022] The observation state space of the protocol agent includes the protocol conversion delay, packet loss rate and channel utilization of the multi-source heterogeneous communication network, while the action space includes dynamic priority scheduling and compression instructions for protocol message processing.

[0023] The observation state space of the specification agent includes real-time photovoltaic power output, state of charge (SOC) of the energy storage system, and real-time power demand of the charging pile. The action space includes adjustment commands for energy storage charging and discharging thresholds, photovoltaic inverter output power, and charging pile power limits.

[0024] Preferably, the protocol agent is configured as follows:

[0025] With protocol conversion delay and packet loss rate as the core state inputs, the system optimizes the output to provide scheduling instructions for data packets of different priorities.

[0026] The specification agent is configured as follows:

[0027] With the energy storage SOC and the power demand of the charging pile as the core state inputs, the output is optimized through entropy increase to provide step-by-step adjustment instructions for the energy storage charging and discharging thresholds.

[0028] Preferably, the protocol agent and the specification agent coordinate through the mutual coupling of reward functions, specifically:

[0029] The instant reward value of the protocol agent is determined by the contribution weight of the real-time action to the reduction rate of line loss and the improvement of voltage qualification rate of the corresponding transformer area controlled by the protocol.

[0030] The instantaneous reward value of the protocol agent is determined by the contribution weight of the real-time action to the reduction rate of average transmission delay and the improvement of packet loss rate of key control commands corresponding to the protocol communication.

[0031] Preferably, the contribution weight is dynamically adjusted based on the optimization priority of communication quality and system energy efficiency under different operating periods.

[0032] Preferably, the collaborative optimization strategy is deployed to the edge control devices in the physical distribution area, specifically as follows:

[0033] The parameters of the trained and stable policy model are encapsulated into a policy update file and distributed from the digital twin to the edge control device through a secure channel;

[0034] After the edge control device loads the policy update file, it outputs protocol conversion and specification control commands in real time based on the current status.

[0035] Preferably, the size of the policy update file is limited to less than 2MB, and the interaction cycle between the digital twin and the edge control device of the physical station area is controlled to less than 100ms.

[0036] Preferably, extreme operating conditions, including communication failures and load surges, are simulated in the digital twin, and corresponding system recovery strategies are generated and integrated into the collaborative optimization strategy.

[0037] Preferably, the method further includes a strategy iteration step:

[0038] Using digital twins, and with preset communication quality index thresholds and system energy efficiency index thresholds as trigger conditions, the collaborative optimization strategy is automatically verified and updated.

[0039] Preferably, a protocol conversion and self-learning system for implementing the above method for coordinated photovoltaic-storage-charging systems in power distribution areas includes:

[0040] Digital twin module, used to build and operate a digital twin of the photovoltaic, energy storage and charging system in a distribution area;

[0041] The agent collaborative training module is used to run and train protocol agents and specification agents in a digital twin;

[0042] The strategy management and deployment module is used to manage collaborative optimization strategies and securely distribute and deploy them to edge control devices in physical areas.

[0043] The beneficial effects of this invention are as follows:

[0044] Compared with the prior art, the technical solution provided by the present invention has the following significant advantages:

[0045] It achieves dynamic adaptation between protocol conversion and collaborative control. Through the dynamic priority scheduling and message compression mechanism of protocol intelligence agents, it can automatically optimize communication strategies according to the system operating status, ensure the timely transmission of key control commands, and solve the problem of the disconnect between traditional static protocol conversion and dynamic control requirements.

[0046] By continuously learning from the operating data of the transformer substation, the regulatory agent can dynamically adjust the energy storage charging and discharging threshold and power allocation weight, thereby effectively coping with photovoltaic power output fluctuations and load changes and overcoming the shortcomings of traditional fixed regulatory strategies in terms of reduced control performance under atypical operating conditions.

[0047] Establish a cross-domain collaborative optimization mechanism. Through the mutual coupling design of reward functions, the protocol agent considers the impact on system energy efficiency when optimizing communication quality, and the regulation agent considers the impact on communication load when optimizing control strategy. This achieves deep collaboration between the communication and control domains and solves the problem of low system efficiency caused by traditional separate optimization.

[0048] To ensure the security and real-time performance of policy deployment, the policy file size is limited to less than 2MB, the interaction cycle is controlled to less than 100ms, and a secure transmission channel is established to ensure that collaborative policies can be deployed to physical stations securely, reliably, and in real time, effectively reducing system operation risks.

[0049] To enhance the robustness and reliability of the system, extreme conditions such as communication failures and load surges are simulated in a digital twin, and corresponding recovery strategies are generated, enabling the system to cope with emergencies and improving the operational reliability of the photovoltaic, energy storage and charging system in the distribution area. Attached Figure Description

[0050] Figure 1 This is a flowchart of the steps of the protocol conversion and self-learning method for photovoltaic-storage-charging coordination in a distribution area provided in Embodiment 1 of the present invention;

[0051] Figure 2 This is a structural block diagram of the protocol conversion and self-learning system for photovoltaic-storage-charging coordination in power distribution areas provided in Embodiment 1 of the present invention. Detailed Implementation

[0052] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description relating to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods and systems consistent with some aspects of this application as detailed in the appended claims.

[0053] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0054] The following detailed description of the specific implementation methods, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided in detail.

[0055] Example 1

[0056] Please refer to Figure 1 This embodiment provides a protocol conversion and self-learning method for photovoltaic-storage-charging coordination in power distribution areas, including the following steps:

[0057] A digital twin of the photovoltaic, energy storage and charging system in the distribution area is constructed. The digital twin is used to perform high-fidelity mapping of the dynamic operation data of the photovoltaic, energy storage and charging equipment, the message sequence of the multi-source heterogeneous communication network and the topology of the distribution area.

[0058] In the digital twin, the protocol intelligent agent and the specification intelligent agent run in parallel. The protocol intelligent agent is used to learn dynamic priority scheduling and message compression strategies by parsing heterogeneous protocol messages such as Modbus, CAN and IEC104. The specification intelligent agent is used to learn power balance and voltage stability control strategies for the distribution area by analyzing photovoltaic power output fluctuations, energy storage SOC and charging pile power demand.

[0059] By using a reinforcement learning framework, the protocol agent and the specification agent can conduct collaborative training and decision-making. The weighted indicators of communication delay, packet loss rate, system network loss, and voltage deviation are used as joint optimization objectives to generate collaborative optimization strategies.

[0060] The collaborative optimization strategy is deployed to the edge control equipment of the physical distribution area. Based on the real-time collected operation status of the distribution area, the protocol command conversion and collaborative control of photovoltaic inverters, energy storage converters and charging piles are executed.

[0061] Furthermore, the protocol agent and the specification agent collaborate using a hierarchical reinforcement learning architecture. In this embodiment, the hierarchical reinforcement learning architecture is built using a Python-based reinforcement learning framework, such as Ray or RLlib. This architecture includes a top-level manager and two parallel agents, namely the protocol agent and the specification agent.

[0062] The top-level manager is responsible for coordinating the training cycle, allocating computing resources, and managing data interaction between the two agents. While the two agents operate in parallel architecturally, they have a collaborative relationship in their decision-making logic, together forming a complete decision-making system, specifically including:

[0063] The protocol-based intelligent agent and the specification-based intelligent agent share some neural network layers;

[0064] To achieve knowledge sharing and collaborative training, this implementation method designs a shared feature extraction network for the two agents;

[0065] Specifically, a shared module containing two fully connected layers, each with 128 neurons, is constructed. The input of this module is the concatenation of the protocol and specification state vectors, and its output serves as the common input of the decision networks of the two agents. Through this structure, the protocol agent can implicitly perceive the power grid operation status when optimizing the communication strategy, while the specification agent can also consider the load of the communication network when formulating the control strategy, thus spontaneously generating cross-layer collaborative behavior.

[0066] The observation state space of the protocol agent includes the protocol conversion delay, packet loss rate and channel utilization of the multi-source heterogeneous communication network, while the action space includes dynamic priority scheduling and compression instructions for protocol message processing.

[0067] The observation state space of the protocol agent is a three-dimensional vector, and its specific structure and data sources are as follows:

[0068] Protocol conversion delay, calculated using timestamps, refers to the time elapsed from when the edge device sends a query command to when the target device, such as a photovoltaic inverter, fully receives and parses the response message; the unit is milliseconds.

[0069] Packet loss rate is calculated by comparing the total number of request packets sent with the total number of valid response packets received within a single sampling period, such as 1 minute, and then determining the missing rate.

[0070] Channel utilization is obtained by monitoring the underlying communication drivers, such as CAN controllers or Ethernet cards, and reflects the percentage of physical channel bandwidth occupied.

[0071] The action space of the protocol agent contains two types of executable instructions:

[0072] The dynamic priority scheduling instruction allows the agent to output a continuous value from 0 to 1. Through a predefined mapping table, this value is converted into a priority identifier for communication messages to different devices, such as photovoltaic data, energy storage status, and charging control. For example, an output value of 0.8 can be mapped to prioritize the processing of energy storage status messages.

[0073] Compression commands allow the agent to output discrete actions, choosing from no compression, moderate compression (e.g., using zlib level 3), or deep compression (e.g., using zlib level 6) to process outbound messages accordingly.

[0074] The observation state space of the specification agent includes real-time photovoltaic power output, state of charge (SOC) of the energy storage system, and real-time power demand of the charging pile; the action space includes adjustment instructions for energy storage charging and discharging thresholds, photovoltaic inverter output power, and charging pile power limits.

[0075] The observation state space of the regulatory agent is a three-dimensional vector, and its specific structure and data sources are as follows:

[0076] The real-time output of photovoltaic power is read directly from the communication register of the photovoltaic inverter, such as the Modbus holding register, and the unit is kilowatts;

[0077] The State of Charge (SOC) of the energy storage system is obtained by parsing the data frames reported by the Battery Management System (BMS) via the CAN bus protocol, and the unit is percentage.

[0078] The real-time power demand of the charging pile is obtained from the current power demand telemetry point transmitted by the charging pile controller through the IEC 104 protocol, and the unit is kilowatts.

[0079] The action space of a specification agent contains three types of executable instructions:

[0080] The energy storage charging and discharging threshold adjustment command outputs a new SOC action point, such as adjusting the charging upper limit of the energy storage system from 90% to 85%;

[0081] The photovoltaic inverter output power adjustment command outputs a limiting value, such as limiting the maximum output power of the photovoltaic inverter to 80% of its rated power;

[0082] The charging pile power limit adjustment command outputs a power value, such as reducing the instantaneous charging power of a certain charging pile from 7kW to 5kW.

[0083] Specifically, through the above design, the present invention achieves deep synergy between communication control and energy control under a unified framework, enabling the system to adaptively respond to dynamic changes in the distribution area, thereby optimizing the overall operational energy efficiency and stability of the distribution area while ensuring the quality of communication services.

[0084] Furthermore, the protocol agent is configured to take protocol conversion delay and packet loss rate as core state inputs, and to output scheduling instructions for data packets of different priorities through policy optimization.

[0085] Core state input refers to prioritizing protocol conversion delay and packet loss rate, the two most direct indicators affecting communication quality, as the main basis for decision-making among the three-dimensional state vector, namely protocol conversion delay, packet loss rate, and channel utilization.

[0086] The protocol agent optimizes the output of scheduling instructions for data packets of different priorities through policy optimization, as specifically implemented as follows:

[0087] The policy optimization algorithm adopts the proximal policy optimization algorithm, which ensures training stability through importance sampling and pruning mechanisms. The agent's policy network receives the core state input and processes it through a three-layer fully connected neural network with 128, 64 and 32 neurons respectively.

[0088] The scheduling instructions are generated, and the policy network outputs an 18-dimensional action probability distribution, corresponding to 18 different priority scheduling schemes. Each scheme defines different priority combinations of three types of data packets: Modbus protocol messages, CAN bus messages, and IEC 104 protocol messages, such as:

[0089] Action 0: Modbus high priority, CAN medium priority, IEC 104 low priority;

[0090] Action 1: CAN high priority, IEC 104 medium priority, Modbus low priority;

[0091] The remaining 16 actions follow the same pattern;

[0092] Upon execution of the instruction, the edge control device sets the corresponding Quality of Service (QoS) parameters in the communication stack according to the selected scheduling scheme to ensure that high-priority packets are transmitted first.

[0093] The protocol agent is configured to take the energy storage SOC and the power demand of the charging pile as the core state inputs, and optimize the output of the energy storage charging and discharging thresholds through entropy increase optimization.

[0094] Core state input refers to the two parameters that have the most decisive impact on the power balance of the distribution area: real-time photovoltaic output, energy storage SOC, and charging pile demand power in the three-dimensional state vector.

[0095] The regulatory agent optimizes the output of stepwise adjustment instructions for the energy storage charging and discharging thresholds through entropy increase, as specifically implemented as follows:

[0096] The entropy-increasing optimization algorithm employs the flexible actor-critic SAC algorithm. This algorithm encourages the agent to maintain a balance between exploration and exploitation by introducing a policy entropy term into the reward function. The formula for calculating policy entropy is:

[0097]

[0098] Where H represents policy entropy, which measures the degree of randomness of policy π in state s, and a represents each action in the action space;

[0099] A stepped adjustment command is generated, dividing the energy storage SOC range [20%, 95%] evenly into 5 stepped intervals. The control agent outputs discrete actions, selecting one from 12 preset stepped adjustment schemes. Each scheme includes:

[0100] The charging threshold can be adjusted, for example, from the current 90% to 85%, 80%, or 75%.

[0101] Discharge threshold adjustment, for example, from the current 30% to 35%, 40%, or 45%;

[0102] The adjustment range is fixed in multiples of 5%, forming a clear stepped characteristic;

[0103] This step-by-step adjustment avoids frequent small changes in charging and discharging thresholds, reduces operational losses on energy storage devices, and provides communication networks with more stable and predictable load characteristics.

[0104] Specifically, through the above-mentioned configuration, the two intelligent agents can make more accurate and stable decisions in their respective domains, while maintaining collaboration through the shared network layer to jointly improve the overall operational efficiency of the photovoltaic-storage-charging system in the distribution area.

[0105] Furthermore, the protocol agent and the protocol agent collaborate through the mutual coupling of reward functions. In the hierarchical reinforcement learning architecture, this is achieved by designing specific reward calculation logic. This coupling mechanism does not rely on additional hardware modules, but rather establishes a bidirectional quantitative relationship between protocol control actions and power grid operation indicators, and between protocol control actions and communication quality indicators, through software algorithms during agent training. Specifically:

[0106] The instant reward value of the protocol agent is determined by the contribution weight of the real-time action to the reduction rate of line loss and the improvement of voltage qualification rate of the corresponding transformer area controlled by the protocol.

[0107] The specific implementation of the instant reward value calculation for the protocol agent is as follows:

[0108] Calculation of the contribution weight of the reduction rate of line loss in the transformer area:

[0109] Line loss reduction rate = (Previous cycle area bus loss - Current cycle area bus loss) / Previous cycle area bus loss;

[0110] Among them, the line loss of the transformer area is calculated by the difference between the metering point at the beginning of the transformer area and the metering points of each branch. The contribution weight of the protocol intelligent agent to the reduction of line loss is determined by the degree of improvement of the timeliness of the transmission of key control messages by the scheduling instructions, specifically as follows:

[0111] Contribution weight = Timely deployment rate of critical control messages after protocol action execution - Timely deployment rate of critical control messages before protocol action execution;

[0112] Calculation of the contribution weight of voltage qualification rate improvement:

[0113] Voltage qualification rate = (Number of monitoring points with qualified voltage within the transformer area / Total number of monitoring points) × 100%;

[0114] Voltage pass rate improvement = Current cycle voltage pass rate - Previous cycle voltage pass rate;

[0115] The contribution weight of the protocol agent to the improvement of voltage qualification rate is determined by the degree of improvement in the reliability of voltage regulation command transmission.

[0116] The instantaneous reward value of the protocol agent is ultimately calculated using the following formula:

[0117]

[0118] Wherein, α and β are weighting coefficients, which are dynamically adjusted according to the current operating status of the transformer area;

[0119] The instant reward value of the protocol agent is determined by the contribution weight of the real-time action to the reduction rate of the average transmission delay corresponding to the protocol communication and the improvement of the packet loss rate of key control commands.

[0120] The specific implementation of calculating the instantaneous reward value for the regulatory agent is as follows:

[0121] Calculation of the contribution weight of the average transmission delay reduction rate:

[0122] Transmission delay reduction rate = (historical average transmission delay - current period transmission delay) / historical average transmission delay;

[0123] The contribution weight of the specification agent to the reduction of transmission delay is determined by the optimization effect of the control strategy on the communication network load, specifically by the proportion of the number of communication packets reduced after changing continuous fine-tuning to step-by-step adjustment.

[0124] Calculation of the contribution weight of the improvement in the packet loss rate of critical control commands:

[0125] Packet loss rate improvement = Packet loss rate of critical instructions in the previous cycle - Packet loss rate of critical instructions in the current cycle;

[0126] The contribution weight of the specification agent to the improvement of packet loss rate is determined by the effect of its power allocation strategy on the degree of network congestion.

[0127] The instantaneous reward value of the regulatory agent is ultimately calculated using the following formula:

[0128]

[0129] Wherein, γ and δ are weighting coefficients that are dynamically adjusted according to the communication network status;

[0130] The two intelligent agents achieve deep coupling through the above reward function. When optimizing communication quality, the protocol intelligent agent will receive additional rewards because its actions indirectly improve the power grid operation indicators.

[0131] When optimizing power grid control, the regulatory agent also receives positive incentives for reducing communication burden due to its actions. This coupling relationship established at the reward function level prompts two agents to actively consider the positive impact on the other domain while pursuing their own primary objectives, thereby achieving cross-layer collaborative optimization.

[0132] Furthermore, the calculation of contribution weights is dynamically adjusted based on the optimization priorities of communication quality and system energy efficiency under different operating periods;

[0133] This implementation divides the operating period of the transformer area into three typical time periods, each with different operating characteristics:

[0134] Peak electricity consumption periods (08:00-12:00, 18:00-22:00): Load is concentrated, voltage fluctuations are significant, and charging demand is strong;

[0135] Communication congestion period (14:00-16:00): Distributed photovoltaic data is reported centrally, and the communication channel load is high;

[0136] During normal operating hours (other times): load is stable, and communication traffic is normal;

[0137] Develop differentiated optimization priority strategies for different runtime scenarios:

[0138] During peak electricity consumption periods, energy efficiency optimization takes precedence over communication quality optimization. The system energy efficiency weighting coefficient is set to 0.7, and the communication quality weighting coefficient is set to 0.3.

[0139] During periods of communication congestion, communication quality optimization takes precedence over system energy efficiency optimization. The weighting coefficient for communication quality is set to 0.8, and the weighting coefficient for system energy efficiency is set to 0.2.

[0140] During normal operation, the weights of both are balanced, with the system energy efficiency weight coefficient set to 0.5 and the communication quality weight coefficient set to 0.5.

[0141] The dynamic adjustment of contribution weights is achieved through a time-period-aware weight calculation module, specifically including:

[0142] Dynamic adjustment of the contribution weights of protocol agents:

[0143] During peak electricity consumption periods, the contribution weight calculation of the protocol agent to protocol control is biased towards the improvement in voltage compliance rate, as shown in the formula:

[0144] Contribution weight = 0.8 × voltage qualification rate improvement + 0.2 × line loss reduction rate;

[0145] During periods of communication congestion, the contribution weight of the protocol agent to protocol control is calculated with an emphasis on the line loss reduction rate, as shown in the formula:

[0146] Contribution weight = 0.3 × voltage qualification rate improvement + 0.7 × line loss reduction rate;

[0147] Dynamic adjustment of the contribution weights of the regulatory agent:

[0148] During peak electricity consumption periods, the contribution weight calculation of the protocol intelligence agent to protocol communication is biased towards the improvement in the packet loss rate of key control commands, as shown in the formula:

[0149] Contribution weight = 0.2 × transmission delay reduction rate + 0.8 × packet loss rate improvement;

[0150] During periods of communication congestion, the contribution weight of the protocol agent to protocol communication is calculated with an emphasis on the average transmission delay reduction rate, as shown in the formula:

[0151] Contribution weight = 0.6 × transmission delay reduction rate + 0.4 × packet loss rate improvement;

[0152] The dynamic adjustment mechanism is implemented through the following steps:

[0153] The system clock module provides information about the current time period;

[0154] The operation status monitoring module collects load distribution and communication traffic data in real time;

[0155] The weight calculation module performs calculations based on time period characteristics and real-time data, according to a preset weight coefficient table.

[0156] The reward function recalculates the agent's immediate reward value based on the dynamically adjusted contribution weight;

[0157] Specifically, through a dynamic adjustment mechanism based on time-period characteristics, this invention can adapt to changes in the operating status of the distribution area, and achieve intelligent energy efficiency of the system while ensuring the quality of communication services. This improves the operational economy and adaptability of the optical-storage-charging collaborative system in the distribution area. This time-period-aware weight adjustment strategy enables the system to maintain optimal collaborative performance under different operating conditions.

[0158] Furthermore, the collaborative optimization strategy will be deployed to the edge control devices in the physical distribution area, specifically as follows:

[0159] The parameters of the trained and stable policy model are encapsulated into a policy update file and distributed from the digital twin to the edge control device through a secure channel;

[0160] The parameters of a stable policy model refer to the final weight parameters of the neural network model after the protocol agent and the specification agent have completed collaborative training in a digital twin. The encapsulation process specifically includes:

[0161] The weight matrix and bias vector are extracted from the policy network PPO algorithm of the protocol agent and the policy network of the regulation agent, namely the flexible actor-critic SAC algorithm.

[0162] Serialization processing uses the Google Protocol Buffers serialization tool to convert the weight parameters into binary format, ensuring the integrity of the data structure and platform independence;

[0163] The serialized parameter data and metadata are packaged together into a policy update file. The metadata includes:

[0164] Model version number, generation timestamp, file checksum, using MD5 algorithm, applicable station area identifier;

[0165] The establishment of the secure channel and the process of distributing documents are implemented as follows:

[0166] Employing a TLS 1.3-based encrypted communication protocol, a secure, two-way authenticated connection is established between the digital twin, deployed on cloud servers, and edge control devices such as the NVIDIA Jetson AGX Xavier.

[0167] Edge control devices verify the legitimacy of digital twins through digital certificates, and vice versa;

[0168] SFTP is used for secure transfer of policy update files, and the data is always encrypted during the transfer process.

[0169] After receiving a file, the edge control device recalculates the MD5 checksum and compares it with the checksum in the metadata to ensure that the file is not damaged or tampered with during transmission.

[0170] After the edge control device loads the policy update file, it outputs protocol conversion and specification control commands in real time according to the current status.

[0171] The specific process of loading policy update files for edge control devices includes:

[0172] Parse metadata to confirm model version compatibility and station identifier matching;

[0173] Using the same Protocol Buffers definition file, binary data is restored to neural network weight parameters;

[0174] The new weight parameters are loaded into the policy network of the protocol agent and the specification agent that are already running in the edge control device, so as to update the model without interrupting the existing service.

[0175] The strategy model from the previous version is retained, and if the new model encounters an anomaly during the validation period, it will automatically revert to the stable version.

[0176] After the edge control device loads the policy update file, the real-time decision-making process is implemented as follows:

[0177] Status acquisition: Continuously collect dynamic operating data of photovoltaic storage and charging equipment and message sequences of multi-source heterogeneous communication networks;

[0178] Real-time reasoning involves inputting the current state into the updated policy network, and forward propagation to calculate the optimal action.

[0179] The protocol agent outputs dynamic priority scheduling and message compression commands;

[0180] The protocol-defined intelligent agent outputs power balance and voltage stability control commands for the distribution area;

[0181] The generated control commands are sent to the specific devices through the corresponding communication interfaces:

[0182] Send power adjustment commands to the photovoltaic inverter via Modbus TCP;

[0183] Send charge and discharge threshold setting commands to the energy storage converter via CAN bus;

[0184] Send power limit instructions to charging piles via the IEC 104 protocol;

[0185] Specifically, through standardized file encapsulation formats, industrial-grade secure transmission protocols, reliable model update mechanisms, and efficient real-time inference processes, the collaborative optimization strategy can be implemented securely, reliably, and efficiently in physical distribution areas, achieving seamless integration and closed-loop optimization between the digital twin environment and the physical distribution area system.

[0186] Furthermore, the size of the policy update file is limited to less than 2MB, and the interaction cycle between the digital twin and the edge control device of the physical station area is controlled to less than 100ms.

[0187] The size limit for policy update files is achieved through the following technical means:

[0188] The neural network weights of the protocol agent and the specification agent are quantized from 32-bit floating-point numbers FP32 to 8-bit integers INT8, which reduces the model size by about 75% while maintaining an inference accuracy loss of less than 3%.

[0189] Parameter pruning and sharing: Parameter pruning is performed on shared neural network layers, removing weight connections whose absolute values ​​are less than the threshold 1e-4;

[0190] The unique decision networks of the protocol agent and the specification agent share the bias terms of the fully connected layer, and Huffman coding is used to further compress the sparse weight matrix.

[0191] When the policy update is small, only the weight change and delta update are transmitted, rather than the complete model parameters, keeping the update packet size within the range of 500KB-1.5MB.

[0192] Precise control of the interaction cycle is achieved through the following technical solutions:

[0193] The clock synchronization system, digital twin and edge control device adopt IEEE 1588 Precision Clock Synchronization Protocol PTP to ensure that the time deviation between the two is less than 1ms, providing an accurate time reference for cycle control;

[0194] Establish a dedicated fiber optic communication link between the digital twin and the edge control device, with a bandwidth of ≥100Mbps and a latency of ≤10ms;

[0195] Zero-copy technology is used to reduce the number of times data is copied in the protocol stack, thereby reducing transmission latency;

[0196] The policy update file is transmitted in chunks, with each chunk not exceeding 512KB in size, and resuming interrupted transmission is supported.

[0197] Assign the highest real-time task priority to the policy loading process in the edge control device, such as the SCHED_FIFO policy in the Linux system;

[0198] The hard time limit for the interaction cycle is set to 100ms, of which file transfer time ≤ 50ms, file verification and loading time ≤ 30ms, and model switching and initialization time ≤ 20ms.

[0199] File size limits and interaction cycle control work together through a unified resource management module:

[0200] The compression level is dynamically adjusted based on the current network conditions, and a more aggressive compression algorithm is used when the network is congested to ensure that the transmission is completed within the periodic time limit.

[0201] In the later stages of training, the digital twin transmits a simplified version of the strategy model, ≤1MB in size, to the edge device. During the formal update, only the differences need to be transmitted.

[0202] The system monitors the execution of the interaction cycle in real time. When the risk of timeout reaches the threshold, it automatically switches to degrade mode to prioritize the transmission of critical control commands.

[0203] Specifically, by using technologies such as model quantization, parameter sharing, and incremental updates, the file size is strictly limited to within 2MB. Combined with precise clock synchronization, transmission link optimization, and real-time scheduling, the interaction cycle is controlled within 100ms. These technical features together constitute an efficient and reliable strategy deployment system that not only meets the real-time requirements of the power grid environment but also adapts to the resource constraints of edge devices. This provides important performance guarantees for the strategy deployment process and ensures that the collaborative optimization strategy can take effect in the physical distribution area in a timely and stable manner.

[0204] Furthermore, extreme operating conditions, including communication failures and load surges, are simulated in the digital twin, and corresponding system recovery strategies are generated and integrated into the collaborative optimization strategy.

[0205] The simulation of extreme operating conditions is carried out in the constructed digital twin, and the specific implementation is as follows:

[0206] Fault types include simulating real communication faults such as CAN bus open circuit, Modbus TCP connection timeout, and IEC 104 message sequence number abnormality.

[0207] Through the fault injection engine of the digital twin, communication faults can be triggered randomly or in a preset pattern at specific time points, such as during peak electricity consumption periods.

[0208] Set the fault duration range to 1-30 seconds and the fault occurrence frequency to 1-5 times per hour;

[0209] The surge mode simulates scenarios such as the simultaneous activation of multiple charging piles in the area and the concentrated commissioning of large electrical equipment.

[0210] The load suddenly increases to 150%-300% of the normal value within 2-5 seconds, and lasts for 5-15 minutes;

[0211] Synchronous simulation of the chain reaction caused by this, such as the voltage drop in the transformer area to below 0.9 pu and line overload;

[0212] The system recovery strategy is generated based on a reinforcement learning agent in the digital twin, and the specific process is as follows:

[0213] State awareness and diagnosis: the protocol agent monitors the communication link status in real time, identifies fault types and their impact range; the protocol agent monitors changes in electrical parameters and assesses the stability of the system.

[0214] In the event of a communication failure, the protocol agent learns and generates a backup communication path switching strategy, such as automatically switching from the CAN bus to the RS485 backup link, or enabling a message retransmission mechanism.

[0215] In the event of a load surge, the specification agent learns and generates a tiered load control strategy, cuts off non-critical loads according to preset priorities, and coordinates the energy storage system to provide power support within 100ms.

[0216] Strategy verification and optimization involves repeatedly rehearsing recovery strategies in the secure environment of the digital twin. The reward function guides the agent to learn the optimal recovery path, ensuring the shortest recovery time and minimal impact.

[0217] The fusion of recovery strategies and collaborative optimization strategies is achieved through the following methods:

[0218] The optimal recovery action learned under extreme conditions is transferred to the policy network under normal conditions through policy distillation, enabling the agent to have predictive optimization capabilities.

[0219] Add a system resilience reward term to the joint optimization objective to encourage agents to consider the system's anti-interference capability during normal optimization:

[0220]

[0221] Wherein, λ1 and λ2 are weighting coefficients, communication redundancy refers to the availability of backup communication channels, and power reserve margin refers to the instantaneous power support capability of the energy storage system.

[0222] The underlying feature extraction layer of the recovery strategy network is shared with the normal optimization strategy network, enabling deep fusion of the two strategies at the feature level.

[0223] Specifically, by systematically simulating extreme conditions such as communication failures and load surges in a digital twin, and by leveraging reinforcement learning agents to autonomously generate recovery strategies, these strategies are ultimately integrated into a collaborative optimization system through knowledge distillation, reward reshaping, and parameter sharing. This enhances the resilience and reliability of the photovoltaic-storage-charging system in the face of actual operational risks. This proactive design philosophy ensures that the system not only performs excellently under normal operating conditions but also maintains safe and stable operation under extreme circumstances.

[0224] Furthermore, the method also includes a strategy iteration step, which uses a digital twin to automatically verify and update the collaborative optimization strategy using preset communication quality index thresholds and system energy efficiency index thresholds as trigger conditions.

[0225] The preset threshold settings specifically include:

[0226] Communication quality indicator thresholds: average transmission delay threshold ≤100ms, critical instruction packet loss rate threshold ≤2%, protocol conversion power threshold ≥99.5%;

[0227] System energy efficiency thresholds: Distribution area line loss rate threshold: ≤3.5%, Voltage qualification rate threshold: ≥98%, Photovoltaic absorption rate threshold: ≥95%;

[0228] The triggering condition is implemented through a real-time monitoring and threshold comparison module:

[0229] Single indicator exceeding the limit trigger: any communication quality indicator or system energy efficiency indicator exceeds the preset threshold for three consecutive sampling cycles;

[0230] The deterioration of composite indicators is triggered when two or more indicators simultaneously reach a critical state, reaching 90% of the threshold.

[0231] Periodic forced triggering, a forced verification is performed every 24 hours, regardless of whether the indicator exceeds the limit;

[0232] The automated verification process is implemented in the digital twin:

[0233] Scenario building for verification:

[0234] Import the actual operating data of the physical transformer area in the last 24 hours as the benchmark scenario, and generate a variety of test scenarios such as typical days, extreme weather, and equipment failures based on historical data;

[0235] The collaborative optimization strategy to be verified is deployed in parallel in the digital twin, all test scenarios are run, key indicators such as protocol conversion power and system energy efficiency are recorded, and the results are compared and analyzed with the original strategy to calculate the percentage performance improvement.

[0236] The automated process for policy updates specifically includes:

[0237] An update is triggered when the verification results show that the new strategy performs no worse than the original strategy in all test scenarios, and improves by more than 5% in at least one scenario;

[0238] Updated decisions must undergo security verification to ensure there are no systemic risks;

[0239] The update and deployment will be phased in, initially deploying the new policy on 10% of the edge control devices and observing the actual operational effects.

[0240] If no abnormalities are found within 24 hours, the deployment scope will be gradually expanded to 30%, 60%, and 100%. During the deployment process, the ability to run the old and new strategies in parallel will be maintained, and rapid rollback will be supported.

[0241] Update effect tracking: After the update is completed, continuously monitor the operation data of the physical transformer area for 7 days, and feed back the actual operation effect to the digital twin for subsequent strategy optimization;

[0242] Data collection: Operational data from the physical distribution area is continuously uploaded to the digital twin.

[0243] Model training involves retraining the protocol agent and the specification agent based on new data.

[0244] Simulation verification: Validating the effectiveness and security of the new strategy within a digital twin;

[0245] Deploy updates by deploying verified policies to edge devices via a secure channel;

[0246] Effectiveness evaluation involves assessing the strategy's effectiveness during actual operation to complete the iterative closed loop.

[0247] Specifically, by establishing a complete closed-loop iterative system of monitoring, evaluation, training, verification, and deployment, the collaborative optimization strategy can be continuously improved, ensuring that the system can adapt to changes in the operating conditions of the transformer area and maintain optimal operating status in the long term. This automated iterative mechanism reduces manual maintenance costs and enhances the system's adaptability and long-term operational stability.

[0248] Please refer to Figure 2 Furthermore, a protocol conversion and self-learning system for coordinated photovoltaic-storage-charging systems in power distribution areas, used to implement the above method, includes:

[0249] Digital twin module, used to build and operate a digital twin of the photovoltaic, energy storage and charging system in a distribution area;

[0250] The digital twin module specifically includes the following sub-modules:

[0251] The physical equipment modeling unit, based on dynamic operating data of photovoltaic, energy storage, and charging equipment, establishes accurate mathematical models of photovoltaic inverters, energy storage converters, and charging piles, including:

[0252] MPPT characteristic curve of photovoltaic inverter, charge / discharge efficiency-SOC relationship curve of energy storage system, constant power / constant current working mode of charging pile;

[0253] The communication network simulation unit realizes the simulation of message sequences in multi-source heterogeneous communication networks, supports the complete communication stack simulation of Modbus, CAN and IEC104 protocols, and can simulate network behaviors such as message transmission delay and packet loss.

[0254] The extreme condition injection unit simulates communication failures and load surges, and provides a graphical interface for configuring the fault type, occurrence time, and duration.

[0255] The agent collaborative training module is used to run and train protocol agents and specification agents in a digital twin;

[0256] The hierarchical architecture management unit implements a hierarchical reinforcement learning architecture, manages the shared neural network layer of the protocol agent and the specification agent, and coordinates the training cycle and parameter synchronization of the two agents.

[0257] The state-action configuration unit sets the core state inputs of the protocol agent, protocol conversion delay, message packet loss rate, and core state inputs of the protocol agent, as well as the power requirements of the energy storage SOC and charging pile, and defines the corresponding action space according to the configuration requirements.

[0258] The reward coupling calculation unit realizes the mutual coupling mechanism of reward functions and calculates the contribution weight of the protocol agent to the protocol control and the contribution weight of the protocol agent to the protocol communication in real time.

[0259] The dynamic weight adjustment unit performs dynamic adjustment of contribution weights and automatically optimizes the priority configuration of communication quality and system energy efficiency according to different runtime periods.

[0260] The strategy management and deployment module is used to manage collaborative optimization strategies and securely distribute and deploy them to edge control devices in physical areas.

[0261] The strategy management and deployment module specifically includes the following sub-modules:

[0262] The policy encapsulation unit encapsulates the parameters of the trained and stable policy model into a policy update file with a size not exceeding 2MB, containing complete version information and verification data;

[0263] The secure transmission unit establishes a secure channel based on TLS 1.3 to ensure the secure transmission of policy update files from the digital twin to the edge control device, and supports breakpoint resumption and integrity verification.

[0264] The cycle control unit strictly manages the data exchange timing between the digital twin and the edge control devices of the physical station area, in accordance with the 100ms interaction cycle requirement.

[0265] The iterative triggering unit enables automated verification and update triggering mechanisms, monitors communication quality indicators and system energy efficiency indicators in real time, and automatically starts the strategy iteration process when a preset threshold is reached.

[0266] The three modules are integrated through a unified service bus to form a complete collaborative workflow:

[0267] Data stream integration: Real-time operational data from the physical transformer area is uploaded to the digital twin module via edge acquisition terminals to drive high-fidelity simulation;

[0268] Training stream integration: Simulation data generated by the digital twin module is provided to the agent collaborative training module for policy optimization;

[0269] Deployment flow integration: The trained collaborative optimization strategy is securely distributed to physical stations through the strategy management and deployment module;

[0270] The system adopts a microservice architecture, where each module can be deployed and scaled independently. It communicates through a RESTful API and supports high-concurrency access and load balancing.

[0271] Specifically, a high-fidelity simulation environment is provided through a digital twin module, a cross-layer optimization algorithm is implemented through an intelligent agent collaborative training module, and a strategy management and deployment module ensures the safe and reliable deployment of strategies. Together, these three components constitute a complete transformer substation optical storage and charging protocol conversion and self-learning system.

[0272] Although alternative embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make further changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0273] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this invention.

Claims

1. A protocol conversion and self-learning method for photovoltaic-storage-charging coordination in power distribution areas, characterized in that: Includes the following steps: A digital twin of the photovoltaic, energy storage and charging system of a distribution area is constructed. The digital twin is used to perform high-fidelity mapping of the dynamic operation data of the photovoltaic, energy storage and charging equipment, the message sequence of the multi-source heterogeneous communication network and the topology of the distribution area. In the digital twin, the protocol agent and the specification agent run in parallel; The protocol agent is used to learn dynamic priority scheduling and message compression strategies by parsing heterogeneous protocol messages such as Modbus, CAN, and IEC104. The specification agent is used to learn power balance and voltage stability control strategies for the distribution area by analyzing photovoltaic power output fluctuations, energy storage SOC, and charging pile power demand. By using a reinforcement learning framework, the protocol agent and the specification agent are trained and make decisions collaboratively. The weighted indicators of communication delay, packet loss rate, system network loss, and voltage deviation are used as joint optimization objectives to generate a collaborative optimization strategy. The collaborative optimization strategy is deployed to the edge control equipment of the physical distribution area, and the protocol command conversion and collaborative control of the photovoltaic inverter, energy storage converter and charging pile are executed based on the real-time collected distribution area operating status.

2. The method according to claim 1, characterized in that, The protocol agent and the specification agent collaborate using a hierarchical reinforcement learning architecture, specifically including: The protocol agent and the specification agent share some neural network layers; The observation state space of the protocol agent includes the protocol conversion delay, packet loss rate and channel utilization of the multi-source heterogeneous communication network, and the action space includes dynamic priority scheduling and compression instructions for protocol packet processing. The observation state space of the protocol intelligent agent includes real-time photovoltaic power output, state of charge (SOC) of the energy storage system, and real-time power demand of the charging pile. The action space includes adjustment commands for the energy storage charging and discharging threshold, photovoltaic inverter output power, and charging pile power limit.

3. The method according to claim 2, characterized in that, The protocol agent is configured as follows: Using the protocol conversion delay and packet loss rate as the core state inputs, the system optimizes and outputs scheduling instructions for data packets of different priorities. The protocol agent is configured as follows: Using the energy storage SOC and the charging pile's required power as the core state inputs, the output is optimized through entropy increase to provide a stepwise adjustment command for the energy storage charging and discharging thresholds.

4. The method according to claim 2, characterized in that, The protocol agent and the specification agent coordinate through the mutual coupling of reward functions, specifically: The instant reward value of the protocol agent is determined by the contribution weight of the real-time action to the reduction rate of line loss in the corresponding transformer area controlled by the protocol, and the contribution weight of the improvement of voltage qualification rate. The instantaneous reward value of the protocol agent is determined by the contribution weight of the real-time action to the reduction rate of the average transmission delay corresponding to the protocol communication and the contribution weight of the improvement of the packet loss rate of key control commands.

5. The method according to claim 4, characterized in that, The calculation of the contribution weight is dynamically adjusted based on the optimization priority of communication quality and system energy efficiency under different operating periods.

6. The method according to claim 1, characterized in that, The collaborative optimization strategy is deployed to the edge control devices of the physical distribution area, specifically as follows: The parameters of the trained and stable policy model are encapsulated into a policy update file and distributed from the digital twin to the edge control device through a secure channel; After loading the policy update file, the edge control device outputs protocol conversion and specification control commands in real time according to the current state.

7. The method according to claim 6, characterized in that, The size of the policy update file is limited to 2MB, and the interaction cycle between the digital twin and the edge control device of the physical station area is controlled to be within 100ms.

8. The method according to claim 1, characterized in that, Extreme operating conditions, including communication failures and load surges, are simulated in the digital twin, and corresponding system recovery strategies are generated and integrated into the collaborative optimization strategy.

9. The method according to claim 1, characterized in that, The method also includes a strategy iteration step: Using the digital twin, the collaborative optimization strategy is automatically verified and updated based on preset communication quality index thresholds and system energy efficiency index thresholds.

10. A protocol conversion and self-learning system for implementing the method described in any one of claims 1 to 9, characterized in that, include: A digital twin module is used to construct and operate a digital twin of the photovoltaic, energy storage, and charging system in the aforementioned area; An agent collaborative training module is used to run and train the protocol agent and the specification agent in the digital twin; The strategy management and deployment module is used to manage the collaborative optimization strategy and securely distribute and deploy it to the edge control devices of the physical substation area.

Citation Information

Patent Citations

  • Protocol conversion and communication adaptation system for optical storage and charging grid-connected device

    CN120729963A

  • Digital twin energy management method and system for source network load storage cooperative scheduling

    CN120749911A