Communication method and device of power distribution network equipment, electronic equipment, medium and product
By acquiring the communication status and action information of power distribution network equipment and optimizing communication strategies using reinforcement learning models, the communication reliability problem of power distribution network equipment in complex environments is solved, multi-link parallel deployment and dynamic decision-making are realized, and the stability and adaptability of data transmission are improved.
Patent Information
- Application Number
- CN202511817175.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-01-13
AI Technical Summary
The communication reliability of power distribution network equipment is poor, and data packet loss and communication link interruption are prone to occur, especially in complex geographical environments where stable data transmission is difficult to achieve.
By acquiring the communication status information of the distribution network equipment and the action information of the previous communication cycle, and combining it with a reinforcement learning model, the system dynamically selects communication links or shuts down data communication, optimizing communication strategies to adapt to quality fluctuations in complex environments.
It improves the communication reliability and scenario adaptability of power distribution network equipment, avoids anomalies caused by single communication link failures, enables parallel deployment of multiple links to compensate for limitations, and allows dynamic decision-making to adapt to environmental changes.
Smart Images

Figure CN121333931A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a communication method, apparatus, electronic device, medium and product for power distribution network equipment. Background Technology
[0002] In modern power grid transmission systems, distribution network equipment undertakes important functions such as monitoring and data acquisition. Because distribution network equipment is typically widely distributed across different geographical locations within the transmission system, stable data transmission is essential.
[0003] In some technologies, environmental interference and communication link failures may lead to communication anomalies in distribution network equipment, resulting in data packet loss, communication link interruptions, and other phenomena. The communication reliability of distribution network equipment in these technologies is poor.
[0004] Therefore, there is an urgent need for a solution that can improve the reliability of data communication in power distribution network equipment. Summary of the Invention
[0005] The communication methods, devices, electronic equipment, media, and products for power distribution network equipment provided in this application are used to improve the reliability of data communication of power distribution network equipment, thereby ensuring the reliable operation of modern power grid transmission systems.
[0006] In a first aspect, embodiments of this application provide a communication method for power distribution network equipment, including:
[0007] Acquire status information and action information; whereby status information represents the communication status of the distribution network equipment in the current communication cycle, and action information represents the communication action performed by the distribution network equipment in the previous communication cycle.
[0008] The state information and action information are input into a preset reinforcement learning model for processing to obtain the target communication action of the distribution network equipment in the current communication cycle; wherein, the target communication action is used to instruct the distribution network equipment to select at least one communication link for data communication in the current communication cycle, or the target communication action is used to instruct the distribution network equipment to shut down data communication in the current communication cycle.
[0009] In one possible implementation, the reinforcement learning model is configured with a set of communication actions, which includes multiple communication actions.
[0010] The state and action information are input into a preset reinforcement learning model for processing to obtain the target communication actions of the distribution network equipment in the current communication cycle, including:
[0011] Based on the reinforcement learning model, the action score of each communication action in the communication action set is determined according to the state information and action information.
[0012] The communication action with the highest action score in the set of communication actions is determined as the target communication action for the current communication cycle.
[0013] In one possible implementation, based on a reinforcement learning model, the action score for each communication action in the communication action set is determined according to state information and action information, including:
[0014] Based on the reinforcement learning model, each communication action in the communication action set is predicted according to the state information to obtain the communication result corresponding to each communication action; where the communication result represents whether the distribution network equipment has successfully completed data communication in the current communication cycle;
[0015] Based on the reinforcement learning model, the negative correlation value corresponding to each communication action in the communication action set is determined according to the action information.
[0016] Based on the communication results and negative correlation values, determine the action score for each communication action.
[0017] In one possible implementation, based on a reinforcement learning model, the negative correlation value corresponding to each communication action in the communication action set is determined according to action information, including:
[0018] Based on action information, each communication action in the communication action set is compared to obtain the matching result corresponding to each communication action; wherein, the matching result indicates whether the communication action selected by the distribution network equipment in the current communication cycle is the same as the communication action executed in the previous communication cycle;
[0019] Based on the matching results and the preset power consumption of each communication action, the negative correlation value corresponding to each communication action is determined.
[0020] In one possible implementation, the status information includes at least one of the following: basic information of the distribution network equipment, the amount of data to be transmitted, environmental information of the distribution network equipment, and channel quality of at least one communication link.
[0021] In one possible implementation, the reinforcement learning model is a reinforcement learning model that includes an actor network and a critic network, wherein the actor network has a first network parameter and the critic network has a second network parameter.
[0022] The method also includes:
[0023] Control the power distribution network equipment to execute target communication actions and detect the actual execution effect of the target communication actions;
[0024] Based on the actual execution results, update the first network parameters and the second network parameters.
[0025] Secondly, embodiments of this application provide a communication device for power distribution network equipment, comprising:
[0026] The acquisition module is used to acquire status information and action information; wherein, the status information represents the communication status of the distribution network equipment in the current communication cycle, and the action information represents the communication action performed by the distribution network equipment in the previous communication cycle.
[0027] The processing module is used to input the state information and action information into a preset reinforcement learning model for processing, so as to obtain the target communication action of the distribution network equipment in the current communication cycle; wherein, the target communication action is used to instruct the distribution network equipment to select at least one communication link for data communication in the current communication cycle, or the target communication action is used to instruct the distribution network equipment to shut down data communication in the current communication cycle.
[0028] In one possible implementation, the reinforcement learning model is configured with a set of communication actions, which includes multiple communication actions.
[0029] The state and action information are input into a preset reinforcement learning model for processing to obtain the target communication action of the distribution network equipment in the current communication cycle. The processing module is used for:
[0030] Based on the reinforcement learning model, the action score of each communication action in the communication action set is determined according to the state information and action information.
[0031] The communication action with the highest action score in the set of communication actions is determined as the target communication action for the current communication cycle.
[0032] In one possible implementation, based on a reinforcement learning model, the action score of each communication action in the communication action set is determined according to state information and action information. The processing module is used for:
[0033] Based on the reinforcement learning model, each communication action in the communication action set is predicted according to the state information to obtain the communication result corresponding to each communication action; where the communication result represents whether the distribution network equipment has successfully completed data communication in the current communication cycle;
[0034] Based on the reinforcement learning model, the negative correlation value corresponding to each communication action in the communication action set is determined according to the action information.
[0035] Based on the communication results and negative correlation values, determine the action score for each communication action.
[0036] In one possible implementation, based on a reinforcement learning model, the negative correlation value corresponding to each communication action in the communication action set is determined according to action information. The processing module is used to:
[0037] Based on action information, each communication action in the communication action set is compared to obtain the matching result corresponding to each communication action; wherein, the matching result indicates whether the communication action selected by the distribution network equipment in the current communication cycle is the same as the communication action executed in the previous communication cycle;
[0038] Based on the matching results and the preset power consumption of each communication action, the negative correlation value corresponding to each communication action is determined.
[0039] In one possible implementation, the status information includes at least one of the following: basic information of the distribution network equipment, the amount of data to be transmitted, environmental information of the distribution network equipment, and channel quality of at least one communication link.
[0040] In one possible implementation, the reinforcement learning model is a reinforcement learning model that includes an actor network and a critic network, wherein the actor network has a first network parameter and the critic network has a second network parameter.
[0041] The processing module is also used for:
[0042] Control the power distribution network equipment to execute target communication actions and detect the actual execution effect of the target communication actions;
[0043] Based on the actual execution results, update the first network parameters and the second network parameters.
[0044] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0045] The memory stores the instructions that the computer executes;
[0046] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0047] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0048] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0049] The communication method, apparatus, electronic device, medium, and product for distribution network equipment provided in this application obtain the communication status information of the distribution network equipment, combine it with the communication actions of the previous communication cycle, and process it using a preset reinforcement learning model. Finally, the communication actions that the distribution network equipment should perform in the current communication cycle are obtained. These communication actions include selecting at least one communication link for data communication or disabling data communication. On the one hand, the ability to select at least one communication link for data communication, through the parallel deployment of multiple communication links, compensates for the limitations of a single communication link, avoids communication anomalies caused by the failure of a single communication link, and improves the communication reliability of the distribution network equipment. On the other hand, by integrating the multi-dimensional communication status information of the distribution network equipment and the communication actions of the previous communication cycle, combined with reinforcement learning for prediction, dynamic decision-making can be achieved, enabling the strategy to adapt to communication quality fluctuations in complex environments, improving the scenario adaptability of the distribution network equipment's communication process, and further enhancing the communication reliability of the distribution network equipment. Attached Figure Description
[0050] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0051] Figure 1 Flowchart of the communication method for the power distribution network equipment provided in this application Figure 1 ;
[0052] Figure 2 Flowchart of the communication method for the power distribution network equipment provided in this application Figure 2 ;
[0053] Figure 3 Flowchart of the communication method for the power distribution network equipment provided in this application Figure 3 ;
[0054] Figure 4 A schematic diagram of the communication device for the power distribution network equipment provided in this application;
[0055] Figure 5 A schematic diagram of the structure of the electronic device provided in this application.
[0056] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0057] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0058] First, let me explain the terms used in this application:
[0059] Communication action: This refers to the action taken by distribution network equipment when selecting a communication link for data communication. For example, distribution network equipment may select one communication link for data communication, select multiple communication links in parallel for data communication, or shut down data communication.
[0060] Communication link: refers to the path established based on a communication protocol to enable data communication between power distribution network equipment. For example, based on the IPSec (IP Security) protocol extension, a separate communication link is established for each access communication method. The IPSec protocol is a network layer security protocol suite of the TCP / IP protocol suite. Its core function is to provide end-to-end encryption, authentication, integrity protection, and anti-replay attack capabilities for IP packets. It does not rely on application layer protocols and is suitable for the secure transmission of sensitive data in public or private networks.
[0061] Reinforcement learning models are machine learning models in which an agent optimizes its decision-making strategy based on reward signals through interaction with its environment. Reinforcement learning models can include various types, such as value-driven reinforcement learning models, policy-driven reinforcement learning models, and actor-critic models.
[0062] Negative correlation value: refers to the combined value of the action cost of communication actions to be performed by distribution network equipment in the current communication cycle and the channel switching penalty.
[0063] In modern power grid transmission systems, distribution network equipment is widely distributed in complex geographical environments such as urban streets, rural fields, and remote mountainous areas, undertaking key functions such as data acquisition, condition monitoring, and remote control of power grid infrastructure. Because power grid infrastructure requires long-distance power transmission and multi-level substations, distribution network equipment must achieve stable data transmission under various communication conditions.
[0064] However, significant communication challenges exist in real-world scenarios: electromagnetic interference in densely populated urban areas, insufficient signal coverage in rural areas, and fragile communication links in mountainous and island regions can easily lead to packet loss, communication link interruptions, and other communication anomalies during data communication. This results in poor communication reliability for power distribution network equipment.
[0065] In some embodiments, data communication of power distribution network equipment primarily relies on communication links using a single communication method, such as fiber optic or 4G. Security is ensured by constructing the communication link using the IPSec protocol. Different communication methods have different characteristics. For example, fiber optic communication is widely used in urban areas due to its high bandwidth and low latency, but its deployment cost is high and it is susceptible to physical damage; while 4G wireless communication offers flexible deployment, its signal strength is significantly affected by geographical environment and weather conditions, leading to fluctuations in link quality. Communication links constructed using a single communication method suffer from low reliability in the data communication process of power distribution network equipment. However, in communication link configurations using multiple communication methods, the relevant embodiments lack proactive management of link redundancy and cannot dynamically adjust the communication method according to the implementation environment.
[0066] Based on the above scenarios, it can be seen that the data communication of power distribution network equipment suffers from poor reliability in related technologies.
[0067] The communication method for distribution network equipment provided in this application obtains the communication status information of the distribution network equipment, combines it with the communication actions of the previous communication cycle, and processes it using a preset reinforcement learning model. The final result is the communication action that the distribution network equipment should perform in the current communication cycle. This communication action includes selecting at least one communication link for data communication or disabling data communication. On the one hand, the ability to select at least one communication link for data communication, through the parallel deployment of multiple communication links, compensates for the limitations of a single communication link, avoiding communication anomalies caused by the failure of a single communication link, and improving the communication reliability of the distribution network equipment. On the other hand, by integrating the multi-dimensional communication status information of the distribution network equipment and the communication actions of the previous communication cycle, combined with reinforcement learning for prediction, dynamic decision-making can be achieved. This allows the strategy to adapt to communication quality fluctuations in complex environments, improving the scenario adaptability of the distribution network equipment's communication process, thereby further enhancing the communication reliability of the distribution network equipment.
[0068] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0069] Figure 1 Flowchart of the communication method for the power distribution network equipment provided in this application Figure 1 ,like Figure 1 As shown, the method includes:
[0070] Step 101. Obtain status information and action information.
[0071] Among them, status information represents the communication status of the distribution network equipment in the current communication cycle, and action information represents the communication action performed by the distribution network equipment in the previous communication cycle.
[0072] For example, the status information of distribution network equipment is obtained. This status information characterizes the communication status of the distribution network equipment in the current communication cycle. Specifically, the communication status may include any one or a combination of the following information: whether the distribution network equipment is ready to send data, whether the distribution network equipment is ready to receive data, the current geographical location of the distribution network equipment, environmental information of the environment in which the distribution network equipment is located, the channel quality of the communication link through which the distribution network equipment performs communication actions, and the amount of data to be transmitted. This information characterizes the relevant equipment status information of the distribution network equipment in the current communication cycle when data communication is to be performed.
[0073] For example, the action information of the distribution network equipment is obtained. This action information represents the communication actions of the distribution network equipment in the previous communication cycle. Specifically, the exact communication actions of the distribution network equipment in the previous communication cycle can be stored in relevant storage components configured within the distribution network equipment. By reading instructions, the time information of the previous communication cycle is determined from the relevant storage components based on the current communication cycle. Based on the time information of the previous communication cycle, the corresponding communication action is queried from the relevant storage components to obtain the action information.
[0074] Step 102. Input the state information and action information into the preset reinforcement learning model for processing to obtain the target communication action of the power distribution network equipment in the current communication cycle.
[0075] The target communication action is used to instruct the distribution network equipment to select at least one communication link for data communication in the current communication cycle, or the target communication action is used to instruct the distribution network equipment to shut down data communication in the current communication cycle.
[0076] For example, the preset reinforcement learning model can be a reinforcement learning model based on an actor-critic network. The state information and action information of the power distribution network equipment are input into the actor-critic network-based reinforcement learning model for processing.
[0077] The preset reinforcement learning model is configured with multiple communication actions. Based on the state information and action information, the reinforcement learning model selects one communication action from the multiple communication actions as the communication action that the power distribution network equipment should execute to transmit data in the current communication cycle.
[0078] In one scenario, the reinforcement learning model, based on state and action information, determines that data communication is not advisable for distribution network equipment during the current communication cycle. Therefore, the target communication action is to shut down data communication for the distribution network equipment during the current communication cycle. For example, in a practical application, if the channel quality is generally poor during the current communication cycle, potentially leading to data communication failures, then data communication can be shut down during the current cycle and resumed once the channel quality recovers.
[0079] In one scenario, the reinforcement learning model, based on state and action information, determines that the distribution network equipment can communicate via a communication link in the current communication cycle. The target communication action is then determined to be that the distribution network equipment selects a communication link for data communication. For example, in a practical application, if the amount of data to be transmitted in the current communication cycle is small and a communication link with good channel quality exists, then that high-quality communication link can be selected to transmit the smaller amount of data.
[0080] In one scenario, the reinforcement learning model, based on state and action information, determines that the distribution network equipment can communicate via multiple communication links during the current communication cycle. Therefore, the target communication action is determined to be that the distribution network equipment selects multiple communication links for data communication. For example, in a practical application scenario where the environment in which the distribution network equipment operates during the current communication cycle is harsh and the channel quality is poor, multiple communication links can be selected for redundant communication to ensure successful and reliable data communication.
[0081] It should be noted that the communication link is an extension of the IPSec (IP Security) protocol, establishing a separate path for data communication between power distribution network equipment for each access communication method. These communication methods can include fiber optic communication, 4G communication, and satellite communication. Specifically, fiber optic communication establishes a site-to-site communication tunnel using a fixed IP address, while 4G and satellite communication utilize NAT traversal technology to construct communication tunnels.
[0082] Among them, the IPSec protocol is used to establish encrypted communication channels between heterogeneous communication links (fiber optic, 4G, satellite) to ensure data transmission security.
[0083] Among them, NAT traversal technology is used to solve the problem of direct communication barriers between devices caused by Network Address Translation (NAT) in 4G communication.
[0084] The communication method for power distribution network equipment provided in this application obtains the communication status information of the power distribution network equipment, combines it with the communication actions of the previous communication cycle, and processes it using a preset reinforcement learning model. Finally, it determines the communication actions that the power distribution network equipment should perform in the current communication cycle. These communication actions include selecting at least one communication link for data communication or disabling data communication.
[0085] On the one hand, during communication operations, at least one communication link can be selected for data communication. By deploying multiple communication links in parallel, the limitations of a single communication link can be compensated for, and communication failures of a single communication link can be avoided, thereby improving the communication reliability of the distribution network equipment.
[0086] On the other hand, by integrating multi-dimensional communication status information of distribution network equipment and combining the communication actions of the previous communication cycle with reinforcement learning for prediction, dynamic decision-making can be achieved, enabling the strategy to adapt to communication quality fluctuations in complex environments, improving the scenario adaptability of the communication process of distribution network equipment, and thus further improving the communication reliability of distribution network equipment.
[0087] Figure 2 Flowchart of the communication method for the power distribution network equipment provided in this application Figure 2 In this embodiment Figure 1 Based on the example, the process of how the reinforcement learning model determines the target communication action based on state information and action information in step 102 is explained in detail.
[0088] In one example, the reinforcement learning model is configured with a set of communication actions, which includes multiple communication actions.
[0089] For example, the communication action set includes multiple communication actions. Each communication action represents an action corresponding to a communication method supported by the distribution network equipment, as well as a combination of actions corresponding to the supported communication methods.
[0090] For example, the set of communication actions includes the following communication actions: using fiber optic communication, using 4G communication, using satellite communication, using both fiber optic and 4G communication, using both fiber optic and satellite communication, using both 4G and satellite communication, using both fiber optic, 4G and satellite communication, and shutting down data communication.
[0091] like Figure 2 As shown, the method includes:
[0092] Step 201. Based on the reinforcement learning model, determine the action score of each communication action in the communication action set according to the state information and action information.
[0093] For example, state information and action information are input into a reinforcement learning model. Based on the reinforcement learning model, each communication action in the communication action set is dynamically scored according to the state information and action information, resulting in an action score for each communication action.
[0094] Action scoring refers to a metric that quantifies the effectiveness of communication action selection. For example, it can be a comprehensive score obtained by subtracting action costs and handover penalties from the transmission success rate.
[0095] Specifically, based on reinforcement learning models and state and action information, the overall state of distribution network equipment in the current communication cycle can be determined, forming the foundation for decision-making. Based on the current communication cycle, the probability distribution of entering the next state after performing different communication actions is calculated, thereby obtaining action scores for the communication actions and providing forward-looking references.
[0096] Step 202. Determine the communication action with the highest action score in the set of communication actions as the target communication action for the current communication cycle.
[0097] For example, based on the action score of each communication action in the set of communication actions, the communication action with the highest action score is output. The communication action with the highest action score is then determined as the target communication action for the current communication cycle.
[0098] It is understandable that a target communication action can be used to instruct distribution network equipment to perform data communication in the current communication cycle. This ensures the success rate and reliability of data communication in the current communication cycle, while balancing the resource consumption of data communication.
[0099] In the above embodiments, a reinforcement learning model is used to dynamically score the communication actions in the set of communication actions configured in the reinforcement learning model, combining the status information and action information of the distribution network equipment. The highest score of the communication action is then used as the communication action of the distribution network equipment in the current communication cycle. This achieves intelligent processing of communication mode selection, enabling the selection of more reliable communication actions based on a machine learning model, combined with the communication actions of the previous communication cycle and the communication status of the current communication cycle, thereby improving the reliability of data communication between distribution network equipment.
[0100] In addition, communication actions can instruct distribution network equipment to select at least one communication link for data communication. This can optimize resource utilization efficiency while taking into account both short-term transmission performance and long-term stability.
[0101] Figure 3 Flowchart of the communication method for the power distribution network equipment provided in this application Figure 3 In this embodiment Figure 2Based on the example, the process of how the reinforcement learning model determines the action score of the communication action according to the state information and action information in step 201 is explained in detail.
[0102] like Figure 3 As shown, the method includes:
[0103] Step 301. Based on the reinforcement learning model, predict each communication action in the communication action set according to the state information to obtain the communication result corresponding to each communication action.
[0104] The communication result characterizes whether the power distribution network equipment has successfully completed data communication in the current communication cycle.
[0105] For example, based on a reinforcement learning model, the communication actions in the communication action set are predicted according to the comprehensive state of the distribution network equipment represented in the state information during the current communication cycle, thus obtaining the communication result for the current communication cycle. The communication result represents whether data communication can be completed by using the communication actions in the communication action set.
[0106] Optionally, the communication result can be represented by the value 1 or 0. A communication result of 1 indicates successful transmission, while a communication result of 0 indicates failure.
[0107] Step 302. Based on the reinforcement learning model, determine the negative correlation value corresponding to each communication action in the communication action set according to the action information.
[0108] Specifically, the negative correlation value corresponding to the communication action can be determined according to the following steps:
[0109] Step 3021. Based on the action information, compare each communication action in the communication action set to obtain the matching result corresponding to each communication action.
[0110] The matching result indicates whether the communication action selected by the distribution network equipment in the current communication cycle is the same as the communication action performed in the previous communication cycle.
[0111] For example, as can be seen from the foregoing embodiments, the action information refers to the communication actions of the distribution network equipment in the previous communication cycle. The action information is matched with each communication action in the communication action set to obtain a matching result. The matching result indicates whether the communication action selected by the distribution network equipment in the current communication cycle is the same as the communication action performed in the previous communication cycle. It can be understood that the matching result indicates whether the distribution network equipment needs to switch communication actions from the previous communication cycle to the current communication cycle.
[0112] Optionally, the matching result can be represented by the value 0 or other non-zero constants. Specifically, when the communication action in the current communication cycle (i.e., from the set of communication actions) is different from the communication action in the previous communication cycle, the matching result is a preset non-zero constant; when they are the same, the matching result is 0.
[0113] Step 3022. Based on the matching results and the preset power consumption of each communication action, determine the negative correlation value corresponding to each communication action.
[0114] For example, each communication action in the set of communication actions corresponds to a power consumption value for the communication method required for that action, which serves as the preset power consumption for the communication action. The power consumption value corresponding to a communication action can be understood as the action cost of the communication action. Since each communication action may correspond to selecting one or more communication links for data communication, it can be understood that the power consumption value corresponding to a communication action that selects multiple communication links is greater than the power consumption value of a communication action that selects only one communication link.
[0115] Furthermore, based on the matching results corresponding to each communication action determined in step 3021 and the preset power consumption value, the results are summed to obtain the negative correlation value corresponding to each communication action.
[0116] The matching results indicate whether a switching occurred between adjacent communication cycles, while the power consumption of the communication action reflects the resource consumption required during its actual execution. The negative correlation value of the communication action is calculated from a more granular perspective, laying the foundation for more accurate action scoring of subsequent communication actions.
[0117] Step 303. Determine the action score for each communication action based on the communication results and negative correlation values.
[0118] For example, the action score corresponding to each communication action is obtained by subtracting the negative correlation value determined in step 3022 from the communication result corresponding to each communication action determined in step 301.
[0119] Optionally, in the process of determining the action score corresponding to each communication action, the weight coefficients of the current action score and the future state score can be dynamically adjusted according to the urgency of the communication environment.
[0120] The weighting coefficient refers to a parameter that dynamically adjusts the ratio of the current action score to the future state score. For example, increasing the weighting coefficient in an emergency scenario allows the reinforcement learning model to focus more on near-term results. Specifically, the weighting coefficient can be dynamically adjusted using preset rules.
[0121] For example, in heavy rain (emergency scenarios), the weighting coefficient approaches 1, prioritizing the success rate of transmission in the current cycle (e.g., activating satellite communication); in normal scenarios, the weighting coefficient approaches 0, optimizing long-term communication stability (e.g., reducing frequent handovers). This dynamic weighting adjustment mechanism allows the communication strategy to flexibly adapt to different scenario requirements, improving the strategy's scenario adaptability and further reducing the risk of service interruption in emergency scenarios.
[0122] Optionally, the weighting coefficient can range from 0 to 1. The closer the value is to 1, the more emphasis is placed on recent effects.
[0123] In the above embodiments, the success of the data communication process for a communication action is reflected in the communication result, demonstrating the reliability of the communication action. Negative correlation values reflect the resource consumption required for the communication action (e.g., whether a communication link needs to be switched, and the actual power consumption of the communication link corresponding to the communication action). Based on these two dimensions, an action score for the communication action is determined. This ensures that the action score balances the required resource consumption while guaranteeing the success rate of data communication, improving the rationality of communication actions of distribution network equipment in the current communication cycle, and ultimately enhancing the reliability of data communication in the distribution network equipment.
[0124] To more accurately determine the communication actions of distribution network equipment in the current communication cycle, the communication status of distribution network equipment can be characterized by diverse status information.
[0125] In one example, the status information includes at least one of the following: basic information of the distribution network equipment, the amount of data to be transmitted, environmental information of the distribution network equipment, and channel quality of at least one communication link.
[0126] For example, the acquired status information may include one or more types of information.
[0127] The basic information of distribution network equipment refers to its current geographical location. By determining the specific environment in which the equipment is located (such as mountainous areas, islands, cities, rural areas, etc.), it is possible to quickly determine whether it is in a remote area with weak communication signals. This basic information is directly related to the prediction of the complexity of the communication environment. For example, remote mountainous areas may face even worse signal coverage, requiring advance consideration of communication methods with stronger anti-interference capabilities.
[0128] The data volume to be transmitted refers to the amount of data that the power distribution network equipment needs to transmit in the current communication cycle. The size of the data volume is a key indicator for selecting an appropriate communication method. For example, when the data volume is large, a communication method with higher bandwidth (such as fiber optic) may be preferred; while when the data volume is small, a more flexible wireless communication method (such as 4G) can be considered, thereby optimizing the utilization efficiency of communication resources while meeting transmission requirements.
[0129] Among them, the environmental information of the distribution network equipment refers to the time-limited information of the environment in which the distribution network equipment is located during the current communication cycle. It focuses on meteorological factors that directly affect communication, such as air humidity and weather conditions (heavy rain, strong winds, etc.). This information is used to predict the degree of environmental interference to wireless communication (such as 4G, satellite communication). For example, heavy rain may weaken 4G signal strength, providing a basis for subsequently selecting communication methods with suitable anti-interference capabilities.
[0130] Among these, the channel quality of at least one communication link refers to the individual communication quality indicators of communication links supporting different communication methods within the distribution network equipment during the current communication cycle. These indicators include, for example, bandwidth, packet loss rate, and latency. These data directly reflect the real-time transmission capability of each communication link and are the core objective standard for determining whether a communication link is suitable.
[0131] In the above example, the basic information, data volume, environmental information, channel quality and other multi-dimensional factors of the distribution network equipment are comprehensively considered. The communication effect is dynamically evaluated in combination with the communication actions of the distribution network equipment in the previous communication cycle. This enables the strategy to adapt to communication quality fluctuations in complex environments and improve its adaptability to remote scenarios.
[0132] Based on any of the foregoing embodiments or examples, reinforcement learning models can use the actor-critic model.
[0133] Specifically, the reinforcement learning model is a reinforcement learning model that includes an actor network and a critic network, with the actor network having a first network parameter and the critic network having a second network parameter.
[0134] For example, in a reinforcement learning model, the actor network is used to output an action probability distribution or deterministic actions (different communication actions) based on a set environmental state (including state information and action information).
[0135] In reinforcement learning models, the critic network is used to evaluate value. Specifically, it calculates action scores for different communication actions by combining the given environmental state (including state information and action information).
[0136] The actor network has a first set of network parameters, including learnable parameters such as network weights and biases. These first set of network parameters determine the output of the communication action.
[0137] The critic network possesses a second set of network parameters, including weights and biases. These second set of network parameters determine the accuracy of the critic network's estimation of action scores for communication actions and are the core object of value assessment calibration.
[0138] Based on the above, the method may also include:
[0139] Control the power distribution network equipment to execute the target communication action, detect the actual execution effect of the target communication action, and update the first network parameter and the second network parameter based on the actual execution effect.
[0140] For example, in practical applications, based on the target communication action output by the reinforcement learning model, the power distribution network equipment is controlled to execute the target communication action in the current communication cycle. This enables the power distribution network equipment to conduct data communication through one or more communication links indicated by the target communication action in the current communication cycle. The actual execution effect of the power distribution network equipment executing the target communication action is then detected.
[0141] The actual performance can include evaluation methods in multiple dimensions, including but not limited to: transmission success rate and channel quality changes.
[0142] Furthermore, the model parameters are updated in reverse based on the actual execution results. Specifically, this includes:
[0143] (1) Update the first network parameters of the actor network. Optionally, the first network parameters are denoted as θ. With the goal of maximizing the first objective function L(θ) and the average action score J(θ) of all agents, the update range is limited by a pruning mechanism to avoid excessive parameter fluctuations.
[0144] Its core is to enable actors to more accurately output actions that can obtain high scores and optimize the execution of strategies.
[0145] For example, the update of the first network parameter needs to satisfy the first objective function and the pruning mechanism expression shown in the following expression:
[0146] ;
[0147] ;
[0148] ;
[0149] Here, N represents the total number of agents. An agent is a decision-making model that includes attribute parameters such as state information, a set of communication actions, and future state distribution.
[0150] It should be noted that the future state distribution of an agent refers to the probability distribution of entering the next state after performing each action. This provides the agent with a forward-looking reference, enabling its decisions to consider not only the present but also the future impact. An agent also has an initial state distribution, which serves as the initial reference benchmark for the agent's decisions, reflecting the probability distribution of the device's states in the initial stage.
[0151] τ is the expected action score of agent c under the first network parameter θ; τ is used to reflect the magnitude of policy update, specifically it can be represented by the ratio of the probability of selecting the target action in the current state using the new first network parameter to the probability of the original first network parameter; A represents the advantage estimation function; σ is a hyperparameter, which can generally take the value of 0.2 or other values.
[0152] Taking the current state s and action a as an example, it can be calculated based on the original first network parameters. Based on this, the probability of choosing action a in state s is calculated as the denominator of the ratio, and the new first network parameters are calculated. Based on this, the probability of choosing action a in state s is calculated as the numerator of the ratio, thus obtaining the value of parameter τ, which reflects the policy update magnitude.
[0153] In one possible implementation, the partial derivative of the first objective function L(θ) with respect to the first network parameters is taken to obtain the update gradient of the first network parameters. To avoid excessive fluctuations in the first network parameters, the update gradient is pruned. Based on the pruned gradient, the first network parameters are updated using the gradient ascent method to obtain new first network parameters.
[0154] (2) Update the second network parameters of the critic network. Optionally, the second network parameter is denoted as φ. The goal is to minimize the mean squared error (MSE) by comparing the actual action scores. By adjusting parameters based on the network's estimated scores, the estimation can be made closer to the actual results, thereby improving the ability to evaluate communication actions.
[0155] For example, the update of the second network parameters needs to satisfy the second objective function as shown in the following representation:
[0156] ;
[0157] in, Indicates the state at time t; This represents the actual action score obtained by the agent starting from time t; The commentator network represents the second network parameter. Next state An estimate of the action score can be obtained. Parameters This refers to the score used to represent the future action score under real-world conditions obtained from the actual interaction between the intelligent agent and the environment.
[0158] Furthermore, parameters It can be calculated using the following formula:
[0159] ;
[0160] in, It is a real-time motion score for time step k. These are future weighting coefficients. It should be noted that the parameters... It is an ideal goal, but it cannot be obtained directly in advance; it can only be obtained through post-event interaction records.
[0161] In one possible implementation, the partial derivative of the inverse function of the second objective function with respect to the second network parameters is taken to obtain the updated gradient of the second network parameters. Based on gradient descent, the second network parameters are updated to obtain new second network parameters.
[0162] In the above embodiments, continuous data processing and strategy updates are achieved through an actor critic network, dynamically optimizing the choice of communication methods, taking into account non-real-time business needs, and ensuring reliable data transmission and efficient interaction.
[0163] The communication method for power distribution equipment provided in this application embodiment can select at least one communication link for data communication during the communication operation. By deploying multiple communication links in parallel, the limitations of a single communication link can be made up for, and communication abnormalities of the power distribution equipment due to communication failure of a single communication link can be avoided, thereby improving the communication reliability of the power distribution equipment.
[0164] By integrating multi-dimensional communication status information of distribution network equipment and communication actions from the previous communication cycle, combined with reinforcement learning for prediction, dynamic decision-making can be achieved. This enables the equipment to adapt to communication quality fluctuations in complex environments, improves the scenario adaptability of the communication process of distribution network equipment, and further enhances the communication reliability of distribution network equipment.
[0165] The action score for communication actions is determined by analyzing communication results and negative correlation values. This ensures that the action score balances the required resource and energy consumption while guaranteeing the success rate of data communication, thereby improving the rationality of communication actions of distribution network equipment in the current communication cycle.
[0166] By comprehensively considering the multi-dimensional factors of distribution network equipment and dynamically evaluating the communication effect based on the communication actions of the distribution network equipment in the previous communication cycle, the strategy can adapt to communication quality fluctuations in complex environments and improve its adaptability to remote scenarios.
[0167] By leveraging an actor critic network, continuous data processing and strategy updates are achieved, dynamically optimizing communication methods while accommodating non-real-time business needs, ensuring reliable data transmission and efficient interaction.
[0168] Figure 4 This is a schematic diagram of the communication device of the power distribution network equipment provided in this application, such as... Figure 4 As shown, the communication device 40 of the power distribution network equipment provided in this embodiment includes:
[0169] The acquisition module 401 is used to acquire status information and action information; wherein, the status information represents the communication status of the distribution network equipment in the current communication cycle, and the action information represents the communication action performed by the distribution network equipment in the previous communication cycle.
[0170] The processing module 402 is used to input the state information and action information into a preset reinforcement learning model for processing to obtain the target communication action of the distribution network equipment in the current communication cycle; wherein, the target communication action is used to instruct the distribution network equipment to select at least one communication link for data communication in the current communication cycle, or the target communication action is used to instruct the distribution network equipment to shut down data communication in the current communication cycle.
[0171] In one possible implementation, the reinforcement learning model is configured with a set of communication actions, which includes multiple communication actions.
[0172] The state information and action information are input into a preset reinforcement learning model for processing to obtain the target communication action of the distribution network equipment in the current communication cycle. The processing module 402 is used for:
[0173] Based on the reinforcement learning model, the action score of each communication action in the communication action set is determined according to the state information and action information.
[0174] The communication action with the highest action score in the set of communication actions is determined as the target communication action for the current communication cycle.
[0175] In one possible implementation, based on a reinforcement learning model, the action score of each communication action in the communication action set is determined according to state information and action information. The processing module 402 is used for:
[0176] Based on the reinforcement learning model, each communication action in the communication action set is predicted according to the state information to obtain the communication result corresponding to each communication action; where the communication result represents whether the distribution network equipment has successfully completed data communication in the current communication cycle;
[0177] Based on the reinforcement learning model, the negative correlation value corresponding to each communication action in the communication action set is determined according to the action information.
[0178] Based on the communication results and negative correlation values, determine the action score for each communication action.
[0179] In one possible implementation, based on a reinforcement learning model and action information, the negative correlation value corresponding to each communication action in the communication action set is determined, and the processing module 402 is used to:
[0180] Based on action information, each communication action in the communication action set is compared to obtain the matching result corresponding to each communication action; wherein, the matching result indicates whether the communication action selected by the distribution network equipment in the current communication cycle is the same as the communication action executed in the previous communication cycle;
[0181] Based on the matching results and the preset power consumption of each communication action, the negative correlation value corresponding to each communication action is determined.
[0182] In one possible implementation, the status information includes at least one of the following: basic information of the distribution network equipment, the amount of data to be transmitted, environmental information of the distribution network equipment, and channel quality of at least one communication link.
[0183] In one possible implementation, the reinforcement learning model is a reinforcement learning model that includes an actor network and a critic network, wherein the actor network has a first network parameter and the critic network has a second network parameter.
[0184] Processing module 402 is also used for:
[0185] Control the power distribution network equipment to execute target communication actions and detect the actual execution effect of the target communication actions;
[0186] Based on the actual execution results, update the first network parameters and the second network parameters.
[0187] The communication device for power distribution equipment provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0188] Figure 5 A schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus 504.
[0189] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.
[0190] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0191] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0192] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0193] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0194] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0195] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0196] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0197] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0198] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0199] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0200] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0201] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0202] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0203] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A communication method for power distribution network equipment, characterized in that, include: Acquire status information and action information; wherein, the status information represents the communication status of the distribution network equipment in the current communication cycle, and the action information represents the communication action performed by the distribution network equipment in the previous communication cycle; The state information and the action information are input into a preset reinforcement learning model for processing to obtain the target communication action of the power distribution network equipment in the current communication cycle; wherein, the target communication action is used to instruct the power distribution network equipment to select at least one communication link for data communication in the current communication cycle, or the target communication action is used to instruct the power distribution network equipment to shut down data communication in the current communication cycle.
2. The method according to claim 1, characterized in that, The reinforcement learning model is configured with a set of communication actions, which includes multiple communication actions; The state information and the action information are input into a preset reinforcement learning model for processing to obtain the target communication action of the power distribution network equipment in the current communication cycle, including: Based on the reinforcement learning model, the action score of each communication action in the communication action set is determined according to the state information and the action information. The communication action with the highest action score in the set of communication actions is determined as the target communication action for the current communication cycle.
3. The method according to claim 2, characterized in that, Based on the reinforcement learning model, and according to the state information and the action information, the action score for each communication action in the set of communication actions is determined, including: Based on the reinforcement learning model, each communication action in the communication action set is predicted according to the state information to obtain the communication result corresponding to each communication action; wherein, the communication result represents whether the power distribution network equipment has successfully completed data communication in the current communication cycle; Based on the reinforcement learning model, and according to the action information, the negative correlation value corresponding to each communication action in the communication action set is determined; Based on the communication results and the negative correlation values, the action score for each communication action is determined.
4. The method according to claim 3, characterized in that, Based on the reinforcement learning model, and according to the action information, the negative correlation value corresponding to each communication action in the communication action set is determined, including: Based on the action information, each communication action in the communication action set is compared to obtain a matching result corresponding to each communication action; wherein, the matching result indicates whether the communication action selected by the power distribution network equipment in the current communication cycle is the same as the communication action executed in the previous communication cycle; Based on the matching results and the preset power consumption of each communication action, the negative correlation value corresponding to each communication action is determined.
5. The method according to claim 1, characterized in that, The status information includes at least one of the following: basic information of the power distribution network equipment, the amount of data to be transmitted, environmental information of the power distribution network equipment, and channel quality of at least one communication link.
6. The method according to any one of claims 1-5, characterized in that, The reinforcement learning model is a reinforcement learning model that includes an actor network and a critic network, wherein the actor network has a first network parameter and the critic network has a second network parameter; The method further includes: Control the power distribution network equipment to execute the target communication action, and detect the actual execution effect corresponding to the target communication action; Based on the actual execution results, the first network parameters and the second network parameters are updated.
7. A communication device for power distribution network equipment, characterized in that, include: An acquisition module is used to acquire status information and action information; wherein, the status information represents the communication status of the distribution network equipment in the current communication cycle, and the action information represents the communication action performed by the distribution network equipment in the previous communication cycle; The processing module is used to input the state information and the action information into a preset reinforcement learning model for processing to obtain the target communication action of the power distribution network equipment in the current communication cycle; wherein, the target communication action is used to instruct the power distribution network equipment to select at least one communication link for data communication in the current communication cycle, or the target communication action is used to instruct the power distribution network equipment to shut down data communication in the current communication cycle.
8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-6.