Communication protocol optimization method for large photovoltaic power stations based on partitioning and layering

Through the optimization method of partitioned and layered communication protocols, combined with deep reinforcement learning and improved Dijkstra algorithm, the problems of network congestion and communication conflicts in large-scale photovoltaic power station communication systems are solved, and efficient and reliable data transmission and monitoring are achieved.

CN120223613BActive Publication Date: 2025-09-02CHINA ENERGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510704292.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-02
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Due to the lack of reasonable planning of the geographical distribution and string topology of photovoltaic modules, the communication system of large photovoltaic power stations has unreasonable network structure, low communication efficiency, and the inability to dynamically adjust communication paths and resource allocation, which is prone to network congestion or interruption, affecting monitoring effect and operational safety.

Method used

The communication protocol optimization method of partitioning and layering is adopted. By obtaining the geographical location and string topological data of the photovoltaic power station, the power station is divided into multiple physical partitions, multiple communication levels are established, and the communication time slot allocation sequence is calculated using deep reinforcement learning algorithms and improved Dijkstra algorithms, and the communication time slot allocation is optimized by combining dynamic time windows and causal graph models.

Benefits of technology

It realizes effective communication scheduling of large-scale equipment in photovoltaic power station systems, reduces network delay and data transmission load, improves the overall performance and reliability of the communication system, avoids data transmission conflicts, and provides a reliable data foundation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223613B_ABST
    Figure CN120223613B_ABST
Patent Text Reader

Abstract

This invention provides a method for optimizing communication protocols for large-scale photovoltaic power plants based on partitioning and layering. This method relates to the field of photovoltaic power plant communication technology. The method involves dividing the photovoltaic power plant into multiple physical partitions and establishing a multi-layer communication hierarchy by acquiring the geographic location and string topology data of photovoltaic modules; collecting communication performance parameters to calculate routing weights and allocate data transmission paths; and using a deep reinforcement learning algorithm to allocate communication time slots to photovoltaic modules. This method can reduce communication latency, improve data transmission reliability, optimize network resource utilization, and achieve adaptive optimization of communication protocols.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to photovoltaic power station communication technology, and in particular to a large-scale photovoltaic power station communication protocol optimization method based on partitioning and layering. Background Art

[0002] Currently, the communication systems of large-scale photovoltaic power plants generally adopt a centralized architecture, with all photovoltaic modules directly connected to a central controller via a communication network. As power plant scale increases, the number of photovoltaic modules increases dramatically, and the communication load and complexity also increase accordingly. PV modules are often installed in complex terrain, with wide geographical distribution and varying communication conditions, making the construction and maintenance of communication networks challenging.

[0003] Existing communication solutions lack sufficient consideration of the physical characteristics of photovoltaic power plants and fail to make reasonable communication network planning based on the geographical distribution of components and the string topology. This leads to an unreasonable network structure and low communication efficiency, especially in large-scale photovoltaic power plants.

[0004] Existing communication protocols use fixed routing and static time slot allocation strategies, which cannot dynamically adjust communication paths and resource allocation according to real-time communication needs and network conditions. When the communication load in certain areas of the photovoltaic power station suddenly increases or the communication quality decreases, it is easy to cause local network congestion or communication interruption.

[0005] The existing communication system lacks an intelligent time slot scheduling mechanism and is unable to perform differentiated processing based on the data importance and communication priority of different photovoltaic modules. Under limited communication resources, it cannot guarantee the timely transmission of key data, affecting the monitoring effect and operational safety of the power station. Summary of the Invention

[0006] The embodiment of the present invention provides a large-scale photovoltaic power station communication protocol optimization method based on partitioning and layering, which can solve the problems in the prior art.

[0007] A first aspect of an embodiment of the present invention provides a method for optimizing a communication protocol of a large photovoltaic power station based on partitioning and layering, comprising:

[0008] Obtain geographic location data and string topology data for PV modules in a PV power station, divide the PV power station into multiple physical partitions and establish multiple communication levels. Set up a regional communication controller in each physical partition to forward data between communication levels.

[0009] Collecting communication performance parameters of each communication layer, calculating routing weight values ​​between each communication layer, allocating data transmission paths for each communication layer, and sending the data transmission paths to regional communication controllers of adjacent communication layers;

[0010] The photovoltaic components in the physical partition are modeled as a communication network graph, a deep reinforcement learning algorithm is used to calculate the communication time slot allocation sequence of each photovoltaic component in the physical partition, and communication time slots are allocated to the photovoltaic components in the physical partition according to the communication time slot allocation sequence;

[0011] The regional communication controller collects photovoltaic module operation data within the communication time slot and forwards the data between the communication levels through the data transmission path.

[0012] In an optional embodiment,

[0013] Collecting communication performance parameters of each communication layer, calculating routing weight values ​​between each communication layer, allocating data transmission paths for each communication layer, and sending the data transmission paths to regional communication controllers of adjacent communication layers includes:

[0014] Communication performance parameters are calculated using a dynamic time window. The current value of the communication performance parameter is obtained by the weighted sum of the historical value and the real-time measurement value, where the weighting coefficient is calculated by the sampling time interval and the preset time constant. The communication performance parameters include data throughput parameter, communication delay parameter, and network congestion parameter. The normalized throughput parameter, delay parameter, and congestion parameter are weighted and combined, and multiplied by the communication layer constraint coefficient to obtain the routing weight. The communication layer constraint coefficient is used to control cross-layer communication.

[0015] The multi-level communication network of the photovoltaic power station is constructed as a directed weighted graph, wherein the nodes of the directed weighted graph correspond to communication nodes, the edges correspond to communication links, and the edge weights correspond to the routing weights; and an improved Dijkstra algorithm is used to search for an optimal path in the directed weighted graph;

[0016] Calculate multiple backup paths with the highest weights and path overlap below a preset overlap threshold, periodically test the connectivity and communication performance of the optimal path, and activate the backup paths in descending order of the product of their routing weights when a link is interrupted or the communication performance falls below a preset change threshold;

[0017] The optimal path and backup path information are sent to a regional communication controller at an adjacent communication level.

[0018] In an optional embodiment,

[0019] Searching for an optimal path in the directed weighted graph using the improved Dijkstra algorithm includes:

[0020] Initialize the distance of the source node to 0 and the distances of other nodes to infinity. Select the unvisited node with the smallest distance as the current node. Update the distance values ​​of the nodes adjacent to the current node and whose level difference is not greater than 1. The distance value is updated by multiplying the routing weights. When all nodes are visited, reconstruct the optimal path based on the distance values ​​of the nodes. The optimal path satisfies that the level difference of adjacent nodes on the path is not greater than 1 and the product of the routing weights of each edge on the path is the largest.

[0021] In an optional embodiment,

[0022] The photovoltaic components in the physical partition are modeled as a communication network graph, and a deep reinforcement learning algorithm is used to calculate a communication time slot allocation sequence for each photovoltaic component in the physical partition. The communication time slot allocation sequence is used to allocate communication time slots to the photovoltaic components in the physical partition, including:

[0023] The photovoltaic components in the physical partition are modeled as a communication network graph, where the nodes of the communication network graph represent the photovoltaic components, and a node feature vector is constructed based on the historical communication data volume, real-time data collection frequency and location encoding information of the photovoltaic components;

[0024] Establishing a causal graph of the communication system within the physical partition, wherein the nodes of the causal graph represent communication state parameters, the edges represent causal relationships between the parameters, and establishing a mapping relationship with the communication network graph;

[0025] Based on the node feature vectors, a multi-layer graph attention network is used to extract features from the communication network graph; based on the mapping relationship, the causal influence strength between the communication state parameters in the causal graph is used to adjust the attention weight calculation of the graph attention network to generate a graph-level feature representation with causal perception capabilities;

[0026] Based on the topological structure of the communication network graph and the historical communication data of nodes, a sliding time window method is used to extract communication behavior features and construct a training sample set;

[0027] A deep reinforcement learning framework is used for training. The graph-level feature representation is input into the policy network to calculate the time slot allocation probability. The priority of the training samples is determined by the temporal difference error and the overall congestion level of the network.

[0028] A greedy strategy is adopted according to the time slot allocation probability to generate a communication time slot allocation sequence to allocate communication time slots to photovoltaic components in the physical partition.

[0029] In an optional embodiment,

[0030] Establishing a causal graph of the communication system within the physical partition, wherein nodes of the causal graph represent communication state parameters, edges represent causal relationships between parameters, and establishing a mapping relationship with the communication network graph includes:

[0031] Acquire communication status parameter data including node level, link level and system level;

[0032] Based on the communication state parameter data, an initial structure of a causal graph is constructed through conditional independence tests and timing information constraints; a nonlinear structural equation is constructed to describe the relationship between each parameter and its direct causal parameter, historical state, and external interference, and the equation parameters are estimated by minimizing the prediction error to obtain a complete causal graph;

[0033] Constructing a bidirectional mapping relationship between communication network graph nodes and causal graph nodes, including: calculating the influence of each communication network node on each communication state parameter based on the nonlinear structural equation to obtain a node-parameter influence matrix; establishing a bidirectional mapping relationship between the communication network graph nodes and the causal graph nodes based on the influence matrix, and the mapping relationship is obtained by calculating a normalized mapping function.

[0034] In an optional embodiment,

[0035] A deep reinforcement learning framework is used for training. The graph-level feature representation is input into the policy network to calculate the time slot allocation probability. The priority of the training samples is determined by the temporal difference error and the overall congestion level of the network.

[0036] Build an Actor-Critic deep reinforcement learning framework, where the Actor network serves as the policy network to generate time slot allocation strategies, and the Critic network is used to evaluate state value;

[0037] The graph-level feature representation is input into the policy network, a first-layer weight matrix and bias vector are linearly transformed to obtain hidden layer features, a nonlinear activation process is performed on the hidden layer features, a second-layer weight matrix and bias vector are linearly transformed to obtain the allocation score of each node in different time slots, and the score is converted into a time slot allocation probability through normalization;

[0038] The node-level communication congestion degree is calculated based on the node communication queue length, data arrival rate and service rate, and the overall network congestion degree is obtained by weighting the node importance weight;

[0039] Calculate the temporal difference error (TDE) including the immediate reward, the current state value, and the next state value, and use the absolute value of the TDE and the weighted combination of the overall network congestion level as the training sample priority;

[0040] Experience replay sampling is performed based on the sample priority, the policy network parameters are updated using the policy gradient method, and the value network parameters are updated using temporal difference learning. During the parameter update process of the policy network and the value network, the importance weights corresponding to the sample priority are used for gradient correction.

[0041] In an optional embodiment,

[0042] Calculating the node-level communication congestion level based on the node communication queue length, data arrival rate and service rate includes:

[0043] The multi-scale congestion characteristics of nodes are calculated using sliding time windows of different sizes, including the instantaneous congestion level in a short-term window, the fluctuation trend in a medium-term window, and the cumulative congestion level in a long-term window.

[0044] Determine the dynamic importance weight of a node based on its degree centrality and communication demand intensity in the communication network;

[0045] The multi-scale congestion features are weighted and combined using the dynamic importance weights to obtain the node-level communication congestion degree.

[0046] According to a second aspect of an embodiment of the present invention, an electronic device is provided, including:

[0047] processor;

[0048] a memory for storing processor-executable instructions;

[0049] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0050] According to a third aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0051] The present invention adopts a partitioned and hierarchical management communication mode for large-scale photovoltaic power stations, thereby achieving effective communication scheduling of large-scale equipment in the photovoltaic power station system, reducing network delay and data transmission load, and improving the overall performance and reliability of the communication system.

[0052] The present invention introduces a physical partitioning mechanism based on geographic location and topological structure and a multi-level communication architecture, combined with communication route optimization and time slot allocation technology, to solve the congestion and communication conflict problems faced by traditional communication methods in large-scale photovoltaic power plants, thereby improving the system communication efficiency.

[0053] The present invention uses a deep reinforcement learning algorithm to intelligently optimize the allocation of communication time slots, so that communication resources are rationally utilized and data transmission conflicts are avoided. At the same time, through the coordination of regional communication controllers, distributed data collection and efficient transmission are realized, providing a reliable data foundation for the operation and maintenance management of photovoltaic power stations. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1This is a flow chart of a method for optimizing communication protocols for large-scale photovoltaic power stations based on partitioning and layering according to an embodiment of the present invention;

[0055] Figure 2 Comparison chart of path recovery time for different methods;

[0056] Figure 3 This is a comparison chart of network throughput of different methods;

[0057] Figure 4 Flowchart of calculation of node-level communication congestion level of the present invention. DETAILED DESCRIPTION

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0059] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0060] Figure 1 Schematic diagram of the process of the embodiment of the present invention, as shown in FIG. Figure 1 As shown, the method includes:

[0061] Obtain geographic location data and string topology data for PV modules in a PV power station, divide the PV power station into multiple physical partitions and establish multiple communication levels. Set up a regional communication controller in each physical partition to forward data between communication levels.

[0062] For example, the GPS coordinate data of each photovoltaic component is obtained from the monitoring system database, such as (119.357621, 31.468935). At the same time, the string topology data is obtained to record the electrical connection relationship between the components. For example, the S001 string in area A contains components P001 to P024. Physical partitioning is performed based on geographical location, and a clustering algorithm is used to group components with a distance of less than 150 meters into one physical partition. For example, a 50MW power station is usually divided into 5-8 physical partitions, and each partition covers an area of ​​approximately 0.8 square kilometers. A communication hierarchy is established based on the geographical distribution characteristics of the physical partitions. The typical configuration is a three-level structure: edge layer (collection equipment), convergence layer (substation controller) and core layer (central controller). Different communication rates are used between adjacent layers, with 60 seconds / cycle from the edge to the convergence layer and 180 seconds / cycle from the convergence to the core layer.

[0063] Collecting communication performance parameters of each communication layer, calculating routing weight values ​​between each communication layer, allocating data transmission paths for each communication layer, and sending the data transmission paths to regional communication controllers of adjacent communication layers;

[0064] The photovoltaic components in the physical partition are modeled as a communication network graph, a deep reinforcement learning algorithm is used to calculate the communication time slot allocation sequence of each photovoltaic component in the physical partition, and communication time slots are allocated to the photovoltaic components in the physical partition according to the communication time slot allocation sequence;

[0065] The regional communication controller collects photovoltaic module operation data within the communication time slot and forwards the data between the communication levels through the data transmission path.

[0066] In an optional embodiment, collecting communication performance parameters of each communication layer, calculating routing weight values ​​between each communication layer, allocating a data transmission path for each communication layer, and sending the data transmission path to a regional communication controller of an adjacent communication layer includes:

[0067] Communication performance parameters are calculated using a dynamic time window. The current value of the communication performance parameter is obtained by the weighted sum of the historical value and the real-time measurement value, where the weighting coefficient is calculated by the sampling time interval and the preset time constant. The communication performance parameters include data throughput parameter, communication delay parameter, and network congestion parameter. The normalized throughput parameter, delay parameter, and congestion parameter are weighted and combined, and multiplied by the communication layer constraint coefficient to obtain the routing weight. The communication layer constraint coefficient is used to control cross-layer communication.

[0068] The multi-level communication network of the photovoltaic power station is constructed as a directed weighted graph, wherein the nodes of the directed weighted graph correspond to communication nodes, the edges correspond to communication links, and the edge weights correspond to the routing weights; and an improved Dijkstra algorithm is used to search for an optimal path in the directed weighted graph;

[0069] Calculate multiple backup paths with the highest weights and path overlap below a preset overlap threshold, periodically test the connectivity and communication performance of the optimal path, and activate the backup paths in descending order of the product of their routing weights when a link is interrupted or the communication performance falls below a preset change threshold;

[0070] The optimal path and backup path information are sent to a regional communication controller at an adjacent communication level.

[0071] For example, in the process of collecting communication performance parameters, the system adopts a dynamic time window mechanism, sets the preset time constant to 300 seconds, and calculates the weighting coefficient each time a sample is taken. For example, when the sampling interval is 30 seconds, the weighting coefficient is calculated to be 0.1, which means that the current measurement value accounts for 10% and the historical value accounts for 90%. The collected communication performance parameters are processed separately according to their types. The data throughput parameter collection cycle is 60 seconds, recording the amount of data successfully transmitted per unit time in kilobytes per second; the communication delay parameter collection cycle is 30 seconds, recording the round-trip time of the data packet from the source node to the destination node in milliseconds; the network congestion parameter collection cycle is 120 seconds, and is comprehensively evaluated by queue length and packet loss rate, with a value range of 0 to 100.

[0072] Performance parameter normalization uses a linear mapping method. For throughput normalization, the raw value is divided by the maximum link bandwidth. For example, if a 2.5 Mbps link measures 1.5 Mbps, the normalized value is 0.6. For latency normalization, an inverse proportional transformation is used, converting a 50 ms delay to a performance index of 0.8. For congestion normalization, the raw value is directly divided by 100. For example, a congestion index of 40 is converted to 0.4. The normalized parameters are weighted: throughput weighted 0.4, latency weighted 0.35, and congestion weighted 0.25.

[0073] The communication layer constraint coefficient is designed to control cross-layer communication. For example, the constraint coefficient for communication within the same layer can be set to 1.0, the constraint coefficient for communication between adjacent layers can be set to 0.8, the constraint coefficient for communication across two layers can be set to 0.5, and the constraint coefficient for communication across three or more layers can be set to 0.2. The final calculation result of the routing weight is accurate to two decimal places.

[0074] When constructing a multi-level communication network for a PV power plant as a directed weighted graph, the communication topology is first modeled. Each communication node is treated as a vertex in the graph and assigned a unique identifier and hierarchical attributes. For example, edge layer nodes are identified as E001 to E350, aggregation layer nodes are identified as A001 to A040, and core layer nodes are identified as C001 to C005. Each node stores information such as its physical location coordinates, the physical partition number to which it belongs, processing capacity parameters, and communication interface type. Communication links are modeled as directed edges in the graph. Each edge contains a source node ID, a destination node ID, and a link type. Link types are categorized as wired links (such as Ethernet and RS485 buses) and wireless links (such as Zigbee and LoRa networks). The system constructs a complete connection topology based on the actual plant deployment. A typical 50MW PV power plant communication network includes approximately 400 nodes and 700-800 communication links. Edge weights are directly mapped to the routing weights calculated above. Weight values ​​range from 0 to 1 and are accurate to three decimal places. Higher weights indicate better communication link quality. The system uses an adjacency list structure to store graph models in memory. Each node maintains a list of outgoing edges, whose elements include the target node ID, edge weight, and link attributes. Furthermore, to accelerate path search, the system establishes node indexes by communication level, enabling fast node lookup within that level.

[0075] Use the improved Dijkstra algorithm to search for the optimal path.

[0076] When calculating backup paths, a multipath search strategy is used. After calculating the optimal path, one link in the optimal path is temporarily removed and the modified Dijkstra algorithm is re-executed to obtain the first backup path. This link is then restored, another link is removed, and the calculation continues until a preset number of backup paths are obtained. Typically, three backup paths are configured.

[0077] Path overlap is assessed using the shared link ratio method. For two paths, the overlap is calculated by dividing the number of links they share by the total number of links in the shorter path. For example, if path A contains links [1, 2, 3, 4, 5] and path B contains links [1, 2, 6, 7, 5], the shared links are [1, 2, 5], and the overlap is 3 / 5 = 0.6. The system sets a preset overlap threshold of 0.7. Backup paths that exceed this threshold with the optimal path or an existing backup path are excluded.

[0078] The connectivity and communication performance of the optimal path are periodically tested every 180 seconds. Link connectivity is checked using the ICMP protocol, and communication performance is evaluated using real-time performance parameters. If a link interruption or a performance drop of more than 30% is detected, a path switching mechanism is triggered. Path switching attempts to find alternative paths in descending order of their routing weight products until a viable path is found.

[0079] When sending optimal and backup path information to the regional communication controller at the adjacent communication level, a compact path encoding format is used. Each path is encoded as a sequence of node IDs and corresponding link performance indicators, encapsulated in JSON format.

[0080] Figure 2 This is a comparison chart of the path recovery time of different methods, showing the comparison of the path recovery time of three different communication routing methods under six fault scenarios. The horizontal axis represents the fault scenario number (1-6), and the vertical axis represents the path recovery time (milliseconds). The method of the present invention (solid line + circular mark) shows extremely low path recovery time in all fault scenarios, ranging from 20 to 40 milliseconds. The single backup path method (dashed line + square mark) performs second best, with a recovery time between 40 and 65 milliseconds, which is about 1.5 to 2 times higher than the method of the present invention. The recalculation path method (dotted line + inverted triangle mark, which refers to a technical method that does not use a pre-calculated backup path when a link fails, but re-executes the complete Dijkstra algorithm in real time to calculate a new optimal path) performs the worst, with a path recovery time between 750 and 920 milliseconds, which is about 25 to 30 times higher than the method of the present invention. This method needs to completely traverse the entire network topology for path search every time a fault occurs, so the recovery time is longer, but it can obtain the optimal path under the current network status. In contrast, the present invention adopts multiple pre-calculated backup paths and takes into account overlap constraints, which can quickly switch when a failure occurs and significantly reduce recovery delay.

[0081] The present invention calculates communication performance parameters through a dynamic time window, combines a hierarchically constrained routing weight calculation mechanism with an improved path search algorithm, and achieves efficient routing optimization of multi-level communication networks in photovoltaic power plants. The introduction of backup paths and a dynamic switching mechanism significantly improves the reliability and adaptability of the communication system. Through rationally designed routing weight calculation and path search strategies, the system can effectively avoid network congestion and single point failures while ensuring data transmission efficiency, providing reliable communication support for the stable operation of photovoltaic power plants and real-time data monitoring.

[0082] In an optional embodiment, searching for an optimal path in the directed weighted graph using an improved Dijkstra algorithm includes:

[0083] Initialize the distance of the source node to 0 and the distances of other nodes to infinity. Select the unvisited node with the smallest distance as the current node. Update the distance values ​​of the nodes adjacent to the current node and whose level difference is not greater than 1. The distance value is updated by multiplying the routing weights. When all nodes are visited, reconstruct the optimal path based on the distance values ​​of the nodes. The optimal path satisfies that the level difference of adjacent nodes on the path is not greater than 1 and the product of the routing weights of each edge on the path is the largest.

[0084] For example, during the algorithm initialization phase, a distance array and an access flag array are first created. Unlike the traditional Dijkstra algorithm, this improved algorithm initializes the distance of the source node to 0 (indicating the optimal distance) and the distances of all other nodes to infinity. For example, for a communication network consisting of nodes N001 (source node), N002, N003, etc., the initial distance value of N001 is 0, and the distance values ​​of the remaining nodes are set to the system maximum value (e.g., 999999).

[0085] Create a priority queue for node selection and initially add source node N001 to the queue. The priority queue is sorted by distance value from smallest to largest, ensuring that each node removed from the queue is the unvisited node with the smallest distance.

[0086] During the main loop, each time the unvisited node with the smallest distance value is taken from the priority queue as the current processing node. For example, initially there is only node N001 in the queue with a distance value of 0, so N001 is selected as the current node and marked as visited.

[0087] For each adjacent node of the current node, check the level constraint. Obtain the level attributes of the current node and the adjacent node, and calculate the absolute value of the difference between the two. Only proceed if the level difference is no greater than 1. For example, if N001 is at level 2 and the adjacent node N002 is at level 3, the level difference is 1, thus satisfying the constraint. If the adjacent node N003 is at level 0 and the level difference is 2, then processing is skipped.

[0088] For adjacent nodes that meet the hierarchical constraints, the distance value update calculation is performed. The product operation of the routing weight is used here instead of the traditional addition operation. The specific implementation is: add the distance value of the current node to the inverse of the routing weight of the connecting edge to obtain a new distance candidate value. Since the routing weight value range is 0-1, the higher the weight, the better the link quality, and the Dijkstra algorithm is looking for the shortest path, the inverse of the weight (1 / weight) is used as the distance increment. For example, the distance value of the current node N001 is 0, and the routing weight of the edge (N001, N002) is 0.8, then the distance increment is 1 / 0.8=1.25, and the new distance candidate value of N002 is 0+1.25=1.25.

[0089] If the calculated new distance value is less than the current distance value of the adjacent node, the distance value of the adjacent node and the predecessor node pointer are updated. For example, if the current distance value of N002 is infinite, it is updated to 1.25 and its predecessor node is recorded as N001.

[0090] After completing the adjacent node update, all updated and unvisited adjacent nodes are added to the priority queue. The above process is repeated until all nodes are visited or no further processing can be done.

[0091] When the algorithm is complete, the distance value for each node represents the cumulative sum of the inverse weights of the optimal path from the source node to that node. The smaller the distance value, the greater the product of the routing weights of the edges on the path.

[0092] In the path reconstruction phase, starting from the target node, the complete optimal path is constructed by tracing back through the predecessor node pointers. For example, if the predecessor node of target node N010 is N007, the predecessor node of N007 is N005, and the predecessor node of N005 is N001, then the optimal path is N001→N005→N007→N010.

[0093] The resulting optimal path satisfies two key conditions: the difference in level between adjacent nodes on the path is no greater than 1, and the product of the routing weights of each edge on the path is maximized. For example, on the path N001→N005→N007→N010, the node levels are 2, 3, 3, and 2, respectively, and the level difference between adjacent nodes does not exceed 1. The routing weights of each edge are 0.9, 0.85, and 0.8, respectively, and the product is 0.612, the maximum value among all possible paths.

[0094] By improving the traditional Dijkstra algorithm, the present invention innovatively introduces a hierarchical constraint check and weight product calculation mechanism, effectively solving the path optimization problem in the multi-level communication network of photovoltaic power plants. The improved algorithm can find the communication path with the largest routing weight product while ensuring reasonable jumps in the communication level, significantly improving data transmission efficiency and reliability. Especially when the network topology changes dynamically, the algorithm can respond quickly and generate the optimal path, providing stable and reliable communication guarantee for real-time monitoring and intelligent operation and maintenance of photovoltaic power plants.

[0095] In an optional embodiment, the photovoltaic components within the physical partition are modeled as a communication network graph, and a deep reinforcement learning algorithm is used to calculate a communication time slot allocation sequence for each photovoltaic component within the physical partition. The communication time slot allocation sequence is used to allocate communication time slots to the photovoltaic components within the physical partition, including:

[0096] The photovoltaic components in the physical partition are modeled as a communication network graph, where the nodes of the communication network graph represent the photovoltaic components, and a node feature vector is constructed based on the historical communication data volume, real-time data collection frequency and location encoding information of the photovoltaic components;

[0097] Establishing a causal graph of the communication system within the physical partition, wherein the nodes of the causal graph represent communication state parameters, the edges represent causal relationships between the parameters, and establishing a mapping relationship with the communication network graph;

[0098] Based on the node feature vectors, a multi-layer graph attention network is used to extract features from the communication network graph; based on the mapping relationship, the causal influence strength between the communication state parameters in the causal graph is used to adjust the attention weight calculation of the graph attention network to generate a graph-level feature representation with causal perception capabilities;

[0099] Based on the topological structure of the communication network graph and the historical communication data of nodes, a sliding time window method is used to extract communication behavior features and construct a training sample set;

[0100] A deep reinforcement learning framework is used for training. The graph-level feature representation is input into the policy network to calculate the time slot allocation probability. The priority of the training samples is determined by the temporal difference error and the overall congestion level of the network.

[0101] A greedy strategy is adopted according to the time slot allocation probability to generate a communication time slot allocation sequence to allocate communication time slots to photovoltaic components in the physical partition.

[0102] For example, modeling the PV panels within a physical partition as a communication network graph is the first step in achieving efficient communication time slot allocation. Taking physical partition A within a 50MW PV power plant as an example, this partition contains 2,500 PV panels. A communication network graph is constructed based on the component deployment locations and communication link relationships. Each PV panel corresponds to a node in the graph, and the communication links between panels correspond to the edges between nodes. A feature vector is constructed for each node, containing the following information: historical communication data volume (average communication volume over the past 24 hours, in KB / hour), real-time data collection frequency (number of collections per minute), and location encoding information (converting the component's two-dimensional coordinates into a one-dimensional code). For example, the feature vector for component PV-A-0123 is [256.5, 6, 0.372], indicating that the component generated an average of 256.5KB of data per hour over the past 24 hours, with a current collection frequency of 6 times per minute, and a location encoding value of 0.372.

[0103] When establishing a cause-and-effect graph for the communication system within a physical partition, key communication status parameters are defined as nodes in the graph, including network throughput, link occupancy, packet loss rate, communication latency, queue length, and channel quality. When establishing a mapping relationship between the communication network graph and the cause-and-effect graph, each PV module node is mapped to its corresponding multiple communication status parameters.

[0104] A three-layered graph attention network (GAN) is used to extract features from communication network graphs. The input layer receives node feature vectors, the intermediate layers perform graph convolution and attention mechanisms, and the output layer generates high-dimensional representations of the nodes. The causal strength value in the causal graph is used as an adjustment factor in the calculation of attention weights. When calculating the attention weight between nodes i and j, the causal strength value between the communication state parameters mapped to these two nodes is retrieved and multiplied by the original attention score. For example, if nodes i and j are mapped to "queue length" and "communication delay," respectively, and the causal strength between these two parameters is 0.85, the original attention score is multiplied by 0.85 to obtain the adjusted attention weight. This adjustment gives node pairs with strong causal relationships higher weights during feature extraction, generating a causally-aware graph-level feature representation.

[0105] To construct the training sample set, a sliding time window method was used to extract communication behavior features. The window size was set to 30 minutes, with a step size of 5 minutes. For each time window, metrics such as the communication data volume, activity, and communication success rate of each component were recorded, along with the overall network congestion level. For example, within the window [10:00-10:30], the communication data volume of component PV-A-0123 was 150KB, the activity was 85%, the communication success rate was 98.5%, and the overall network congestion level was 28%. Using a sliding window, a large number of training samples were generated from historical data. Each sample contained input features (component communication behavior and network status) and an output label (optimal time slot allocation scheme). Training was performed using a deep reinforcement learning framework.

[0106] Traditional methods usually use fixed time slot allocation or priority-based dynamic allocation, which is difficult to adapt to the complex and changing communication needs of photovoltaic power plants. Figure 3This graph compares network throughput across different methods. The horizontal axis represents the number of PV modules, ranging from 0 to 3000; the vertical axis represents network throughput (Mbps), ranging from 0 to 70 Mbps. The three algorithms compared include: the causal-aware deep reinforcement learning algorithm of our invention (solid line + circular markers), the traditional DRL algorithm (dashed line + square markers), and the fixed time slot allocation algorithm (dash-dotted line + triangle markers). As the graph shows, the network throughput of all three algorithms increases with the number of PV modules, but the algorithm of our invention performs best at all node scales. At the maximum scale of 3000 modules, the algorithm of our invention achieves a network throughput of 31.5 Mbps, approximately 23% higher than the traditional DRL algorithm (25.6 Mbps) and approximately 94% higher than the fixed time slot allocation algorithm (16.2 Mbps), demonstrating the significant advantages of causal-aware features in large-scale PV module communication systems. In particular, in large-scale networks of 2500-3000 nodes, the throughput growth curve of our algorithm maintains a high slope, demonstrating its good scalability and suitability for deployment in expanding PV power plant communication systems.

[0107] This paper introduces a method that combines causal graphs with deep reinforcement learning. By establishing a causal model of the communication system, causal knowledge is integrated into the feature extraction process, enabling the learning algorithm to understand the inherent relationships between communication parameters. By modeling photovoltaic components within a physical partition as a communication network graph and combining the causal graph with a deep reinforcement learning algorithm, the paper achieves efficient communication time slot allocation. This approach not only considers the topological structure and historical communication behavior of the communication network, but also incorporates knowledge of the causal relationships between communication parameters, significantly improving the accuracy and adaptability of time slot allocation.

[0108] In an optional embodiment, establishing a causal graph of the communication system within the physical partition, where nodes of the causal graph represent communication state parameters and edges represent causal relationships between the parameters, and establishing a mapping relationship with the communication network graph includes:

[0109] Acquire communication status parameter data including node level, link level and system level;

[0110] Based on the communication state parameter data, an initial structure of a causal graph is constructed through conditional independence tests and timing information constraints; a nonlinear structural equation is constructed to describe the relationship between each parameter and its direct causal parameter, historical state, and external interference, and the equation parameters are estimated by minimizing the prediction error to obtain a complete causal graph;

[0111] Constructing a bidirectional mapping relationship between communication network graph nodes and causal graph nodes, including: calculating the influence of each communication network node on each communication state parameter based on the nonlinear structural equation to obtain a node-parameter influence matrix; establishing a bidirectional mapping relationship between the communication network graph nodes and the causal graph nodes based on the influence matrix, and the mapping relationship is obtained by calculating a normalized mapping function.

[0112] For example, when acquiring communication status parameter data, information is collected at three levels: node-level parameters include the data generation rate, cache utilization, and power status of each photovoltaic module; link-level parameters include channel quality, link occupancy, signal strength, packet loss rate, and transmission delay; and system-level parameters include network throughput, global congestion, and scheduling efficiency. These parameters are collected every five minutes via a distributed sensor network and stored in a time series database. For example, in physical partition A of a 50MW photovoltaic power plant, node PV-A-0128 has a data generation rate of 45KB / minute and a cache utilization rate of 67%; the link occupancy between it and the gateway is 72%, with a packet loss rate of 2.3%; and the system-level congestion is 48%.

[0113] When constructing the initial structure of a causal graph based on collected parameter data, a conditional independence test is applied to identify causal relationships between parameters. This test determines whether a direct causal relationship exists between parameters X and Y by calculating the conditional correlation between them while controlling for parameter Z. If the conditional correlation decreases significantly, it indicates that Z is a mediating variable in the relationship between X and Y. For example, when testing the conditional correlation between "data generation rate" and "network throughput" while controlling for "link utilization," the correlation coefficient dropped from 0.86 to 0.21, indicating that "link utilization" is a mediating variable in the relationship. To avoid ambiguity in the direction of causal relationships, timing constraints are introduced. The causal direction is determined by analyzing the sequence of parameter changes. For example, an increase in "link utilization" typically precedes an increase in "packet loss rate" by approximately 30 seconds. Based on this, the causal direction is determined to be from "link utilization" to "packet loss rate." Combining the results of the conditional independence test and the timing constraints, the initial structure of the causal graph is constructed, consisting of 15 communication state parameter nodes and 23 causal edges.

[0114] When constructing nonlinear structural equations, an additive noise model is used to describe the relationship between each parameter and its direct causal parameters, historical state, and external interference. For parameter Y, its structural equation is expressed as a function of the current values ​​of its direct causal parameters X1, X2, ...Xn, Y's own historical values, and random noise. This function takes the form of a piecewise linear function, using different linear combination coefficients for different parameter value ranges. For example, the structural equation for the "packet loss rate" parameter includes "link occupancy" and "signal strength" as causal inputs. When "link occupancy" is below 60%, its influence coefficient is 0.3; when it exceeds 60%, the influence coefficient increases sharply to 0.8, reflecting the strong impact of high occupancy on packet loss. To estimate the equation parameters by minimizing the prediction error, a batch gradient descent algorithm is used, using the past week's historical data as the training set. All causal strength parameters are initialized to random values ​​(between 0 and 1), and then they are iteratively updated to minimize the mean squared error between the model predictions and the actual observations. For example, after 1000 iterations, the causal strength between "link utilization" and "packet loss rate" converges to 0.72, indicating a strong causal relationship; whereas the causal strength between "battery status" and "packet loss rate" is only 0.08, indicating a weak causal relationship. This ultimately results in a complete causal graph, with each edge labeled with its corresponding causal strength value.

[0115] When constructing a bidirectional mapping between nodes in the communication network graph and nodes in the causal graph, the nonlinear structural equation is used to calculate the influence of each communication network node on each communication state parameter. While holding the parameters of other nodes constant, the parameters of the target node are slightly perturbed (increased or decreased by 5%). The magnitude of change in each communication state parameter is observed and used as a quantitative indicator of the degree of influence. For example, perturbing the data generation rate of node PV-A-0128 results in a 3.2% change in the "link occupancy rate" parameter and a 1.7% change in the "system congestion level." Based on this, the node's influence on these two parameters is quantified to 0.64 and 0.34, respectively. This process is repeated for all nodes to obtain a complete node-parameter influence matrix.

[0116] Establishing a bidirectional mapping relationship between nodes in the communication network graph and nodes in the causal graph based on the influence matrix quantifies the influence of a node on a parameter into a specific mapping strength value. This step converts the previously calculated node-parameter influence matrix into a standardized bidirectional mapping relationship. The mapping relationship is calculated using a normalized mapping function: for each node i in the communication network graph and each node j (communication state parameter) in the causal graph, the forward mapping strength M_forward(i, j) is equal to the influence of node i on parameter j divided by the sum of the influences of all nodes on parameter j. For example, if the influence of the photovoltaic module PV-A-0128 on the "link occupancy" parameter is 0.64, and the sum of the influences of all modules on this parameter is 27.8, then the mapping strength from this module to "link occupancy" is 0.64 / 27.8 = 0.023. Similarly, the reverse mapping strength M_backward(j, i) is equal to the influence of parameter j on the communication behavior of node i divided by the sum of the influences of all parameters on node i. For example, the impact of the "Packet Loss Rate" parameter on component PV-A-0128 is 0.38, and the total impact of all parameters on this component is 5.07. Therefore, the mapping strength of "Packet Loss Rate" to this component is 0.38 / 5.07 = 0.075. This normalization ensures the comparability and rationality of the mapping strengths. The mapping strength value ranges from 0 to 1, and for each parameter or node, the sum of all mapping strengths is 1. The final bidirectional mapping relationship is stored as two matrices, providing critical causal knowledge support in subsequent feature extraction and decision-making.

[0117] The present invention constructs a causal graph of the communication system and its mapping relationship with the communication network through a systematic approach, revealing the inherent causal mechanism between parameters in the communication network of a complex photovoltaic power station. It not only identifies the direct and indirect causal relationships between key communication status parameters, but also quantifies the degree of mutual influence between photovoltaic components and communication status parameters, providing a causal knowledge basis for subsequent communication time slot allocation. It improves the accuracy and adaptability of communication resource allocation, fundamentally improves communication efficiency and reliability, and provides solid support for the efficient operation and management of photovoltaic power stations.

[0118] In an optional embodiment, a deep reinforcement learning framework is used for training, the graph-level feature representation is input into the policy network to calculate the time slot allocation probability, and the priority of the training sample is determined by the temporal difference error and the overall congestion level of the network, including:

[0119] Build an Actor-Critic deep reinforcement learning framework, where the Actor network serves as the policy network to generate time slot allocation strategies, and the Critic network is used to evaluate state value;

[0120] The graph-level feature representation is input into the policy network, a first-layer weight matrix and bias vector are linearly transformed to obtain hidden layer features, a nonlinear activation process is performed on the hidden layer features, a second-layer weight matrix and bias vector are linearly transformed to obtain the allocation score of each node in different time slots, and the score is converted into a time slot allocation probability through normalization;

[0121] The node-level communication congestion degree is calculated based on the node communication queue length, data arrival rate and service rate, and the overall network congestion degree is obtained by weighting the node importance weight;

[0122] Calculate the temporal difference error (TDE) including the immediate reward, the current state value, and the next state value, and use the absolute value of the TDE and the weighted combination of the overall network congestion level as the training sample priority;

[0123] Experience replay sampling is performed based on the sample priority, the policy network parameters are updated using the policy gradient method, and the value network parameters are updated using temporal difference learning. During the parameter update process of the policy network and the value network, the importance weights corresponding to the sample priority are used for gradient correction.

[0124] For example, the Actor network consists of a three-layer fully connected neural network: the input layer has the same dimensions as the graph-level feature representation (e.g., 128 dimensions), the hidden layer contains 256 neurons, and the output layer has the dimensions of the number of nodes multiplied by the number of time slots. The Critic network has a similar structure, but the output layer contains only one neuron, which outputs the state value estimate. For example, for a scenario with 2,500 photovoltaic panels and a total of 128 time slots, the output layer of the Actor network has dimensions of 320,000 (2,500 × 128), representing the assigned score of each panel in each time slot.

[0125] The process of inputting the graph-level feature representation into the policy network for forward computation involves multiple transformation steps: The input features are linearly transformed using the first-layer weight matrix and bias vector to obtain hidden features. The 128-dimensional input features are multiplied by the 256×128-dimensional weight matrix and then added to the 256-dimensional bias vector to obtain 256-dimensional hidden features. For example, if an input feature element has a value of 0.75 and a corresponding weight of 0.5, its contribution is 0.375. Accumulating the contributions of all input elements and adding a bias of 0.2 yields a hidden feature element value of 6.8. The hidden features are nonlinearly activated using the ReLU function, which sets all negative values ​​to 0 while leaving positive values ​​unchanged. For example, an element with a value of -2.3 in the hidden feature becomes 0 after ReLU processing, while an element with a value of 6.8 remains unchanged. The activated features are then linearly transformed using the second-layer weight matrix (320,000×256 dimensions) and bias vector (320,000 dimensions) to obtain the allocation score for each node at different time slots. For example, the allocation score for node PV-A-0128 in time slot 37 is 8.5, and the allocation score for time slot 58 is 5.2. The scores are converted into time slot allocation probabilities through normalization. A Softmax function is applied to all time slot scores for each node so that the sum of all time slot allocation probabilities for each node is 1. For example, after Softmax processing of the scores for node PV-A-0128 across 128 time slots, the allocation probability for time slot 37 is 0.12, the allocation probability for time slot 58 is 0.06, and the allocation probabilities for the remaining time slots range from 0.001 to 0.08.

[0126] The node-level communication congestion degree is calculated based on the node communication queue length, data arrival rate and service rate.

[0127] When calculating the temporal difference error (TDE), the immediate reward, current state value, and next state value are obtained. The immediate reward is calculated based on a weighted combination of communication throughput, average latency, and fairness metrics, with weights of 0.5, 0.3, and 0.2, respectively. The state value is obtained through the critic network. For example, if the immediate reward of a training example is 0.65, the current state value is 1.2, the discount factor is 0.95, and the next state value is 1.3, the calculated TDE is 0.65 + 0.95 × 1.3 - 1.2 = 0.685.

[0128] The training sample priority is calculated as a weighted combination of the absolute value of the TDE and the overall network congestion level: the priority is equal to the absolute value of the TDE multiplied by 0.7 plus the overall network congestion level multiplied by 0.3. For example, if the absolute value of the TDE is 0.685 and the overall network congestion level is 0.58, the calculated sample priority is 0.685 × 0.7 + 0.58 × 0.3 = 0.6535. A higher priority value indicates more important information contained in the sample and a greater probability of being selected during training.

[0129] When replaying experience based on sample priority, a prioritized experience replay mechanism is used. A 10,000-point experience buffer is maintained, with each experience entry containing the current state, action performed, reward received, next state, and sample priority. The sampling probability is proportional to the sample priority, but a base probability of 0.1 is introduced to ensure that all samples have a chance of being selected. For example, a sample with a priority of 0.6535 has a probability of being selected of approximately 0.01, which is five times that of random sampling.

[0130] When updating policy network parameters using the policy gradient method, the policy gradient is calculated and multiplied by the advantage function (the immediate reward plus the discounted next state value minus the current state value) as the gradient direction. The learning rate is set to 0.001, and 256 samples are sampled from the experience buffer at a time for batch updates. For example, if the current value of a policy network parameter is 0.45, the calculated gradient is 0.08, and the advantage function value is 0.685, the updated parameter value is 0.45 + 0.001 × 0.08 × 0.685 = 0.4504.

[0131] When using temporal difference learning to update the value network parameters, the goal is to minimize the square of the temporal difference error. The learning rate is set to 0.002, and a batch size of 256 samples is used. For example, if the current value of a value network parameter is 0.32, the corresponding gradient is -0.15, and the temporal difference error is 0.685, the updated parameter value is 0.32-0.002×(-0.15)×0.685=0.3202.

[0132] During the parameter update process, gradient correction is performed using importance weights corresponding to sample priorities to compensate for bias caused by non-uniform sampling. Importance weights are equal to the inverse of the probability of a sample being selected and are normalized. For example, a sample with a priority of 0.6535 has an importance weight of approximately 0.2, indicating that its gradient contribution needs to be reduced accordingly.

[0133] By combining the Actor-Critic deep reinforcement learning framework with a priority sample training mechanism, the present invention effectively solves the problem of time slot allocation optimization in photovoltaic power station communication networks. Taking into account the topology and traffic characteristics of the communication network, the present invention introduces the network congestion status as a priority factor, significantly improving the training efficiency and strategy quality. Through a comprehensive evaluation mechanism of temporal difference error and network congestion level, the system can prioritize learning decision-making strategies for important or difficult scenarios and quickly adapt to changes in the communication environment.

[0134] In an optional embodiment, calculating the node-level communication congestion level based on the node communication queue length, data arrival rate, and service rate includes:

[0135] The multi-scale congestion characteristics of nodes are calculated using sliding time windows of different sizes, including the instantaneous congestion level in a short-term window, the fluctuation trend in a medium-term window, and the cumulative congestion level in a long-term window.

[0136] Determine the dynamic importance weight of a node based on its degree centrality and communication demand intensity in the communication network;

[0137] The multi-scale congestion features are weighted and combined using the dynamic importance weights to obtain the node-level communication congestion degree.

[0138] For example, combined Figure 4 The present invention's node-level communication congestion calculation flow chart illustrates this: When calculating a node's multi-scale congestion characteristics using sliding time windows of varying sizes, the system simultaneously maintains three time windows: a short window (30 seconds) for capturing instantaneous congestion states, a medium window (5 minutes) for identifying fluctuation trends, and a long window (30 minutes) for assessing cumulative congestion levels. Within each window, data on the node's communication queue length, data arrival rate, and service rate are collected.

[0139] To calculate the instantaneous congestion level within a short window, a queue theory model is used: the queue utilization is calculated by dividing the current queue length by the queue buffer capacity, and the traffic intensity is calculated by dividing the data arrival rate by the service rate. The weighted sum of the queue utilization and traffic intensity is then calculated, with weights of 0.6 and 0.4, respectively. For example, the average queue length of node PV-A-0128 over the last 30 seconds is 256 KB, the buffer capacity is 512 KB, and the queue utilization is 0.5; the average data arrival rate is 45 KB / minute, the average service rate is 60 KB / minute, and the traffic intensity is 0.75. The calculated instantaneous congestion level is 0.5 × 0.6 + 0.75 × 0.4 = 0.6.

[0140] When calculating the fluctuation trend within the medium-time window, the 5-minute window is divided into 10 30-second subwindows. The instantaneous congestion level is calculated for each subwindow, and then trend characteristics are extracted: the slope of the congestion level is calculated using linear regression, and the fluctuation amplitude (the difference between the maximum and minimum values) is calculated. For example, the instantaneous congestion level sequence of node PV-A-0128 within the 10 subwindows is [0.45, 0.52, 0.58, 0.63, 0.60, 0.65, 0.68, 0.72, 0.69, 0.64]. The calculated slope is 0.025 (an upward trend) and the fluctuation amplitude is 0.27. The fluctuation trend indicator is derived by weighting the normalized slope (ranging from -1 to 1) and the normalized fluctuation amplitude (ranging from 0 to 1), with weights of 0.7 and 0.3, respectively, resulting in a value of 0.6 × 0.7 + 0.27 × 0.3 = 0.5010.

[0141] When calculating the cumulative congestion level within a long-term window, an exponentially weighted method is used to process historical data within 30 minutes, with more recent data given a higher weight. The 30-minute window is divided into six 5-minute segments, and the average congestion level for each segment is calculated. The weight coefficients are [0.4, 0.25, 0.15, 0.1, 0.06, 0.04] from the most recent segment to the most distant segment. For example, the average congestion level for node PV-A-0128 over the six time segments is [0.65, 0.58, 0.52, 0.49, 0.45, 0.43]. The calculated cumulative congestion level is 0.65 × 0.4 + 0.58 × 0.25 + 0.52 × 0.15 + 0.49 × 0.1 + 0.45 × 0.06 + 0.43 × 0.04 = 0.5767.

[0142] To determine a node's dynamic importance weight based on its degree centrality and communication demand intensity in a communication network, a node's degree centrality is calculated as the number of directly connected nodes divided by the maximum possible number of connections in the network. For example, node PV-A-0128 is directly connected to 32 other nodes, and the maximum possible number of connections in the network is 2499. The calculated degree centrality is 32 / 2499 = 0.0128. Communication demand intensity is calculated by combining the node's historical communication data volume and current data generation rate. The communication demand indicator is the weighted sum of the node's average communication data volume over the past 24 hours and its current data generation rate, with weights of 0.4 and 0.6, respectively. For example, if node PV-A-0128 generated an average of 220KB of data per hour over the past 24 hours and a current data generation rate of 45KB / minute (i.e., 2700KB / hour), the calculated communication demand intensity is 220 × 0.4 + 2700 × 0.6 = 1708. This value is normalized by dividing it by the maximum value of the communication demand intensity of all nodes in the entire network, 2500, and the standardized communication demand intensity is 1708 / 2500=0.6832.

[0143] A node's dynamic importance weight is calculated as the weighted sum of its degree centrality and normalized communication demand intensity, with weights of 0.3 and 0.7, respectively. For example, node PV-A-0128 has a degree centrality of 0.0128 and a normalized communication demand intensity of 0.6832. The calculated dynamic importance weight is 0.0128 × 0.3 + 0.6832 × 0.7 = 0.4821. This weight represents the relative importance of the node in the network, with higher values ​​indicating a greater impact on network congestion.

[0144] When combining multi-scale congestion features using dynamic importance weights to calculate node-level communication congestion, the short-term instantaneous congestion level is weighted 0.5, the medium-term fluctuation trend is weighted 0.3, and the long-term cumulative congestion level is weighted 0.2. For example, the instantaneous congestion level of node PV-A-0128 is 0.6, the fluctuation trend is 0.5010, and the cumulative congestion level is 0.5767. The calculated comprehensive congestion index is 0.6 × 0.5 + 0.5010 × 0.3 + 0.5767 × 0.2 = 0.5706.

[0145] Multiplying a node's comprehensive congestion index by its dynamic importance weight yields the node-level communication congestion level. For example, node PV-A-0128 has a comprehensive congestion index of 0.5706 and a dynamic importance weight of 0.4821. The calculated node-level communication congestion level is 0.5706 × 0.4821 = 0.2751. This value reflects the severity of the node's congestion in the current network state and its impact on overall network performance, serving as an important input for subsequent sample priority calculations.

[0146] The present invention achieves accurate characterization of the node communication congestion status through multi-scale congestion analysis and a dynamic importance weight mechanism. This multi-dimensional congestion measurement method improves the system's perception of network status changes, provides a more accurate priority basis for deep reinforcement learning training, effectively guides the intelligent allocation of communication resources, and improves the overall performance and reliability of photovoltaic power station communication networks.

[0147] According to a second aspect of an embodiment of the present invention, an electronic device is provided, including:

[0148] processor;

[0149] a memory for storing processor-executable instructions;

[0150] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0151] According to a third aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0152] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A communication protocol optimization method for a large photovoltaic power station based on partitioning and layering, characterized in that: include: Obtain geographic location data and string topology data for PV modules in a PV power station, divide the PV power station into multiple physical partitions and establish multiple communication hierarchies. Set up a regional communication controller in each physical partition to forward data between communication hierarchies. Collecting communication performance parameters of each communication layer, calculating routing weight values ​​between each communication layer, allocating data transmission paths for each communication layer, and sending the data transmission paths to regional communication controllers of adjacent communication layers; The photovoltaic components in the physical partition are modeled as a communication network graph, and a deep reinforcement learning algorithm is used to calculate the communication time slot allocation sequence of each photovoltaic component in the physical partition. Communication time slots are allocated to the photovoltaic components in the physical partition according to the communication time slot allocation sequence, including: modeling the photovoltaic components in the physical partition as a communication network graph, the nodes of the communication network graph represent the photovoltaic components, and constructing node feature vectors according to the historical communication data volume, real-time data acquisition frequency and position coding information of the photovoltaic components; establishing a causal graph of the communication system in the physical partition, the nodes of the causal graph represent communication state parameters, the edges represent the causal relationship between the parameters, and establishing a mapping relationship with the communication network graph; based on the node feature vectors, using a multi-layer graph attention network to extract features from the communication network graph; based on the mapping relationship, the causal influence strength between the communication state parameters in the causal graph is used to adjust the attention weight calculation of the graph attention network to generate a graph-level feature representation with causal perception capability; Based on the topological structure of the communication network graph and historical node communication data, a sliding time window method is used to extract communication behavior features and construct a training sample set. A deep reinforcement learning framework is used for training, and the graph-level feature representation is input into the policy network to calculate the time slot allocation probability. The priority of the training sample is determined by the time series difference error and the overall congestion level of the network. Based on the time slot allocation probability, a greedy strategy is used to generate a communication time slot allocation sequence to allocate communication time slots to photovoltaic modules within the physical partition. The regional communication controller collects photovoltaic module operation data within the communication time slot and forwards the data between the communication levels through the data transmission path.

2. The method according to claim 1, characterized in that Collecting communication performance parameters of each communication layer, calculating routing weight values ​​between each communication layer, allocating a data transmission path for each communication layer, and sending the data transmission path to a regional communication controller of an adjacent communication layer includes: Communication performance parameters are calculated using a dynamic time window. The current value of the communication performance parameter is obtained by the weighted sum of the historical value and the real-time measurement value, where the weighting coefficient is calculated by the sampling time interval and the preset time constant. The communication performance parameters include data throughput parameter, communication delay parameter, and network congestion parameter. The normalized throughput parameter, delay parameter, and congestion parameter are weighted and combined, and multiplied by the communication layer constraint coefficient to obtain the routing weight. The communication layer constraint coefficient is used to control cross-layer communication. The multi-level communication network of the photovoltaic power station is constructed as a directed weighted graph, wherein the nodes of the directed weighted graph correspond to communication nodes, the edges correspond to communication links, and the edge weights correspond to the routing weights; and an improved Dijkstra algorithm is used to search for an optimal path in the directed weighted graph; Calculate multiple backup paths with the highest weights and path overlap below a preset overlap threshold, periodically test the connectivity and communication performance of the optimal path, and activate the backup paths in descending order of the product of their routing weights when a link is interrupted or the communication performance falls below a preset change threshold; The optimal path and backup path information are sent to a regional communication controller at an adjacent communication level.

3. The method according to claim 2, characterized in that Searching for an optimal path in the directed weighted graph using the improved Dijkstra algorithm includes: Initialize the distance of the source node to 0 and the distances of other nodes to infinity. Select the unvisited node with the smallest distance as the current node. Update the distance values ​​of the nodes adjacent to the current node and whose level difference is not greater than 1. The distance value is updated by multiplying the routing weights. When all nodes are visited, reconstruct the optimal path based on the distance values ​​of the nodes. The optimal path satisfies that the level difference of adjacent nodes on the path is not greater than 1 and the product of the routing weights of each edge on the path is the largest.

4. The method according to claim 1, wherein Establishing a causal graph of the communication system within the physical partition, wherein nodes of the causal graph represent communication state parameters, edges represent causal relationships between parameters, and establishing a mapping relationship with the communication network graph includes: Acquire communication status parameter data including node level, link level and system level; Based on the communication state parameter data, an initial structure of a causal graph is constructed through conditional independence tests and timing information constraints; a nonlinear structural equation is constructed to describe the relationship between each parameter and its direct causal parameter, historical state, and external interference, and the equation parameters are estimated by minimizing the prediction error to obtain a complete causal graph; Constructing a bidirectional mapping relationship between communication network graph nodes and causal graph nodes, including: calculating the influence of each communication network node on each communication state parameter based on the nonlinear structural equation to obtain a node-parameter influence matrix; establishing a bidirectional mapping relationship between the communication network graph nodes and the causal graph nodes based on the influence matrix, and the mapping relationship is obtained by calculating a normalized mapping function.

5. The method according to claim 1, wherein A deep reinforcement learning framework is used for training. The graph-level feature representation is input into the policy network to calculate the time slot allocation probability. The priority of the training samples is determined by the temporal difference error and the overall congestion level of the network. Build an Actor-Critic deep reinforcement learning framework, where the Actor network serves as the policy network to generate time slot allocation strategies, and the Critic network is used to evaluate state value; The graph-level feature representation is input into the policy network, a first-layer weight matrix and bias vector are linearly transformed to obtain hidden layer features, a nonlinear activation process is performed on the hidden layer features, a second-layer weight matrix and bias vector are linearly transformed to obtain the allocation score of each node in different time slots, and the score is converted into a time slot allocation probability through normalization; The node-level communication congestion degree is calculated based on the node communication queue length, data arrival rate and service rate, and the overall network congestion degree is obtained by weighting the node importance weight; Calculate the temporal difference error (TDE) including the immediate reward, the current state value, and the next state value, and use the absolute value of the TDE and the weighted combination of the overall network congestion level as the training sample priority; Experience replay sampling is performed based on the training sample priority, the policy network parameters are updated using the policy gradient method, and the value network parameters are updated using temporal difference learning. During the parameter update process of the policy network and the value network, the importance weights corresponding to the training sample priority are used for gradient correction.

6. The method according to claim 5, characterized in that Calculating the node-level communication congestion level based on the node communication queue length, data arrival rate and service rate includes: The multi-scale congestion characteristics of nodes are calculated using sliding time windows of different sizes, including the instantaneous congestion level in a short-term window, the fluctuation trend in a medium-term window, and the cumulative congestion level in a long-term window. Determine the dynamic importance weight of a node based on its degree centrality and communication demand intensity in the communication network; The multi-scale congestion features are weighted and combined using the dynamic importance weights to obtain the node-level communication congestion degree.

7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Equipment integration diagnosis system of photovoltaic power station

    CN117458993A

  • Distributed photovoltaic hierarchical regulation and control method based on cloud edge-end integrated cooperation

    CN117691753A