Heterogeneous graph-oriented characterization learning and RL strategy adaptive cooperation power transmission line equipment intelligent type selection method and system
By employing a representation learning approach for heterogeneous graphs and an adaptive collaborative method for RL policies, this study addresses the high-order interdependence and structural constraints of multi-source heterogeneous data in the selection of transmission line equipment. This approach achieves accuracy, economy, robustness, and feasibility in the selection of transmission line equipment, adapting to complex terrain and extreme weather conditions.
Patent Information
- Application Number
- CN202511160820.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies struggle to effectively utilize multi-source heterogeneous data, especially in the selection of power transmission line equipment. They cannot fully express the high-order interdependencies and structural constraints between equipment, between equipment and geography, and between equipment and meteorology, resulting in insufficient adaptability and accuracy under complex terrain and extreme weather conditions. Furthermore, they lack an integrated closed loop of HGNN representation, constraint perception loss, and adaptive RL policy, which fails to meet the requirements for robust and economical optimization across different scenarios.
This paper adopts a collaborative approach of representation learning and adaptive RL policy for heterogeneous graphs. Through standardized modeling of multi-source heterogeneous data, relation-aware representation of heterogeneous graph neural networks, joint optimization of self-supervised reinforcement and price constraints, and engineering rule embedding for consistency verification, combined with reinforcement learning-driven adaptive closed-loop iteration, the paper achieves accuracy, economy, robustness and feasibility in the selection of transmission line equipment.
It significantly improves the accuracy, economy, robustness and feasibility of power transmission line equipment selection under cross-scenario conditions, solves the defects such as insufficient multi-source information fusion, lack of relationship modeling capabilities, neglect of economy and fixed parameters, and achieves the optimal synergy between robustness and economy across scenarios.
Smart Images

Figure CN121094401A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system planning technology, and in particular to a method and system for intelligent selection of transmission line equipment based on heterogeneous graph representation learning and adaptive collaboration of RL strategy. Background Technology
[0002] The selection of transmission line equipment is crucial in the planning and construction of power systems, directly impacting the stability, transmission efficiency, safety, durability, and economy of the lines. Traditional methods often rely on human experience and simplified computational models, or use Geographic Information System (GIS) scenarios with equipment location, load-bearing capacity, and material as input, employing linear regression or rule bases to provide solutions; or they introduce supervised learning models such as Support Vector Machine (SVM) and Random Forest (RF) for category / material prediction. While these methods improve efficiency to some extent, they are generally based on "independent and identically distributed" vectorized features, making it difficult to express higher-order interdependencies and structural constraints between equipment, between equipment and geography, and between equipment and weather conditions. This results in insufficient adaptability and accuracy under complex terrain and extreme weather conditions.
[0003] With the widespread adoption of smart grids and multi-source data, data-driven structured modeling has become a hot topic. Graph Neural Networks (GNNs) can perform "message passing-feature aggregation" on graph-structured data, effectively capturing irregular and non-Euclidean relationships. Given the inherent heterogeneity of power transmission scenarios (equipment, geography, and meteorology are different types of nodes; physical connections, spatial affiliation, and environmental influences are different types of edges), heterogeneous graph neural networks (HGNNs), which can distinguish between "node types / relationship types," are more suitable for joint modeling of "equipment-geography-meteorology." Meanwhile, Reinforcement Learning (RL) has demonstrated its advantages in closed-loop optimization of "policy-environment-reward" in complex decision-making, combinatorial optimization, and adaptive control in recent years. It has been used to automatically adjust "policy hyperparameters" such as loss weights, thresholds, and resource allocation, and to balance robustness and economy in multiple scenarios. Coupled with the structure-aware representation of HGNNs and the policy adaptive optimization of RL, it is expected to dynamically meet engineering constraints such as price and schedule while ensuring safety and performance, and improve generalization ability under extreme weather and complex terrain conditions.
[0004] Existing technologies still have shortcomings: First, rule-based, linear models, and traditional machine learning (ML) (SVM, RF) schemes struggle to utilize multiple relationships, heterogeneous graph structures, and high-order neighborhood information. Second, when directly transferring homogeneous graph GNNs to power transmission selection, semantic and type differences in relationships are often ignored, limiting representational capabilities and interpretability. Third, economic and safety constraints in engineering practice (such as price caps and wind / icing margins) are often not explicitly incorporated into the objective function. Fourth, while there have been explorations of introducing reinforcement learning (RL) into power scenarios, these have mostly remained at the level of homogeneous state encoding or static reward design, resulting in insufficient reward shaping, unknown constraints, low sample efficiency, and difficulty in guaranteeing safety constraints. Fifth, the lack of an integrated closed loop of HGNN representation—constraint-aware loss—RL policy adaptation means that the sensitivity of policy updates to heterogeneous relationships, extreme weather disturbances, and price constraints is not systematically characterized.
[0005] In summary, there is an urgent need for a method that integrates representation learning and adaptive RL policy collaboration for heterogeneous graphs, which can fuse multi-source information and explicit modeling constraints within a unified framework, and achieve optimal collaboration between robustness and economy across scenarios through policy updates. Summary of the Invention
[0006] To overcome the shortcomings of existing technologies, the present invention aims to provide a method and system for intelligent selection of transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration. Through standardized modeling of multi-source heterogeneous data, relation-aware representation of heterogeneous graph neural networks, joint optimization of self-supervised enhancement and price constraints, engineering rule embedding for consistency verification, and policy adaptive closed-loop iteration driven by reinforcement learning, the method significantly improves the accuracy, economy, robustness, and feasibility of transmission line equipment selection under cross-scenario conditions. It effectively overcomes the defects of existing technologies, such as insufficient multi-source information fusion, lack of relation modeling capabilities, neglect of economic efficiency, and fixed parameters.
[0007] To achieve the above objectives, the present invention provides the following solution:
[0008] A method for intelligent selection of transmission line equipment based on the adaptive collaboration of heterogeneous graph representation learning and RL policy includes:
[0009] Collect sample equipment data, sample geographic information data, and sample static meteorological data, and uniformly encode and standardize the numerical and categorical characteristics of the sample equipment data, the sample geographic information data, and the sample static meteorological data to form a structured multi-source data table associated with the equipment sample;
[0010] Based on the device, geographic, and meteorological fields in the multi-source data table, device nodes, geographic nodes, and meteorological nodes are generated respectively, and device-device edges, device-geographic edges, and device-meteorological edges are generated;
[0011] The device node, the geographic node, the meteorological node, the device-device edge, the device-geographic edge, and the device-meteorological edge are identified as a sample heterogeneous graph that distinguishes relation types.
[0012] Based on the heterogeneous graph neural network, the sample heterogeneous graph is performed with typified embedding and parameter sharing, relation-aware multi-head attention and multi-scale aggregation, and node embedding enhancement is used to randomly mask and restore some node features during the training phase to construct node self-supervised loss.
[0013] The initial model parameters and inference threshold are obtained by jointly optimizing the comprehensive classification loss with equipment category as the supervision signal, the price loss corresponding to the price constraint, and the node self-supervision loss.
[0014] Based on the target equipment data, target geographic information data and target static meteorological data of the target project, a corresponding target heterogeneous map is constructed. Based on the initial model parameters and the inference threshold, the target heterogeneous map is inferred to obtain the category probability and preliminary selection results of each target equipment node.
[0015] Based on price constraints and safety / performance rules, a consistency check is performed between the category probabilities and the preliminary selection results to generate a candidate selection list and constraint deviation and robustness indicators;
[0016] The strategy parameters are updated online using reinforcement learning, with the graph representation of the target heterogeneous graph, the constraint bias of the candidate selection list, and the robustness index constituting the state. The strategy parameters include at least one or more of the loss weight and the selection confidence threshold.
[0017] The updated strategy parameters are used to adjust the joint optimization and inference / verification process, with the rewards being the accuracy of equipment selection, satisfaction of price constraints, and robustness to extreme weather conditions. The process is iterated repeatedly until the convergence criteria are met, and the final equipment selection list and key indicators are output.
[0018] Preferably, the heterogeneous graph neural network performs typified embedding and parameter sharing of the sample heterogeneous graph, relation-aware multi-head attention and multi-scale aggregation, and constructs a node self-supervised loss by randomly masking and restoring some node features during the training phase through node embedding enhancement, including:
[0019] For different types of nodes in the heterogeneous graph of the samples, the corresponding embedding function is called to map the original node features into an initial low-dimensional embedding vector, and the weight matrix of each embedding function is decomposed into a low-rank value to achieve parameter sharing.
[0020] The attention weights of the neighbor node features are calculated according to the edge type of the heterogeneous graph of the sample. The neighbor information within the same relation type is fused using a multi-head attention mechanism, and the node representations are independently aggregated between relation types to obtain the node representation updated in the first stage.
[0021] Based on the node representation in the first stage, node features from at least two neighborhoods are fused to form a multi-scale aggregation result, which is then added to the previous stage representation through residual connections to maintain feature stability.
[0022] During the training phase, the original features of some nodes in the heterogeneous graph of the samples are randomly masked, and only the neighbor node information of some nodes is retained as input. The feature vectors of the masked nodes are recovered based on the multi-scale aggregation results. The difference between the recovered results and the original features is calculated to obtain the node self-supervised loss.
[0023] Preferably, the comprehensive classification loss with equipment category as the supervision signal, the price loss corresponding to the price constraint, and the node self-supervised loss are jointly optimized to obtain the initial model parameters and inference threshold, including:
[0024] The final node representation of each device node in the sample heterogeneous graph is input into the classification output layer, and the cross-entropy loss corresponding to the device category label is calculated as the comprehensive classification loss.
[0025] Based on the price field corresponding to the device node in the sample heterogeneous graph, the deviation between the predicted category price of the heterogeneous graph neural network and the preset price constraint is calculated to obtain the price loss;
[0026] The comprehensive classification loss, the price loss, and the node self-supervised loss are weighted and summed according to preset loss weights to form a joint optimization objective function;
[0027] Based on the joint optimization objective function, the model parameters of the heterogeneous graph neural network are iteratively updated using the backpropagation algorithm until the convergence condition is met, thus obtaining the initial model parameters.
[0028] During the training set validation phase, the inference threshold used in the inference phase is determined based on the probability distribution of device category predictions and the validation set accuracy curve.
[0029] Preferably, a corresponding target heterogeneity map is constructed based on the target equipment data, target geographic information data, and target static meteorological data of the target project. The target heterogeneity map is then inferred based on the initial model parameters and the inference threshold to obtain the category probability and preliminary selection results for each target equipment node, including:
[0030] The target equipment data, target geographic information data, and target static meteorological data of the target project are uniformly encoded and standardized to form a target multi-source data table associated with the target equipment sample.
[0031] Based on the target multi-source data table, feature vectors of target device nodes, target geographic nodes, and target meteorological nodes are generated. Target device-device edges, target device-geographic edges, and target device-meteorological edges are generated according to physical adjacency or connection relationships, spatial affiliation relationships, and statistical partition affiliation relationships, respectively, thus constructing a target heterogeneous graph that distinguishes relationship types.
[0032] The node features of the target heterogeneous graph are input into the device selection model based on the heterogeneous graph neural network, the initial model parameters are loaded, and the probability of each target device node belonging to each candidate category is calculated.
[0033] Based on the inference threshold, the category probability of each target device node is screened and determined to identify the predicted category of each target device node, thus forming the preliminary selection result.
[0034] Preferably, a consistency check is performed between the category probabilities and the preliminary selection results based on price constraints and safety / performance rules to generate a candidate selection list and constraint deviation and robustness indicators, including:
[0035] For each target device node, read the predicted category and corresponding price from the preliminary selection results, combine them with the upper limit of the price constraint of the target device node, calculate the price deviation and mark whether it exceeds the limit;
[0036] For each target device node, read the predicted category and safety / performance indicators from the preliminary selection results, combine them with the minimum requirement value of the target device node, calculate the safety / performance deviation, and mark whether it meets the standard.
[0037] Based on the category probability, price deviation, and safety / performance deviation, a consistency score is calculated, and the inclusion in the candidate selection list is determined according to the consistency score.
[0038] For target equipment nodes included in the candidate selection list, output a comprehensive constraint deviation index of price and safety / performance, and evaluate the robustness index under a preset extreme weather disturbance set;
[0039] The target device nodes that pass the consistency determination are summarized into the candidate selection list, along with the constraint deviation and robustness index;
[0040] The consistency score is calculated using the following formula:
[0041] ;
[0042] and with The consensus was established.
[0043] The robustness index is:
[0044]
[0045] in, For target device node The class probability corresponding to the predicted class; For target device node Predicting categories in preliminary selection results The corresponding price; For target device node Price constraint ceiling; For target device node In prediction category The following safety / performance index values; For target device node Minimum safety / performance requirements; This is the threshold for consistency determination; This is a set of extreme weather disturbances used for verification. In case of disturbance The predicted safety / performance index value for this category; A consistency score is assigned, with a higher score indicating better consistency under category probability, price constraints, and safety / performance rules. For robustness indicators, This indicates that the minimum requirements are met under all disturbance conditions.
[0046] Preferably, the strategy parameters are updated online using reinforcement learning, with the graph representation of the target heterogeneous graph, the constraint biases of the candidate selection list, and robustness indicators constituting the state. This includes:
[0047] The mean vector of the node embedding vectors obtained by the heterogeneous graph neural network for each target device node in the target heterogeneous graph is extracted, and combined with the mean global constraint deviation and mean robustness index of the candidate selection list, and concatenated into a state vector and input into the reinforcement learning environment.
[0048] In the reinforcement learning environment, forward reasoning is performed on the input state vector based on the current policy network, and the output is an action vector containing the loss weight adjustment and the selection confidence threshold adjustment.
[0049] Each component of the action vector is added to the current loss weight and the selection confidence threshold to obtain an updated parameter set. The joint optimization and consistency verification process is then re-executed using the updated parameter set to calculate the reward value.
[0050] The policy network parameters are updated according to the following formula:
[0051]
[0052] in, Let the objective function of the policy network be denoted as . The number of interactions in a training batch; For the first The reward value obtained from this interaction; In the strategy parameters Below, State Corresponding actions The probability of; For the first The action vector output by the next interaction; For the first The state vector of the second interaction; the reward value Defined as
[0053]
[0054] in, This represents the improvement in selection accuracy compared to the previous iteration. This represents the increase in the price constraint satisfaction rate. This represents the improvement value of the robustness index; These are the non-negative coefficients used to balance the contribution of the three factors; These are the policy network parameters.
[0055] Preferably, the updated strategy parameters are used to adjust the joint optimization and inference / verification process, with rewards based on selection accuracy, price constraint satisfaction, and robustness to extreme weather conditions. This process is repeated iteratively until the convergence criterion is met, outputting the final equipment selection list and key indicators, including:
[0056] After each round of strategy update, based on the updated loss weights and selection confidence thresholds, the joint optimization step and inference / verification step are re-executed to obtain the selection accuracy, price constraint satisfaction rate and robustness index of the current round.
[0057] Calculate the reward value and trigger the policy network parameter update according to the following formula:
[0058]
[0059] in, For the first The reward value of each iteration; For the first The selection accuracy obtained from rounds of iteration; For the first The price constraint satisfaction rate obtained from rounds of iteration; For the first Robustness metrics obtained from rounds of iteration; , , The non-negative balance coefficients contributing to the three improvements; , , These are the values corresponding to the previous iteration;
[0060] The updated strategy parameters based on the reward value are written back to the joint optimization process to adjust the loss weight, and written back to the inference / verification process to adjust the selection confidence threshold.
[0061] If continuous In each iteration, the absolute value of the change in the reward value is less than the preset threshold. If the condition is met, then convergence is determined.
[0062] When the convergence condition is met, the final result is output as the equipment selection list, price constraint deviation, robustness index, and selection accuracy of the current iteration.
[0063] Preferably, the key indicators of the final equipment selection list for the current iteration include: the selection accuracy rate, price constraint satisfaction rate, robustness index under extreme weather disturbance conditions, and comprehensive constraint deviation index corresponding to the equipment selection list in the convergence iteration round; wherein, the selection accuracy rate is the proportion of the number of equipment nodes in the final equipment selection list whose predicted category matches the actual category to the total number of equipment nodes; the price constraint satisfaction rate is the proportion of the number of equipment nodes in the final equipment selection list whose price meets the preset price constraint to the total number of equipment nodes; the robustness index is the minimum ratio of the safety / performance index value of each equipment node to the corresponding minimum requirement value under the preset extreme weather disturbance set; and the comprehensive constraint deviation index is the mean of the weighted sum of price deviation and safety / performance deviation on the set of equipment nodes.
[0064] Preferably, the sample equipment data includes the model, main design parameters, load-bearing capacity parameters, material information, and price per unit or set of the sample equipment; the sample geographic information data includes the elevation, slope, aspect, landform type, and land use type corresponding to the installation location of the sample equipment; the sample static meteorological data includes the extreme maximum temperature, extreme minimum temperature, average wind speed, maximum wind speed, maximum instantaneous wind speed, annual precipitation, and seasonal distribution of precipitation in the sample equipment installation area for a representative month or season.
[0065] A smart selection system for transmission line equipment based on the adaptive collaboration of heterogeneous graph representation learning and RL policy includes:
[0066] The sample multi-source data acquisition and preprocessing unit is used to acquire sample equipment data, sample geographic information data and sample static meteorological data, and to uniformly encode and standardize the numerical and categorical characteristics of the sample equipment data, the sample geographic information data and the sample static meteorological data to form a structured multi-source data table associated with the equipment sample.
[0067] The sample heterogeneous graph construction unit is used to generate device nodes, geographic nodes and meteorological nodes respectively based on the device, geographic and meteorological fields in the multi-source data table, and generate device-device edges, device-geographic edges and device-meteorological edges;
[0068] A heterogeneous graph definition and storage unit is used to determine the device nodes, geographic nodes, meteorological nodes, device-device edges, device-geographic edges, and device-meteorological edges as sample heterogeneous graphs that distinguish relation types;
[0069] The heterogeneous graph neural network training unit is used to perform typified embedding and parameter sharing of the sample heterogeneous graph based on the heterogeneous graph neural network, relation-aware multi-head attention and multi-scale aggregation, and to construct node self-supervised loss by randomly masking and restoring some node features during the training phase through node embedding enhancement.
[0070] The joint optimization parameter acquisition unit is used to jointly optimize the comprehensive classification loss with equipment category as the supervision signal, the price loss corresponding to the price constraint, and the node self-supervision loss to obtain the initial model parameters and inference threshold.
[0071] The target heterogeneity map inference unit is used to construct a corresponding target heterogeneity map based on the target equipment data, target geographic information data and target static meteorological data of the target project, and to infer the target heterogeneity map based on the initial model parameters and the inference threshold to obtain the category probability and preliminary selection results of each target equipment node.
[0072] The consistency verification and indicator generation unit is used to perform consistency verification on the category probability and the preliminary selection result based on price constraints and safety / performance rules, and generate a candidate selection list and constraint deviation and robustness indicators;
[0073] The reinforcement learning policy update unit is used to update the policy parameters online using reinforcement learning, based on the graph representation of the target heterogeneous graph, the constraint bias of the candidate selection list, and the robustness index. The policy parameters include at least one or more of the loss weight and the selection confidence threshold.
[0074] The closed-loop iteration and result output unit is used to provide rewards based on selection accuracy, price constraint satisfaction, and extreme weather robustness. It uses the updated strategy parameters to adjust the joint optimization and inference / verification process, repeating the iteration until the convergence criterion is met, and outputs the final equipment selection list and key indicators.
[0075] The present invention discloses the following technical effects:
[0076] (1) This invention collects and uniformly encodes sample equipment data, sample geographic information data and sample static meteorological data to form a structured multi-source data table, thereby realizing the standardized integration of multi-source heterogeneous information related to the selection of transmission line equipment. This overcomes the defects of the existing technology, such as inconsistent formats of multiple data sources and low information fusion, which lead to insufficient quality of model training data.
[0077] (2) By generating device nodes, geographic nodes and meteorological nodes based on multi-source data tables and constructing heterogeneous graph structures of device-device edge, device-geographic edge and device-meteorological edge, this invention can simultaneously capture the physical association between devices, the spatial affiliation between devices and the geographic environment and the statistical association between devices and meteorological conditions, thus making up for the shortcomings of traditional methods in being unable to characterize the complex interactions of multiple types of nodes and multiple relationship edges.
[0078] (3) This invention adopts typed embedding, parameter sharing, relation-aware multi-head attention and multi-scale aggregation of heterogeneous graph neural networks, and introduces node embedding enhancement and self-supervised loss. This invention significantly improves the generalization ability and robustness under the conditions of small sample device categories and missing data, and solves the problem of insufficient representation ability of homogeneous graph models under heterogeneous data.
[0079] (4) The present invention jointly optimizes the comprehensive classification loss of equipment category, the price loss corresponding to the price constraint, and the node self-supervised loss. This not only ensures the accuracy of the selection, but also explicitly introduces economic constraints during the model training stage. This overcomes the shortcomings of the prior art, which only uses classification accuracy as the optimization target and ignores economic requirements, and achieves a multi-objective balance of safety, performance and investment cost.
[0080] (5) The present invention performs a comprehensive verification of the inference results by combining price constraints and safety / performance rules through the consistency verification process, and generates a candidate selection list, constraint deviation and robustness index. The present invention establishes a complete filtering mechanism from prediction output to engineering available solutions, avoiding selection results that do not meet engineering requirements due to simply relying on model output, and enhancing the feasibility and reliability of the solution.
[0081] (6) This invention introduces reinforcement learning to use the graph representation of the target heterogeneous graph, constraint bias and robustness index as the state, and updates the loss weight and selection confidence threshold online. It also uses selection accuracy, price constraint satisfaction and extreme weather robustness as rewards to form a closed-loop iterative optimization mechanism, which enables this invention to adapt to the comprehensive performance requirements of different geographical and meteorological scenarios and overcomes the problems of poor adaptability and insufficient generalization ability of traditional static parameter models in cross-scenario applications. Attached Figure Description
[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0083] Figure 1 A flowchart of the method provided in an embodiment of the present invention;
[0084] Figure 2 This is a schematic diagram of the power transmission line equipment selection process provided in an embodiment of the present invention;
[0085] Figure 3 This is a heterogeneous diagram provided for an embodiment of the present invention. Detailed Implementation
[0086] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0087] The purpose of this invention is to provide a method and system for intelligent selection of transmission line equipment based on heterogeneous graph representation learning and adaptive collaboration of RL strategies, thereby achieving precise and adaptive optimization of transmission line equipment selection.
[0088] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0089] Figure 1 The method flowchart provided in the embodiments of the present invention is as follows: Figure 1 As shown, this invention provides a method for intelligent selection of transmission line equipment based on the adaptive collaboration of heterogeneous graph representation learning and RL policy, including:
[0090] Step 100: Collect sample equipment data, sample geographic information data and sample static meteorological data, and uniformly encode and standardize the numerical and categorical characteristics of the sample equipment data, sample geographic information data and sample static meteorological data to form a structured multi-source data table associated with the equipment sample.
[0091] Step 200: Generate device nodes, geographic nodes, and meteorological nodes based on the device, geographic, and meteorological fields in the multi-source data table, and generate device-device edges, device-geographic edges, and device-meteorological edges respectively;
[0092] Step 300: Identify the device nodes, geographic nodes, meteorological nodes, device-device edges, device-geographic edges, and device-meteorological edges as sample heterogeneous graphs that distinguish relation types;
[0093] Step 400: Based on the heterogeneous graph neural network, perform typified embedding and parameter sharing of sample heterogeneous graphs, relation-aware multi-head attention and multi-scale aggregation, and construct node self-supervised loss by randomly masking and restoring some node features during the training phase through node embedding enhancement.
[0094] Step 500: Jointly optimize the comprehensive classification loss with equipment category as the supervision signal, the price loss corresponding to the price constraint, and the node self-supervision loss to obtain the initial model parameters and inference threshold;
[0095] Step 600: Construct a corresponding target heterogeneous map based on the target equipment data, target geographic information data and target static meteorological data of the target project. Infer the target heterogeneous map based on the initial model parameters and inference threshold to obtain the category probability and preliminary selection results of each target equipment node.
[0096] Step 700: Based on price constraints and safety / performance rules, perform consistency checks on the category probabilities and preliminary selection results, and generate a candidate selection list and constraint deviation and robustness indicators;
[0097] Step 800: Using the graph representation of the target heterogeneous graph, the constraint bias of the candidate selection list, and the robustness index to form the state, reinforcement learning is used to update the policy parameters online; the policy parameters include at least one or more of the loss weight and the selection confidence threshold.
[0098] Step 900: With the rewards being selection accuracy, price constraint satisfaction, and extreme weather robustness, the updated strategy parameters are used to adjust the joint optimization and inference / verification process. This process is repeated iteratively until the convergence criteria are met, and the final equipment selection list and key indicators are output.
[0099] Specifically, in step 100 of this embodiment, the category features refer to attribute fields that are non-continuous numerical types and whose values belong to a finite discrete category. These fields need to be encoded (such as one-hot encoding or embedding encoding) during the data preprocessing stage before they can be used for model training. Category features in sample equipment data include: equipment model, material information, equipment category or type; category features in sample geographic information data include: landform type, land use type; category features in sample static meteorological data include: seasonal type (such as four seasons, wet and dry seasons, etc.), and seasonal precipitation distribution category (such as concentrated, uniform, bimodal, etc.).
[0100] Optionally, step 100 of this embodiment further includes a data collection and preprocessing process, which specifically includes:
[0101] 1.1 Equipment Data Acquisition
[0102] Collect relevant information on equipment to be constructed or used in power transmission lines. This data includes, but is not limited to: equipment model (e.g., tower type, conductor type); main design parameters (height, etc.); load-bearing capacity parameters (maximum load, current carrying capacity, etc.); material information (e.g., steel pipe, angle steel, aluminum stranded wire, etc.); and unit / set price. See Table 1 for details.
[0103] Table 1. Detailed information on sample device data
[0104] Data types File format unit Data Description Equipment Model Excel - Tower type, conductor type, etc. Design parameters Excel meters (m) and kg Tower height, span, etc. Bearing capacity parameters Excel kN, A Maximum load, current carrying capacity Material Information Excel - Materials such as steel pipes, angle steel, and aluminum stranded wire Basic engineering quantity Excel Foundation concrete, foundation steel reinforcement (kg), etc. Price per unit / set Excel - Price of a single device
[0105] 1.2 Geographic Information Data Acquisition
[0106] To obtain the geographic information of the spatial coordinates of each device (such as towers and line segments), first, DEM raster elevation data (such as SRTM DEM, ASTER GDEM, ALOS World 3D-30m, etc.) is acquired. Then, it is processed using ArcGIS or other GIS software to obtain slope and aspect data. Its geomorphic classification type is obtained through other methods. Finally, the raster data is converted to an Excel spreadsheet, and the processed raster data of the device is extracted into an Excel spreadsheet for subsequent processing, as shown in Table 2.
[0107] Table 2 Detailed information on geographic information data
[0108] Data types Data source File format unit Data Description latitude and longitude coordinates - Excel Degree(°) WGS-84 coordinates Elevation (DEM) SRTM DEM, ASTER GDEM, ALOS World 3D TIFF Meter (m) altitude slope ArcGIS / QGIS DEM Processing TIFF Degree(°) Slope value Slope ArcGIS / QGIS DEM Processing TIFF Degree(°) Slope azimuth Landform classification types ArcGIS / QGIS DEM Processing TIFF category Plains, hills, and mountains Land use types Land use type data (GlobeLand30 2020 version) TIFF category Current land use types
[0109] 1.3 Static meteorological data acquisition
[0110] Meteorological data acquisition and geographic information data acquisition are carried out in the same way. The data source is ERA5 meteorological data. Historical meteorological statistics of the grid where the device is located are extracted through ArcGIS and other methods. The main data include: extreme maximum temperature and extreme minimum temperature of each season (spring, summer, autumn and winter) or representative month; average and maximum wind speed, maximum instantaneous wind speed; annual precipitation, seasonal precipitation distribution, etc.
[0111] Through the detailed data collection and processing procedures described above, this invention achieves accurate and efficient acquisition and integration of geographic and meteorological information, providing a reliable basic data guarantee for the intelligent selection of subsequent power transmission line equipment.
[0112] 1.4 Data Normalization
[0113] To eliminate the influence of different data types and units on subsequent model training, all numerical features are uniformly normalized, commonly using the min-max normalization method, specifically:
[0114] ;
[0115] in, The original data, and These are the minimum and maximum values of this feature in the entire dataset. These are the normalized eigenvalues.
[0116] 1.5 Data Structure Organization
[0117] The processed equipment data, geographic information data, and meteorological data are organized into structured tables or databases to ensure that each equipment sample can be associated with its geographic and meteorological information, providing basic data support for the subsequent construction of heterogeneous map structures.
[0118] This completes the collection, cleaning, normalization, and structuring of multi-source data, laying a solid foundation for the subsequent heterogeneous graph modeling and intelligent selection in this invention.
[0119] Furthermore, step 200 in this embodiment is the construction process of the heterogeneous graph structure, such as... Figure 2 and Figure 3 As shown, after completing multi-source data acquisition and preprocessing, this embodiment constructs a heterogeneous graph data structure for transmission lines based on different types of data, providing a foundation for subsequent heterogeneous graph neural network modeling. The specific process includes:
[0120] 2.1 Node Type Definition
[0121] Based on the preprocessed data described above, the following three types of nodes are defined:
[0122] Equipment Node: Each critical piece of equipment on a transmission line (such as each tower, conductor segment, insulator, etc.) corresponds to an equipment node. The attributes of an equipment node include various engineering parameters such as model, load-bearing capacity, material, and price.
[0123] Geographic Node: Obtain the DEM, slope, aspect, and other attributes of the DEM data raster corresponding to the latitude and longitude of the device node.
[0124] Meteorological nodes: Based on the spatial and temporal characteristics of meteorological data, each area where the transmission line equipment is located corresponds to a meteorological node. The attributes of meteorological nodes include statistical information such as extreme maximum temperature, minimum temperature, average wind speed, maximum instantaneous wind speed, annual precipitation, and seasonal precipitation distribution for each season or month.
[0125] 2.2 Edge Type Definition
[0126] This invention effectively integrates multi-source heterogeneous information such as power transmission equipment, geography, and meteorology by constructing a heterogeneous graph of power transmission lines, thereby representing more comprehensively and accurately the various factors and complex relationships upon which equipment selection depends. To achieve the above objectives, this invention constructs three different types of edges: equipment-equipment edges, equipment-geography edges, and equipment-meteorology edges. The specific definitions, construction reasons, and necessity analysis of each type of edge are as follows:
[0127] 1) Equipment-equipment edge: If two devices (such as towers, conductors or insulators) are directly adjacent or have a clear physical connection relationship in the actual transmission line topology (such as the suspension relationship between conductors and towers, the connection relationship between towers, etc.), then an equipment-equipment edge is constructed between the corresponding two equipment nodes.
[0128] Transmission lines are essentially network systems with well-defined physical structures. The physical adjacency or connection relationships between equipment such as towers, conductors, and insulators directly affect the transmission or interaction of mechanical and electrical properties between these devices. Therefore, clarifying these relationships and representing them in heterogeneous diagrams is an important prerequisite for models to capture real-world scenarios.
[0129] Through device-to-device edges, graph neural networks can directly aggregate information from adjacent device nodes. This is beneficial for the model to learn the possible "local correlation" patterns between towers and conductors, such as the mutual influence relationship between load characteristics and current carrying capacity of adjacent towers.
[0130] 2) Device-Geographic Edge: When a device node (e.g., a pole) is located within a specific geographic grid or geographic cell (DEM grid), a device-geographic edge is constructed between the device node and the corresponding geographic node.
[0131] The geographical environment of equipment (such as elevation, slope, and landform type) significantly affects its stress condition, construction difficulty, and equipment selection requirements. Through equipment-geographic edges, the feature information (DEM height, slope, etc.) in geographic nodes can be effectively transmitted to equipment nodes, enabling graph neural networks to more accurately understand how the geographical environment affects the technical requirements of equipment.
[0132] Geographical factors have a long-term impact on equipment performance and lifespan. Clearly connecting equipment and geographical nodes enables the model to clearly identify the intrinsic relationship between equipment selection and geographical factors during the training process, thereby improving the scientific nature and accuracy of equipment selection decisions.
[0133] 3) Device-meteorological edge: If the location of the device node is within a meteorological statistical area (such as a meteorological grid divided by administrative boundaries or meteorological statistical units), then a device-meteorological edge is constructed between the device node and the corresponding meteorological node.
[0134] Meteorological factors (such as wind speed, temperature, and precipitation) significantly affect the stability, reliability, and durability of power transmission equipment. Through the equipment-meteorological edge, meteorological condition information can be explicitly transmitted to equipment nodes, thereby enabling accurate modeling of the differences in equipment selection under different climatic conditions.
[0135] Equipment selection requirements vary significantly under different meteorological environments (e.g., in areas with high wind speeds, tower selection tends to prioritize mechanical strength). Constructing an equipment-meteorological boundary helps graph neural networks fully perceive the differences in equipment environments, thereby enabling the selection model to better adapt to complex and variable meteorological factors.
[0136] 2.3 Heterogeneous Graph Structured Storage
[0137] To facilitate efficient model computation and training, this embodiment stores the heterogeneous graph structure in a unified structured storage format (a combination of graph database, adjacency matrix, and node feature matrix). The specific storage structure is as follows:
[0138] Each node type constitutes an independent node feature matrix. The rows represent node instances, and the columns represent node features.
[0139] Edge information is uniformly stored in either an adjacency list or an adjacency matrix. The adjacency matrix is defined as follows: .in, , and These represent the number of device nodes, the number of geographical nodes, and the number of meteorological nodes, respectively. It is the set of real numbers.
[0140] Furthermore, steps 300 and 400 of this embodiment are for the design and training of a heterogeneous graph neural network model. After the heterogeneous graph structure is constructed, this embodiment uses a heterogeneous graph neural network to model and train the constructed heterogeneous graph of the transmission line, thereby realizing intelligent prediction of the equipment node category. The heterogeneous graph neural network model includes: (1) Data input layer: the node features after data preprocessing are used as input; (2) Node embedding layer: the input data is converted into an initial node feature vector through a node-specific type embedding function; (3) Multi-head heterogeneous graph attention layer: the relationship between nodes is calculated using a dynamic attention mechanism, attention weights are generated and node feature information is transmitted; (4) Multi-scale aggregation layer: information at different scales is aggregated and node features are updated through residual connections; (5) Classification output layer: the node category prediction is output through a fully connected layer after the final aggregated node features are processed. In the model training, the data is first preprocessed and then input into the node embedding layer to generate initial node features. Then, in the multi-head heterogeneous graph attention layer, the node features are transmitted and aggregated through dynamic attention weights, and intermediate layer features are output. Next, a multi-scale aggregation layer further integrates node information, enhancing the model's ability to perceive local and global features. Finally, the aggregated node features are mapped to specific categories through a classification output layer, completing the node classification task. During training, model parameters are jointly optimized using node self-supervised loss and classification loss, achieving efficient and accurate data-driven transmission line equipment selection decisions.
[0141] Specifically, to effectively represent node features in heterogeneous graphs and improve the robustness and generalization ability of node features during model learning, this embodiment employs node embedding and node embedding enhancement mechanisms. Based on device type, a specific feature embedding function is implemented by replacing the embedding layer weight matrix with low-rank decomposition to achieve parameter sharing, specifically including:
[0142]
[0143] in, The vector embedding representation of node v; A feature embedding function specific to the type; These are the original node features; For the current node; It is a set of nodes.
[0144] The weight matrix of the embedding layer is composed of a low-rank decomposition matrix:
[0145]
[0146] in, , , , It is the input feature dimension. It is the embedded feature dimension. It is the rank of the decomposition. By using low-rank decomposition to achieve parameter sharing, the overfitting problem of small sample equipment types in the power grid dataset can be solved.
[0147] Preferably, to further improve the generalization performance and robustness of node embedding, this embodiment proposes a node embedding enhancement mechanism, namely, introducing a self-supervised learning method: by randomly masking some node features during the training phase, the masked node features are predicted and recovered using their neighborhood information. This self-supervised mechanism can enhance the generalization of node features, enabling node feature embedding to not only include the node's own features but also fully integrate the effective information of neighboring nodes, thereby improving the network's generalization ability to complex scenarios and unknown nodes. The loss function of this node self-supervised process is defined as:
[0148]
[0149] in, This is the self-monitored loss value; For the set of feature-masked nodes; For shielded nodes The set of neighboring nodes.
[0150] Node self-supervision enhances the model's robustness to node embedding, enabling node features to more comprehensively reflect local structural information and contextual relationships, thereby improving the overall model's generalization performance.
[0151] Specifically, a multi-head heterogeneous graph attention layer is used. To effectively model the complex relationships between nodes in a heterogeneous graph and avoid relying on manually defined static relationships, this embodiment employs a graph attention mechanism. Nodes in the graph not only need to aggregate information from their neighbors but also need to dynamically adjust the weights of different edges based on the relationship types between nodes (such as weather conditions, geographical location, equipment performance, etc.). The attention weights of heterogeneous edges (multi-head) are calculated, for each head... ,node For neighboring nodes The attention weights are calculated as follows:
[0152]
[0153] Among them, among them, To indicate the head Middle node pair The initial attention coefficient; Indicates the activation function; For the head Middle node type and The relationship weight vector; For the head Middle node type The corresponding feature mapping weight matrix; For nodes Features; This indicates feature splicing.
[0154] Then normalize the attention weights:
[0155]
[0156] The formula above calculates the attention weight of the relationship between nodes and uses this weight to perform weighted aggregation of the features of neighboring nodes. For nodes The neighboring nodes; For nodes The set of neighboring nodes.
[0157] In this way, the relationships between nodes can be dynamically modeled, and the update process of each node takes into account its actual relationship with its neighbors. For the selection of transmission line equipment, factors such as different meteorological conditions, geographical locations, and equipment characteristics can be effectively expressed through this relationship modeling method, enabling the selection decision to fully consider various complex environmental factors.
[0158] Preferably, multi-scale neighborhood aggregation is employed, and the graph neural network uses graph convolution operations to achieve information propagation between nodes. Each node not only receives information from its direct neighbors but also aggregates this information in a weighted manner to update its features. Let... If is the network layer number, then the _th The update formula for layer node features is:
[0159]
[0160] in, For the first Layer nodes Update features; Represents a nonlinear activation function; For the first Layer residual connection weight matrix; For the first Characteristics of layer nodes; For the first Layer Middle node With nodes Attention weights between them; For the first Layer Corresponding node type A specific weight matrix; For the number of attention heads.
[0161] Specifically, the key to graph convolution lies in updating the features of the target node by aggregating the features of neighboring nodes. This process effectively transmits information from neighboring nodes to the target node and dynamically adjusts the influence of different neighboring nodes' information on the target node's feature update based on attention weights. Through graph convolution, nodes can receive information from their neighbors and fuse this information in each layer's update. This ensures that each node's features not only reflect its own attributes but also include the influence of its neighbors. Furthermore, through the self-attention mechanism, each edge (relationship between nodes) in the graph can be dynamically assigned different weights based on its actual importance. The edge weights allow nodes with strong relationships to have a greater impact on the target node's feature update, thereby improving the model's adaptability in complex environments.
[0162] Further, signal convergence and node classification. After multiple rounds of graph convolution and information aggregation, the final features of a node will contain information about the complex relationships between it and its neighboring nodes. These features play a crucial role in the node classification task. Specifically, after multi-layer attention aggregation in a graph neural network, the node's features will contain information from multiple neighborhoods. Finally, a fully connected layer (FC layer) maps these features to the node's class space, thereby achieving node classification.
[0163] Optionally, step 500 in this embodiment specifically involves an equipment selection decision and optimization output process. After completing the training of the heterogeneous graph neural network model, this embodiment uses the equipment node category prediction results output by the model, combined with the actual economic requirements of the project, to make selection decisions and optimization outputs for transmission line equipment. The specific process is as follows:
[0164] 4.1 Obtaining Equipment Category Prediction Results
[0165] Specifically, for the task of selecting power grid towers, the classification objective is to predict the most suitable selection category for each unknown tower node. The specific process includes: Feature mapping: through multi-layer graph convolution and attention mechanisms, nodes... Feature vectors obtained in the final layer It gathers information about all neighboring nodes and their relationships. Class prediction: The final features are mapped to a class space through a fully connected layer. Typically, the class space is a discrete set, for example, it can be used to classify different types of power grid equipment. Specifically, the formula for class prediction is:
[0166]
[0167] in, For nodes The predicted category probability distribution; This represents the normalized activation function; This is the weight matrix; Let i be the vector representation of node i in the Lth layer; This is a bias term.
[0168] In this way, the node classification task not only considers the characteristics of the nodes themselves, but also dynamically incorporates the complex relationships between nodes through graph convolutional networks and attention mechanisms, thereby making more accurate classification decisions. In the scenario of power grid tower selection, the model can accurately predict suitable equipment selection based on the relationship between unknown nodes and other nodes, as well as the obtained characteristics.
[0169] Furthermore, to fully consider the impact of equipment price on the model output, this embodiment adds a price optimization loss function during the model optimization stage, specifically defined as follows:
[0170]
[0171] in, To synthesize the classification loss function, This is the weighting coefficient for the price constraint term, used to adjust the importance of price; in this embodiment, it is set to 1. The price loss function is specifically defined as follows:
[0172]
[0173] in, Indicates price loss, Indicates the total number of nodes. Represents a node Corresponding equipment price, Represents a node Output the probability of the selected category.
[0174] The existing loss function is the cross-entropy loss function, which is defined as follows:
[0175]
[0176] in: For equipment selection losses, C represents the total number of candidate categories for transmission line equipment. Represents a node Actual one-hot encoded labels, if the node Category The value is 1 if the value is 1, otherwise it is 0. Represents a node Predict the probability output belonging to node c, which is the Softmax probability value of the model output layer.
[0177] Furthermore, this embodiment comprehensively considers the node self-supervision loss ( ) and comprehensive classification loss ( ), through weighting coefficients By adjusting the relationship between the two losses, joint optimization can be achieved, thereby simultaneously improving the robustness of node feature representation and the accuracy and economy of device classification.
[0178] Specifically, the combined joint loss function is:
[0179]
[0180] in: The overall loss function for model training. The overall classification loss includes both classification accuracy loss and price loss. This is the node self-supervised loss, used to improve the generalization ability of node feature representation. This is the self-supervised loss weight coefficient, used to control the degree of influence of the self-supervised task on the overall training task. In this embodiment, its position is set to 1.
[0181] Specifically, in step 600 of this embodiment, the process of constructing a corresponding target heterogeneous map based on the target equipment data, target geographic information data, and target static meteorological data of the target project, and inferring the category probability and preliminary selection results of each target equipment node based on the initial model parameters and the inference threshold, is as follows: First, the numerical and categorical features of the target equipment data, target geographic information data, and target static meteorological data collected for the target project are uniformly encoded and standardized. The target equipment data includes at least the equipment model, main design parameters, load-bearing capacity parameters, material information, and price per unit or set. The target geographic information data includes at least the elevation, slope, aspect, landform type, and land use type corresponding to the equipment installation location. The target static meteorological data includes at least the extreme maximum temperature, extreme minimum temperature, average wind speed, maximum wind speed, maximum instantaneous wind speed, annual precipitation, and seasonal precipitation distribution of the equipment installation area. The aforementioned uniform encoding refers to converting categorical features into numerical vectors that can be input into the model through one-hot encoding or embedding mapping. Standardization refers to scaling the numerical features to a uniform scale using zero-mean unit variance or normalization methods. The processed data is associated with unique device identifiers to generate a target multi-source data table that corresponds one-to-one with the target device samples. Subsequently, feature vectors for target device nodes, target geographic nodes, and target meteorological nodes are generated based on this target multi-source data table. Then, according to the physical adjacency or connection relationships between devices, the spatial affiliation between devices and geographic locations, and the statistical zoning affiliation between devices and meteorological conditions, target device-device edges, target device-geographic edges, and target device-meteorological edges are constructed respectively, thus obtaining a target heterogeneous graph that distinguishes relationship types. Here, a heterogeneous graph refers to a graph structure where nodes and edges have different types of graph structures, capable of simultaneously representing multiple entities such as devices, geographic, and meteorological entities, and their various relationship types. The node features of the target heterogeneous graph are input into a device selection model based on a heterogeneous graph neural network, loading the initial model parameters obtained during the sample training phase, and calculating the probability of each target device node belonging to each candidate category. Finally, the category probabilities of each target device node are filtered and determined based on an inference threshold. The inference threshold controls the confidence level of the classification results; categories with probabilities higher than this threshold are used as the predicted categories of the node, forming the preliminary selection results for the target project.
[0182] In one specific implementation, step 700 of this embodiment is as follows: A consistency check is performed on the category probability and the preliminary selection result based on price constraints and safety / performance rules. Specifically, for each target equipment node, firstly, the predicted category and corresponding price of the node are read from the preliminary selection result, and compared with the upper limit of the price constraint for that node to obtain the price deviation; simultaneously, the safety / performance index of the predicted category is read, and compared with the minimum requirement value for that node to obtain the safety / performance deviation; then, a consistency score is constructed using the category probability, price deviation, and safety / performance deviation. The score is calculated by subtracting the normalization penalty of the price deviation from the probability contribution, and then subtracting the normalization penalty of the safety / performance deviation. The system employs a "uniform penalty" function and uses a consistency threshold to determine the scores. Nodes with scores not lower than the threshold are included in the candidate selection list. For nodes included in the list, the system calculates the weighted composite of price and safety / performance deviations and averages them across the node set to obtain a comprehensive constraint deviation index. Furthermore, under a preset set of extreme weather disturbances, the system re-evaluates the safety / performance index for each candidate node's prediction category and takes the minimum ratio relative to the minimum requirement value as the node's robustness index. Finally, the system aggregates the target equipment nodes that pass the consistency determination to generate a candidate selection list, along with the comprehensive constraint deviation index and robustness index for subsequent adaptive optimization of the strategy.
[0183] Furthermore, the source, values, and functional descriptions of the parameters involved in the above consistency score and robustness index are as follows:
[0184] The category probability is derived from the inference output of the equipment selection model based on the heterogeneous graph neural network on the target heterogeneous graph, with a value ranging from 0 to 1, used to characterize the confidence of the predicted category; the price corresponding to the predicted category comes from the price field in the target equipment data, and the upper limit of the price constraint comes from the engineering budget or design constraints, used to constrain the upper limit of investment; the safety / performance indicators of the predicted category come from the selection rule base or equipment technical parameter table, and the minimum requirement value comes from the engineering specifications or operation and maintenance safety boundary, used to ensure that safety and performance meet the standards; the consistency judgment threshold is a fixed threshold used to screen candidate nodes, which can be set as a certain value according to the performance of the validation set. The constant is between 0 and 1, and 0.2 is selected in this embodiment; the extreme weather disturbance set is a set of representative cases used to verify robustness, which may include disturbance scenarios that can be achieved by the project, such as a certain percentage increase in wind speed, a certain degree Celsius decrease in temperature, and a certain percentage increase in precipitation; the robustness index is the minimum ratio of the safety / performance index to the minimum requirement value under all disturbance conditions, and a value of not less than 1 indicates that the minimum requirement is met under all disturbance conditions; the comprehensive constraint deviation index is the average value of the price deviation and the safety / performance deviation on the candidate node set after equal weighting, which is used to measure the overall deviation level of the candidate scheme on the project constraints.
[0185] In one specific implementation, step 800 of this embodiment is as follows: using the graph representation of the target heterogeneous graph, the constraint biases and robustness indicators of the candidate selection list to form a state, reinforcement learning is used to update the policy parameters online, specifically including the following four stages:
[0186] 7.1 State Construction: Node embedding vectors for each target device node are extracted from the target heterogeneous graph and obtained through a heterogeneous graph neural network. An arithmetic mean vector is obtained by averaging these embedding vectors. Simultaneously, the arithmetic mean of price constraint deviations in the candidate selection list is calculated to obtain the global constraint deviation mean, and the arithmetic mean of robustness indices is calculated to obtain the robustness index mean. The mean vector, the global constraint deviation mean, and the robustness index mean are concatenated in a fixed order to form a state vector, which is then input into the reinforcement learning environment. The aforementioned "node embedding vector" is a low-dimensional numerical representation of node features and relational structures aggregated by the heterogeneous graph neural network's encoding layer, used to characterize the device-geography-meteorological relationship information. The "reinforcement learning environment" refers to a closed-loop execution system that takes the state vector as input, feeds back the actions output by the policy network to the joint optimization and consistency verification process, and generates a reward value.
[0187] 7.2 Policy Inference and Action Generation: In the reinforcement learning environment, the current policy network is invoked, and forward inference is performed on the input state vector to obtain the action vector. The action vector contains two components: the adjustment amount of the loss weight and the adjustment amount of the selection confidence threshold. To ensure training stability, the two components are restricted to preset intervals (for example, the adjustment amount of the loss weight is restricted to between -0.2 and 0.2, and the adjustment amount of the selection confidence threshold is restricted to between -0.1 and 0.1. The above intervals are used in this embodiment). The resulting action vector is the online update suggestion of the policy for the parameters in this round.
[0188] 7.3 Environment Execution and Reward Calculation: The two components of the action vector are added to the current loss weight and the selection confidence threshold respectively to obtain the updated parameter set; the joint optimization and consistency verification process is re-executed with this parameter set to obtain the selection accuracy, price constraint satisfaction rate and robustness index of the new round; the difference between the corresponding index of this round and the previous round is summed in a weighted manner to obtain the reward value. The weights are used to balance the contribution of the three improvements. In this embodiment, the three weights are 0.5, 0.3 and 0.2 respectively; when the reward value is non-negative, it indicates that the overall performance has not deteriorated, and when the reward value is positive, it indicates that the overall performance has improved. The magnitude of the reward value is used to guide the policy network parameters to update in the desired direction.
[0189] 7.4 Policy Update and Convergence Output: Within a training batch, states, actions, rewards, and action occurrence probabilities are collected sequentially at a fixed number of interactions. A policy gradient-based update rule is used to iteratively optimize the policy network parameters to maximize the expected reward. After parameter updates, the new loss weights and selection confidence thresholds are written back to the joint optimization and inference / verification process, and the next batch iteration begins. The convergence criterion is set as follows: the absolute value of the reward change within 5 consecutive rounds is less than 0.001, or the selection accuracy, price constraint satisfaction rate, and robustness index on the validation set no longer significantly improve (the improvement is less than 0.1%). At this point, iteration stops, and the final equipment selection list, along with the corresponding selection accuracy, price constraint satisfaction rate, robustness index, and comprehensive constraint deviation index, are output as key indicators.
[0190] In one specific implementation, step 900 of this embodiment is: using selection accuracy, price constraint satisfaction, and extreme weather robustness as rewards, the updated strategy parameters are used to adjust the joint optimization and inference / verification process and iterated cyclically, specifically including:
[0191] 9.1 After each round of strategy update, based on the recently written-back loss weights and selection confidence threshold, the joint optimization step and inference / verification step are re-executed to obtain the selection accuracy, price constraint satisfaction rate, and robustness index for the current round. The improvement values of these three indicators relative to the previous round are weighted and summed using fixed coefficients to obtain the scalar reward value (in this embodiment, the three balance coefficients are 0.5, 0.3, and 0.2, respectively). When any indicator deteriorates, a corresponding deduction is generated. Among them, the "selection confidence threshold" refers to the lower probability limit used to make class determination from the class probability distribution, the "reward value" refers to the single evaluation quantity used to guide the direction of reinforcement learning updates, the "robustness index" refers to the minimum ratio of equipment safety / performance relative to the minimum requirements under the preset extreme weather disturbance set, the "joint optimization step" refers to the joint training process that integrates classification loss, price loss, and node self-supervised loss, and the "inference / verification step" refers to the process of generating class probabilities on the target heterogeneous graph and performing consistency verification according to price constraints and safety / performance rules.
[0192] 9.2 A policy gradient-based update rule is used to iteratively optimize the policy network parameters to maximize the expected reward. Within each training batch, several sets of "state-action-reward-action occurrence probability" samples are sequentially collected (the number of interactions per batch is set to 16 in this embodiment). After the parameter update is completed, the new loss weights are written back to the joint optimization step, and the new selection confidence thresholds are written back to the inference / verification step before entering the next iteration. To ensure stability, the single-round adjustment of the loss weights is limited to the range of -0.2 to 0.2, and the single-round adjustment of the selection confidence thresholds is limited to the range of -0.1 to 0.1 to prevent excessive step size from causing training divergence. The "policy network" refers to a function approximator that receives the state vector and outputs the parameter adjustment in the reinforcement learning environment. Its output is the action vector, corresponding to the loss weight adjustment and selection confidence threshold adjustment in this embodiment.
[0193] 9.3 Set convergence criteria to terminate iteration and output results: When the absolute value of the change in the return value is less than 0.001 for 5 consecutive rounds, or when the improvement of the selection accuracy, price constraint satisfaction rate, and robustness index is less than 0.1% for 5 consecutive rounds, convergence is determined; when the convergence condition is met, the equipment selection list for the current round is output as the final result, and key indicators, including selection accuracy, price constraint satisfaction rate, robustness index, and comprehensive constraint deviation index, are output at the same time; the calculation method of the above key indicators is consistent with that of the aforementioned consistency verification stage to ensure the consistency between the training loop and the engineering evaluation method.
[0194] Furthermore, in this embodiment, the node-level indicators of the final equipment selection list are shown in Table 3. Each row corresponds to a target equipment node, listing the node's predicted category, actual category, predicted category probability, actual price, upper limit of price constraint, price deviation, safety / performance index value, minimum safety / performance requirements, safety / performance deviation, robustness ratio, consistency judgment result (including price satisfaction, prediction accuracy, and compliance robustness), and consistency score. The category probability is inferred by the equipment selection model based on heterogeneous graph neural networks and is used to characterize the confidence level of the predicted category. The price deviation and safety / performance deviation respectively characterize the degree of deviation of the predicted category from the engineering constraints in terms of price and performance. The robustness ratio reflects the performance stability of the node under preset extreme weather disturbance conditions. The consistency score is a single-node comprehensive consistency index derived from the category probability, price constraint, and performance constraint, and is used for the subsequent selection and ranking of candidate equipment.
[0195] Table 3 Equipment Selection List
[0196] Device Node Prediction Category Real Category Category Probability price Price constraint cap Price deviation Safety performance indicators Minimum safety performance requirements Safety performance deviation Robustness ratio Price meets Prediction correct Robustness to meet standards Consistency score N01 Type A Type A 0.91 9500 10000 0.0000 1.12 1.00 0.0000 1.08 yes yes yes 0.9100 N02 TypeB TypeB 0.88 11000 12000 0.0000 1.05 1.00 0.0000 1.02 yes yes yes 0.8800 N03 Type A Type C 0.72 10200 10000 0.0200 0.98 1.00 0.0200 0.95 no no no 0.6800 N04 Type C Type C 0.84 9800 10000 0.0000 1.10 1.00 0.0000 1.03 yes yes yes 0.8400 N05 TypeB TypeB 0.79 12500 12000 0.0417 1.01 1.00 0.0000 1.00 no yes yes 0.7483 N06 Type A Type A 0.93 8700 9000 0.0000 1.07 1.00 0.0000 1.04 yes yes yes 0.9300 N07 Type C Type C 0.76 9200 9500 0.0000 0.99 1.00 0.0100 0.97 yes yes no 0.7500 N08 TypeB TypeB 0.89 11800 12000 0.0000 1.03 1.00 0.0000 1.01 yes yes yes 0.8900
[0197] In this embodiment, as shown in Table 4, the overall performance under convergence iteration rounds is listed, including selection accuracy, price constraint satisfaction rate, robustness index, and comprehensive constraint deviation index. Selection accuracy is the proportion of equipment nodes in the final equipment selection list whose predicted category matches the actual category out of the total number of equipment nodes; price constraint satisfaction is the proportion of equipment nodes that meet budget or design price constraints out of the total number of equipment nodes; robustness index is the minimum ratio of the safety / performance index values of each equipment node to the minimum required value under a preset extreme weather disturbance set, reflecting the overall adaptability under disturbance conditions; comprehensive constraint deviation index is the weighted average of price deviation and safety / performance deviation, used to measure the degree of deviation of the overall scheme from engineering constraints, with a lower value indicating closer proximity to the constraint requirements.
[0198] Table 4 Calculation Results of Key Indicators
[0199] Key Indicators Value Selection accuracy 87.50% Price constraint satisfaction rate 75.00% Robustness index (minimum ratio) 0.95 Comprehensive constraint deviation index (mean) 0.0102
[0200] Corresponding to the above method, this embodiment also provides an intelligent selection system for transmission line equipment based on the adaptive collaboration of heterogeneous graph representation learning and RL policy, including:
[0201] The sample multi-source data acquisition and preprocessing unit is used to acquire sample equipment data, sample geographic information data and sample static meteorological data, and to uniformly encode and standardize the numerical and categorical characteristics of the sample equipment data, the sample geographic information data and the sample static meteorological data to form a structured multi-source data table associated with the equipment sample.
[0202] The sample heterogeneous graph construction unit is used to generate device nodes, geographic nodes and meteorological nodes respectively based on the device, geographic and meteorological fields in the multi-source data table, and generate device-device edges, device-geographic edges and device-meteorological edges;
[0203] A heterogeneous graph definition and storage unit is used to determine the device nodes, geographic nodes, meteorological nodes, device-device edges, device-geographic edges, and device-meteorological edges as sample heterogeneous graphs that distinguish relation types;
[0204] The heterogeneous graph neural network training unit is used to perform typified embedding and parameter sharing of the sample heterogeneous graph based on the heterogeneous graph neural network, relation-aware multi-head attention and multi-scale aggregation, and to construct node self-supervised loss by randomly masking and restoring some node features during the training phase through node embedding enhancement.
[0205] The joint optimization parameter acquisition unit is used to jointly optimize the comprehensive classification loss with equipment category as the supervision signal, the price loss corresponding to the price constraint, and the node self-supervision loss to obtain the initial model parameters and inference threshold.
[0206] The target heterogeneity map inference unit is used to construct a corresponding target heterogeneity map based on the target equipment data, target geographic information data and target static meteorological data of the target project, and to infer the target heterogeneity map based on the initial model parameters and the inference threshold to obtain the category probability and preliminary selection results of each target equipment node.
[0207] The consistency verification and indicator generation unit is used to perform consistency verification on the category probability and the preliminary selection result based on price constraints and safety / performance rules, and generate a candidate selection list and constraint deviation and robustness indicators;
[0208] The reinforcement learning policy update unit is used to update the policy parameters online using reinforcement learning, based on the graph representation of the target heterogeneous graph, the constraint bias of the candidate selection list, and the robustness index. The policy parameters include at least one or more of the loss weight and the selection confidence threshold.
[0209] The closed-loop iteration and result output unit is used to provide rewards based on selection accuracy, price constraint satisfaction, and extreme weather robustness. It uses the updated strategy parameters to adjust the joint optimization and inference / verification process, repeating the iteration until the convergence criterion is met, and outputs the final equipment selection list and key indicators.
[0210] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0211] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for intelligent selection of transmission line equipment based on the adaptive collaboration of representation learning and RL policy for heterogeneous graphs, characterized in that, include: Collect sample equipment data, sample geographic information data, and sample static meteorological data, and uniformly encode and standardize the numerical and categorical characteristics of the sample equipment data, the sample geographic information data, and the sample static meteorological data to form a structured multi-source data table associated with the equipment sample; Based on the device, geographic, and meteorological fields in the multi-source data table, device nodes, geographic nodes, and meteorological nodes are generated respectively, and device-device edges, device-geographic edges, and device-meteorological edges are generated; The device node, the geographic node, the meteorological node, the device-device edge, the device-geographic edge, and the device-meteorological edge are identified as a sample heterogeneous graph that distinguishes relation types. Based on the heterogeneous graph neural network, the sample heterogeneous graph is performed with typified embedding and parameter sharing, relation-aware multi-head attention and multi-scale aggregation, and node embedding enhancement is used to randomly mask and restore some node features during the training phase to construct node self-supervised loss. The initial model parameters and inference threshold are obtained by jointly optimizing the comprehensive classification loss with equipment category as the supervision signal, the price loss corresponding to the price constraint, and the node self-supervision loss. Based on the target equipment data, target geographic information data and target static meteorological data of the target project, a corresponding target heterogeneous map is constructed. Based on the initial model parameters and the inference threshold, the target heterogeneous map is inferred to obtain the category probability and preliminary selection results of each target equipment node. Based on price constraints and safety / performance rules, a consistency check is performed between the category probabilities and the preliminary selection results to generate a candidate selection list and constraint deviation and robustness indicators; The strategy parameters are updated online using reinforcement learning, with the graph representation of the target heterogeneous graph, the constraint bias of the candidate selection list, and the robustness index constituting the state. The strategy parameters include at least one or more of the loss weight and the selection confidence threshold. The updated strategy parameters are used to adjust the joint optimization and inference / verification process, with the rewards being the accuracy of equipment selection, satisfaction of price constraints, and robustness to extreme weather conditions. The process is iterated repeatedly until the convergence criteria are met, and the final equipment selection list and key indicators are output.
2. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 1, characterized in that, Based on a heterogeneous graph neural network, the sample heterogeneous graphs are subjected to typified embedding and parameter sharing, relation-aware multi-head attention and multi-scale aggregation. During the training phase, node embedding enhancement is used to randomly mask and restore some node features, constructing a node self-supervised loss, including: For different types of nodes in the heterogeneous graph of the samples, the corresponding embedding function is called to map the original node features into an initial low-dimensional embedding vector, and the weight matrix of each embedding function is decomposed into a low-rank value to achieve parameter sharing. The attention weights of the neighbor node features are calculated according to the edge type of the heterogeneous graph of the sample. The neighbor information within the same relation type is fused using a multi-head attention mechanism, and the node representations are independently aggregated between relation types to obtain the node representation updated in the first stage. Based on the node representation in the first stage, node features from at least two neighborhoods are fused to form a multi-scale aggregation result, which is then added to the previous stage representation through residual connections to maintain feature stability. During the training phase, the original features of some nodes in the heterogeneous graph of the samples are randomly masked, and only the neighbor node information of some nodes is retained as input. The feature vectors of the masked nodes are recovered based on the multi-scale aggregation results. The difference between the recovered results and the original features is calculated to obtain the node self-supervised loss.
3. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 1, characterized in that, The initial model parameters and inference thresholds are obtained by jointly optimizing the comprehensive classification loss (using equipment category as the supervision signal), the price loss corresponding to the price constraint, and the node self-supervised loss, including: The final node representation of each device node in the sample heterogeneous graph is input into the classification output layer, and the cross-entropy loss corresponding to the device category label is calculated as the comprehensive classification loss. Based on the price field corresponding to the device node in the sample heterogeneous graph, the deviation between the predicted category price of the heterogeneous graph neural network and the preset price constraint is calculated to obtain the price loss; The comprehensive classification loss, the price loss, and the node self-supervised loss are weighted and summed according to preset loss weights to form a joint optimization objective function; Based on the joint optimization objective function, the model parameters of the heterogeneous graph neural network are iteratively updated using the backpropagation algorithm until the convergence condition is met, thus obtaining the initial model parameters. During the training set validation phase, the inference threshold used in the inference phase is determined based on the probability distribution of device category predictions and the validation set accuracy curve.
4. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 1, characterized in that, Based on the target equipment data, target geographic information data, and target static meteorological data of the target project, a corresponding target heterogeneity map is constructed. Based on the initial model parameters and the inference threshold, the target heterogeneity map is inferred to obtain the category probability and preliminary selection results for each target equipment node, including: The target equipment data, target geographic information data, and target static meteorological data of the target project are uniformly encoded and standardized to form a target multi-source data table associated with the target equipment sample. Based on the target multi-source data table, feature vectors of target device nodes, target geographic nodes, and target meteorological nodes are generated. Target device-device edges, target device-geographic edges, and target device-meteorological edges are generated according to physical adjacency or connection relationships, spatial affiliation relationships, and statistical partition affiliation relationships, respectively, thus constructing a target heterogeneous graph that distinguishes relationship types. The node features of the target heterogeneous graph are input into the device selection model based on the heterogeneous graph neural network, the initial model parameters are loaded, and the probability of each target device node belonging to each candidate category is calculated. Based on the inference threshold, the category probability of each target device node is screened and determined to identify the predicted category of each target device node, thus forming the preliminary selection result.
5. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 1, characterized in that, Based on price constraints and safety / performance rules, a consistency check is performed between the category probabilities and the preliminary selection results to generate a candidate selection list and constraint deviation and robustness indicators, including: For each target device node, read the predicted category and corresponding price from the preliminary selection results, combine them with the upper limit of the price constraint of the target device node, calculate the price deviation and mark whether it exceeds the limit; For each target device node, read the predicted category and safety / performance indicators from the preliminary selection results, combine them with the minimum requirement value of the target device node, calculate the safety / performance deviation, and mark whether it meets the standard. Based on the category probability, price deviation, and safety / performance deviation, a consistency score is calculated, and the inclusion in the candidate selection list is determined according to the consistency score. For target equipment nodes included in the candidate selection list, output a comprehensive constraint deviation index of price and safety / performance, and evaluate the robustness index under a preset extreme weather disturbance set; The target device nodes that pass the consistency determination are summarized into the candidate selection list, along with the constraint deviation and robustness index; The consistency score is calculated using the following formula: ; and with The consensus was established. The robustness index is: ; in, For target device node The class probability corresponding to the predicted class; For target device node Predicting categories in preliminary selection results The corresponding price; For target device node Price constraint ceiling; For target device node In prediction category The following safety / performance index values; For target device node Minimum safety / performance requirements; This is the threshold for consistency determination; This is a set of extreme weather disturbances used for verification. In case of disturbance The predicted safety / performance index value for this category; A consistency score is assigned, with a higher score indicating better consistency under category probability, price constraints, and safety / performance rules. For robustness indicators, This indicates that the minimum requirements are met under all disturbance conditions.
6. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 1, characterized in that, Using the graph representation of the target heterogeneous graph, the constraint biases of the candidate selection list, and robustness indicators to form the state, reinforcement learning is employed to update the policy parameters online, including: The mean vector of the node embedding vectors obtained by the heterogeneous graph neural network for each target device node in the target heterogeneous graph is extracted, and combined with the mean global constraint deviation and mean robustness index of the candidate selection list, and concatenated into a state vector and input into the reinforcement learning environment. In the reinforcement learning environment, forward reasoning is performed on the input state vector based on the current policy network, and the output is an action vector containing the loss weight adjustment and the selection confidence threshold adjustment. Each component of the action vector is added to the current loss weight and the selection confidence threshold to obtain an updated parameter set. The joint optimization and consistency verification process is then re-executed using the updated parameter set to calculate the reward value. The policy network parameters are updated according to the following formula: ; in, Let the objective function of the policy network be denoted as . The number of interactions in a training batch; For the first The reward value obtained from this interaction; In the strategy parameters Below, State Corresponding actions The probability of; For the first The action vector output by the next interaction; For the first The state vector of the second interaction; the reward value Defined as ; in, The improvement in selection accuracy compared to the previous iteration; This represents the increase in the price constraint satisfaction rate. This represents the improvement value of the robustness index; These are the non-negative coefficients used to balance the contribution of the three factors; These are the policy network parameters.
7. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 1, characterized in that, The rewards are determined by the accuracy of equipment selection, satisfaction of price constraints, and robustness to extreme weather conditions. Updated strategy parameters are used to adjust the joint optimization and inference / verification process, iterating repeatedly until convergence criteria are met. The final equipment selection list and key indicators are then output, including: After each round of strategy update, based on the updated loss weights and selection confidence thresholds, the joint optimization step and inference / verification step are re-executed to obtain the selection accuracy, price constraint satisfaction rate and robustness index of the current round. Calculate the reward value and trigger the policy network parameter update according to the following formula: ; in, For the first The reward value of each iteration; For the first The selection accuracy obtained from rounds of iteration; For the first The price constraint satisfaction rate obtained from rounds of iteration; For the first Robustness metrics obtained from rounds of iteration; , , The non-negative balance coefficients contributing to the three improvements; , , These are the values corresponding to the previous iteration; The updated strategy parameters based on the reward value are written back to the joint optimization process to adjust the loss weight, and written back to the inference / verification process to adjust the selection confidence threshold. If continuous In each iteration, the absolute value of the change in the reward value is less than the preset threshold. If the condition is met, then convergence is determined. When the convergence condition is met, the final result is output as the equipment selection list, price constraint deviation, robustness index, and selection accuracy of the current iteration.
8. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 7, characterized in that, The key metrics for the final equipment selection list in the current iteration include: the selection accuracy rate, price constraint satisfaction rate, robustness under extreme weather disturbances, and comprehensive constraint deviation rate of the equipment selection list at the convergence iteration round; wherein, the selection accuracy rate is the proportion of the number of equipment nodes in the final equipment selection list whose predicted category matches the actual category to the total number of equipment nodes; the price constraint satisfaction rate is the proportion of the number of equipment nodes in the final equipment selection list whose price meets the preset price constraint to the total number of equipment nodes; the robustness rate is the minimum ratio of the safety / performance index value of each equipment node to the corresponding minimum requirement value under the preset extreme weather disturbance set; and the comprehensive constraint deviation rate is the mean of the weighted sum of price deviation and safety / performance deviation over the set of equipment nodes.
9. The intelligent selection method for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration as described in claim 1, characterized in that, The sample equipment data includes the model, main design parameters, load-bearing capacity parameters, material information, and price per unit or set of the sample equipment; the sample geographic information data includes the elevation, slope, aspect, landform type, and land use type corresponding to the installation location of the sample equipment; the sample static meteorological data includes the extreme maximum temperature, extreme minimum temperature, average wind speed, maximum wind speed, maximum instantaneous wind speed, annual precipitation, and seasonal distribution of precipitation in the sample equipment installation area for a representative month or season.
10. A smart selection system for transmission line equipment based on heterogeneous graph representation learning and adaptive RL policy collaboration, characterized in that, include: The sample multi-source data acquisition and preprocessing unit is used to acquire sample equipment data, sample geographic information data and sample static meteorological data, and to uniformly encode and standardize the numerical and categorical characteristics of the sample equipment data, the sample geographic information data and the sample static meteorological data to form a structured multi-source data table associated with the equipment sample. The sample heterogeneous graph construction unit is used to generate device nodes, geographic nodes and meteorological nodes respectively based on the device, geographic and meteorological fields in the multi-source data table, and generate device-device edges, device-geographic edges and device-meteorological edges; A heterogeneous graph definition and storage unit is used to determine the device nodes, geographic nodes, meteorological nodes, device-device edges, device-geographic edges, and device-meteorological edges as sample heterogeneous graphs that distinguish relation types; The heterogeneous graph neural network training unit is used to perform typified embedding and parameter sharing of the sample heterogeneous graph based on the heterogeneous graph neural network, relation-aware multi-head attention and multi-scale aggregation, and to construct node self-supervised loss by randomly masking and restoring some node features during the training phase through node embedding enhancement. The joint optimization parameter acquisition unit is used to jointly optimize the comprehensive classification loss with equipment category as the supervision signal, the price loss corresponding to the price constraint, and the node self-supervision loss to obtain the initial model parameters and inference threshold. The target heterogeneity map inference unit is used to construct a corresponding target heterogeneity map based on the target equipment data, target geographic information data and target static meteorological data of the target project, and to infer the target heterogeneity map based on the initial model parameters and the inference threshold to obtain the category probability and preliminary selection results of each target equipment node. The consistency verification and indicator generation unit is used to perform consistency verification on the category probability and the preliminary selection result based on price constraints and safety / performance rules, and generate a candidate selection list and constraint deviation and robustness indicators; The reinforcement learning policy update unit is used to update the policy parameters online using reinforcement learning, based on the graph representation of the target heterogeneous graph, the constraint bias of the candidate selection list, and the robustness index. The policy parameters include at least one or more of the loss weight and the selection confidence threshold. The closed-loop iteration and result output unit is used to provide rewards based on selection accuracy, price constraint satisfaction, and extreme weather robustness. It uses the updated strategy parameters to adjust the joint optimization and inference / verification process, repeating the iteration until the convergence criterion is met, and outputs the final equipment selection list and key indicators.
Citation Information
Patent Citations
Graph neural network-based power grid dispatching decision-making method and large model
CN119294872A
Intelligent monitoring method and system for operation state of electric power system
CN119474804A
Virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning
CN119494521A
Electric power data analysis method and device based on graph calculation
CN119669730A
Power line health state evaluation and prediction method and system based on big data
CN120146319A