Industrial data mining method for power engineering design

By generating topology sets through data mining systems and graph neural networks, and combining topological cognitive entropy values ​​and distribution drift states, the problems of data black holes and cognitive security in power engineering design are solved, achieving dual protection of physical efficiency and cognitive security in power grid design.

CN122113320APending Publication Date: 2026-05-29SHANDONG DINGXIN ENERGY ENG CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG DINGXIN ENERGY ENG CO LTD
Filing Date
2026-02-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In power engineering design, existing technologies are unable to effectively address the data black hole problem caused by missing or damaged drawings, and fail to ensure the cognitive security and physical efficiency of the system during data drift.

Method used

By acquiring historical engineering unstructured data and real-time load flow data through a data mining system, an initial topology set is generated using a graph neural network. Combined with topology cognitive entropy and distribution drift state, the target topology scheme is dynamically determined, and a decision control strategy that balances physical efficiency and cognitive safety is constructed.

Benefits of technology

It effectively fills data gaps, quantifies the complexity of human-computer interaction, ensures that the system pursues the ultimate physical efficiency when the data deviation is small, and prioritizes the simplest structural solution when the deviation is large, thereby reducing system complexity, ensuring cognitive security and logical certainty, and achieving self-improvement and evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113320A_ABST
    Figure CN122113320A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electric power engineering and data processing, in particular to an industrial data mining method for electric power engineering design, comprising a data acquisition step: obtaining historical engineering unstructured data and real-time load flow data; a topology generation step: generating an initial topology set containing multiple candidate architectures by using a graph neural network model; a cognitive entropy evaluation step: calculating the topology cognitive entropy value representing the complexity of human-computer interaction for the candidate architecture; a decision control step: determining the target topology scheme according to the distribution drift state of the real-time load flow data and combining the topology cognitive entropy value; the present application effectively overcomes the data black hole problem of missing underground pipe gallery engineering drawings, and can dynamically adjust and optimize the target according to the data reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of power engineering and data processing technology, specifically to an industrial data mining method for power engineering design. Background Technology

[0002] With the development of digital twin technology and intelligent power engineering, industrial data mining for power engineering design has become an important research direction in the energy field; in order to achieve effective management and operation of the power grid in the target area, the topology generation and decision control of the system have become particularly important. In topology inference and decision-making in power engineering, graph neural network technology has shown great potential due to its powerful implicit architecture inference capabilities. However, the construction of artificial intelligence models relies on a large amount of unstructured historical engineering data. In scenarios such as underground utility tunnels, missing or damaged drawings can cause data black holes, making data cleaning and utilization extremely difficult. In existing dispatching decision-making systems, the theoretical maximization of power supply capacity is often pursued in isolation, ignoring the distribution and drift of real-time load flow data relative to historical data. Furthermore, the psychological burden on human dispatchers caused by complex topology structures under extreme fault conditions is not considered, making it difficult to ensure the cognitive security of the system when data confidence declines. Therefore, how to mine and process existing historical engineering unstructured data and real-time load flow data, calculate the topological cognitive entropy value that characterizes the complexity of human-computer interaction, and dynamically determine the target topology scheme based on the data drift state in order to construct a decision control strategy that takes into account both physical efficiency and cognitive safety is crucial to ensuring the safe and efficient operation of power engineering systems. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention provides an industrial data mining method for power engineering design. Specifically, the technical solution of this invention includes: The data mining system used in this method includes a data acquisition end, a topology generation end, a cognitive entropy evaluation end, and a decision control end. The method includes: the data acquisition end acquiring historical unstructured engineering data and real-time load flow data of the target area; the topology generation end generating an initial topology set containing multiple candidate architectures based on the historical unstructured engineering data and using a preset graph neural network model; the cognitive entropy evaluation end calculating a topology cognitive entropy value representing the complexity of human-computer interaction for each candidate architecture in the initial topology set; and the decision control end determining a target topology scheme from the initial topology set based on the distribution drift state of the real-time load flow data relative to the historical unstructured engineering data and the topology cognitive entropy value.

[0004] Preferably, determining the target topology scheme from the initial topology set based on the distribution drift state of the real-time load flow data relative to the historical engineering unstructured data, combined with the topology cognitive entropy value, includes: when the data deviation value represented by the distribution drift state is less than a preset drift threshold, the decision control terminal filters a first target topology from the initial topology set based on a preset power supply capacity maximization criterion; the decision control terminal determines the first target topology as the target topology scheme.

[0005] Preferably, determining the target topology scheme from the initial topology set based on the distribution drift state of the real-time load flow data relative to the historical engineering unstructured data, combined with the topology cognitive entropy value, includes: when the data deviation value represented by the distribution drift state is greater than or equal to the drift threshold, the decision control terminal filters a second target topology from the initial topology set based on the criterion of minimizing the topology cognitive entropy value; the decision control terminal determines the second target topology as the target topology scheme.

[0006] Preferably, the method further includes: the decision control terminal, upon determining the second target topology, generating a dimensionality reduction control instruction sequence that matches the second target topology; and the decision control terminal sending the dimensionality reduction control instruction sequence to a remote control device of the target area.

[0007] Preferably, the method further includes: the remote control device receiving the dimensionality reduction control command sequence; the remote control device performing switching operations on the switching equipment in the target area according to the dimensionality reduction control command sequence to construct a physical network.

[0008] Preferably, the step of calculating the topological cognitive entropy value, which characterizes the complexity of human-computer interaction, for each candidate architecture in the initial topology set includes: the cognitive entropy evaluation end extracting the node connection relationship features and operation logic depth features of the candidate architecture; and the cognitive entropy evaluation end calculating the topological cognitive entropy value based on the node connection relationship features and the operation logic depth features using a preset entropy weight algorithm.

[0009] Preferably, the decision control terminal, based on the distribution drift state of the real-time load flow data relative to the historical engineering unstructured data, includes: the decision control terminal calculating the statistical distribution characteristics of the real-time load flow data; the decision control terminal calculating the divergence value between the statistical distribution characteristics and the baseline distribution characteristics corresponding to the historical engineering unstructured data, and determining the divergence value as the distribution drift state.

[0010] Preferably, the method further includes: after the remote control device completes the switching operation, the decision control terminal collects the steady-state feedback data of the target area; the decision control terminal corrects the weight parameters of the graph neural network model based on the steady-state feedback data.

[0011] Compared with the prior art, the present invention has the following beneficial effects: 1. This method acquires historical unstructured engineering data of the target area and drives a pre-defined graph neural network model to perform implicit architecture inference. Combined with constraint sampling operations based on spectral graph theory, it generates a physically feasible initial topology set, which can effectively fill the data gaps caused by missing or damaged drawings and establish a digital mapping of the physical world. By calculating the topological cognitive entropy value representing the complexity of human-computer interaction for each candidate architecture, the abstract psychological cognitive load is quantified into a mathematical index that includes node connection relationships and operational logic depth. This provides a quantitative basis for evaluating the information processing capacity of human dispatchers in power grid design, realizing a dual consideration of physical performance and human-computer interaction friendliness. 2. This method quantifies the distribution drift state by calculating the divergence between the statistical distribution characteristics of real-time load flow data and the historical baseline distribution characteristics. It can keenly capture the degree of deviation of the current power grid operating environment from historical experience. By establishing an adversarial mechanism between engineering design and operation and maintenance cognition, it pursues the ultimate physical efficiency based on the principle of maximizing power supply capacity when the data deviation is less than a preset threshold, and prioritizes the simple and intuitive solution based on the principle of minimizing topological cognitive entropy when the data deviation is large. This ensures that the system complexity is actively reduced when facing unknown risks or attacks, and protects the cognitive safety and intervention capabilities of human dispatchers in emergency situations. 3. This method generates a dimensionality-reduced control instruction sequence by using state difference matrix analysis after determining the target topology, and uses an instruction sorting algorithm to force the closing instruction to be placed before the opening instruction. This transforms the complex network topology reconstruction process into a safe and orderly operation flow. Through this dimensionality-reduced control strategy, the state space of the physical network is mapped from a high-dimensional coupled state with multiple degrees of freedom to a low-dimensional radial state, which significantly reduces the cyclic complexity of the system, ensures the logical determinism from digital decision-making to physical execution, and prevents cascading failures caused by excessive system complexity. 4. This method collects steady-state feedback data after the switching operation is completed by the remote control device, and uses state estimation and topology error identification techniques to construct a truth matrix reflecting the real connection relationship, which can physically verify the prediction results of the graph neural network. By correcting the weight parameters of the graph neural network model based on the feedback data, a closed-loop learning system is constructed, which enables the model to continuously absorb real feedback from the physical world to correct the initial deviation of historical data. This achieves self-improvement and evolution of the system during long-term operation and continuously improves the accuracy of topology generation. Attached Figure Description

[0012] The present invention will be further explained below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0014] Example 1: Please see Figure 1 An industrial data mining method for power engineering design is proposed. The data mining system applied by the method includes a data acquisition end, a topology generation end, a cognitive entropy evaluation end, and a decision control end. The method includes: the data acquisition end acquiring historical unstructured engineering data and real-time load flow data of the target area; the topology generation end generating an initial topology set containing multiple candidate architectures based on the historical unstructured engineering data and using a preset graph neural network model; the cognitive entropy evaluation end calculating the topology cognitive entropy value, which represents the complexity of human-computer interaction, for each candidate architecture in the initial topology set; and the decision control end determining the target topology scheme from the initial topology set based on the distribution drift state of the real-time load flow data relative to the historical unstructured engineering data and the topology cognitive entropy value.

[0015] This embodiment details the specific operational logic of the aforementioned industrial data mining method. This method operates within a system comprised of a data acquisition terminal, a topology generation terminal, a cognitive entropy evaluation terminal, and a decision control terminal. Its aim is to establish a digital mapping of the physical world and overcome the data black hole problem caused by missing underground utility tunnel engineering drawings. The data acquisition terminal performs historical data acquisition operations, specifically including scanning damaged or poorly labeled paper engineering drawings, CAD fragments, and historical maintenance logs. This weakly labeled data, containing a large amount of noise, constitutes the system's basic knowledge base. Simultaneously, this terminal collects voltage, current phasor, and power flow data through intelligent sensing terminals deployed at key nodes of the power distribution network, such as miniature synchronous phasor measurement units, to obtain real-time load flow data at the current moment. The topology generation module utilizes artificial intelligence to fill in the gaps in historical data. Based on unstructured historical engineering data, it drives a pre-defined graph neural network model to perform implicit architecture inference. To construct the graph structure input required for graph neural network operations, the system performs a preprocessing step of latent adjacency graph construction: calculating the Euclidean distance between nodes using the spatial coordinate vectors of the drawing fragments, and employing the K-Nearest Neighbors (K-NN) algorithm, where the number of neighbors is set. For each node, an initial physical connection assumption is established to form a base graph structure BaseGraph that carries the transmission of features. For the edges in this base graph, if there are clear line records in the historical fragment data, their historical impedance values ​​are mapped as edge features. If not, high impedance initial values ​​are assigned based on spatial distance to simulate potential connections. Building upon this, the graph neural network model employs the relational graph convolutional network R-GCN architecture, with the input layer constructing a high-dimensional feature tensor: the normalized spatial coordinate vector of the drawing fragments. The vector is concatenated with the One-Hot encoded vector of the device semantic category and used as the input feature vector of the node. Simultaneously extract the resistance of historical circuits. With reactance Composition of edge feature vectors The model contains two hidden layers using the ReLU activation function, and the output layer uses the Sigmoid function to predict the probability matrix of connections between potential node pairs. To transform the continuous probability matrix into a discrete and physically feasible initial topology set, the system performs a constraint sampling operation based on spectral graph theory: Each element is sampled independently multiple times by Bernoulli to generate the original set of adjacency matrices; For each sampled adjacency matrix Calculate its Laplace matrix. ,in, Given the degree matrix, solve for... The system retains only the eigenvalues ​​of those eigenvalues, where the number of 0s is exactly 1, to ensure that the generated topology is simply connected. Meanwhile, to avoid the risk of infinite loops caused by sparse probability distribution, the system sets a maximum number of sampling attempts. If a sufficient number of connected topologies are not generated after reaching this number of iterations, the system will initiate a deterministic completion mechanism: the probability matrix will be... reciprocal As edge weights, the Prim algorithm is invoked to force the construction of a minimum spanning tree (MST) to ensure that each topology in the output set is physically connected. Depth-first search (DFS) is used to detect loops and filter out illegal short-circuit loop structures other than the pre-allowed connecting loop networks. The filtered valid topologies are then used to form an initial topology set. This model learns the primitive features and spatial relationships in historical drawings to infer the potential node connection probabilities, thereby generating an initial topology set that is feasible within the constraints of physical space, covering various possible connection methods such as single-ring networks and piezoresistant networks. Based on this, the cognitive entropy assessment end quantifies the psychological load on human dispatchers caused by the design scheme. For each candidate architecture in the set, the topological cognitive entropy value is calculated. This entropy value, as a dimensionless index, measures the amount of information processing required for human dispatchers to understand the topology and formulate the correct switching strategy when a power grid failure occurs. As the central hub of the system, the decision-making and control terminal establishes a counter-mechanism between engineering design and operation and maintenance cognition. Based on the distribution and drift of real-time load flow data relative to historical data, it dynamically combines the topology cognition entropy value to determine the target topology scheme from the initial topology set. This process no longer simply pursues the theoretical maximization of power supply capacity, but dynamically adjusts and optimizes the target based on the reliability of the data.

[0016] Example 2: Based on the distribution drift state of real-time load flow data relative to historical engineering unstructured data, and combined with the topology cognitive entropy value, a target topology scheme is determined from the initial topology set. This includes: when the data deviation value represented by the distribution drift state is less than a preset drift threshold, the decision control terminal selects a first target topology from the initial topology set based on a preset power supply capacity maximization criterion; and the decision control terminal determines the first target topology as the target topology scheme.

[0017] This embodiment details the decision-making logic when the data environment is relatively stable; the decision control terminal performs drift state determination, checking whether the data deviation value representing the distribution drift state is less than a preset drift threshold; this threshold... Set to 3 times the standard deviation of the mean of the historical normal operation data distribution divergence. This ensures that aggressive optimization is triggered only under extremely high confidence levels. When the deviation is less than this value, it indicates that the historical data is basically consistent with the current physical reality, meaning that the data confidence level is high. In response to the data deviation value being less than this threshold, the system determines that the environment is predictable and then enters the aggressive optimization mode. The decision control end selects the first target topology from the initial topology set based on the preset power supply capacity maximization criterion. Specifically, the implementation process of this criterion is as follows: the system performs power flow simulation on each architecture in the initial topology set, uses the Newton-Raphson method to iteratively solve the power balance equation, and obtains the actual load current of each branch. Simultaneously, the system retrieves the physical parameters of each branch conductor from the equipment asset database, including conductor diameter. Surface heat absorption coefficient and emission coefficient The IEEE 738 standard heat balance equation is invoked; specifically, this equation establishes the balance between heat gain and loss in the conductor: ,in, For convective heat dissipation power, For radiative heat dissipation power, The solar heat absorption power, For the conductor at the maximum allowable operating temperature AC resistance below; The system performs the calculation based on the following formula: Calculate radiative heat dissipation. Here, the constant coefficient 0.0178 is a conversion correction value for the Stefan-Boltzmann constant in engineering units, and its physical source is... ,in, Designed to adapt to wire diameter The calculation of radiated power is performed using millimeters (mm) as the unit to ensure the accuracy of the results. The dimension is watts per meter (W / m); calculate convective heat dissipation. The system calculates the natural convection heat dissipation power based on the simplified formula recommended by IEEE 738. , The constant coefficient 0.0205 is derived from the empirical Nusselt number correlation of natural convection heat dissipation in a horizontal cylinder. This coefficient has pre-set standard gravitational acceleration and fluid thermophysical property parameters for use in situations involving wire diameter... For millimeters, air density When the value is kilograms per cubic meter, the convective heat dissipation power can be directly calculated; compared with the forced convection heat dissipation power... ; It should be noted that, in order to correct the Reynolds number term in the formula Due to the dimensional imbalance caused by the unit of measurement of the parameters, the system automatically converts the wire diameter (unit: millimeter) to mm before inputting it into the calculation. Converted to meters (m) to ensure that the combined quantity is a dimensionless value Re, thereby guaranteeing the physical correctness of subsequent exponentiation operations; where, At the membrane temperature The air density, viscosity, and thermal conductivity are obtained from the table below. The larger value selected after comparison is taken as the final convective heat dissipation power. Substitute into subsequent calculations; Using formula Solve this problem; based on this, the system combines the real-time ambient temperature obtained from the data acquisition terminal. Wind speed And real-time solar intensity collected by the light sensor integrated into the intelligent sensing terminal. And according to the formula Convert it to line-average heat absorption power Real-time calculation of the dynamic thermal rating of equipment on each line. ; Calculate the current margin index of critical branches In order to accurately reflect the bottleneck effect of the architecture, The operands of the function are explicitly defined as the set of all critical branches in the current candidate architecture. This refers to the set of lines in the system that are under heavy load or connected to core nodes. To eliminate ambiguity in the technical definition, this embodiment clearly defines... The specific rules for composition are as follows: If the real-time load rate of a certain line exceeds 75% (based on the line attribute definition), or if any endpoint connected to the line belongs to the top 20% of nodes in the entire network in terms of betweenness centrality (based on the node attribute definition), then the line is determined to be a critical branch and included in the set. The system iterates through every branch in the set. Calculate their respective margin values ​​and take the minimum value among them. ; The current margin index is used here to clearly indicate that it refers to current carrying capacity rather than power capacity, and the calculation result is a dimensionless value, representing the proportion of the equipment's remaining current carrying capacity; the system screens out The architecture with the highest value is selected as the first target topology. This criterion aims to prioritize the topology that can maximize the utilization of the device's dynamic thermal rating and carry the maximum load current, while meeting the N-1 safety criterion. In this case, the system is allowed to select a complex but highly efficient solution, such as a complex flexible interconnect structure. Finally, the decision control terminal locks the first target topology as the target topology solution.

[0018] Example 3: Based on the distribution drift state of real-time load flow data relative to historical engineering unstructured data, and combined with the topology cognitive entropy value, a target topology scheme is determined from the initial topology set. This includes: when the data deviation value represented by the distribution drift state is greater than or equal to the drift threshold, the decision control end selects a second target topology from the initial topology set based on the criterion of minimizing the topology cognitive entropy value; the decision control end determines the second target topology as the target topology scheme.

[0019] This embodiment details the defensive decision-making logic when the data environment experiences severe disturbances. The system performs a high-risk drift judgment, detecting that the data deviation value representing the distribution drift state is greater than or equal to the drift threshold. This situation typically corresponds to asymmetric interference attacks or large-scale distributed energy, such as the disorderly access of electric vehicle supercharging stations, causing the patterns mined from historical data to become invalid and the confidence of AI prediction models to drop sharply. In response to this high-risk state, the system immediately switches to a safe mode. The decision control end selects a second target topology from the initial topology set based on the criterion of minimizing the topology cognitive entropy value. The core of this criterion is to prioritize the topology structure that is the simplest in structure, the most intuitive in logic, and the most in line with traditional human operation and maintenance habits, such as a standard radial network, even if the power supply capacity of this scheme may be lower than the theoretical optimal value. Finally, the decision control end determines the second target topology as the target topology scheme. This embodiment establishes an active defense mechanism. When severe data drift is detected, the system actively abandons the pursuit of extreme physical efficiency and instead pursues cognitive security. By forcibly selecting a low-entropy topology, it ensures that when the system faces unknown extreme failure risks, human schedulers can quickly understand the network structure and intervene within the golden window of opportunity based on intuition, preventing a chain reaction of collapses caused by excessive system complexity.

[0020] Example 4: The method also includes: when the decision control terminal determines the second target topology, it generates a dimensionality reduction control command sequence that matches the second target topology; the decision control terminal sends the dimensionality reduction control command sequence to the remote control device of the target area.

[0021] The method also includes: a remote control device receiving a sequence of dimensionality reduction control instructions; and the remote control device performing switching operations on the switching equipment in the target area according to the sequence of dimensionality reduction control instructions to construct a physical network.

[0022] This embodiment describes a closed-loop control process from digital mining decision-making to physical world execution. After determining to adopt a low-entropy second target topology, the decision control end executes an instruction generation step to generate a matching dimensionality-reduced control instruction sequence. This generation process specifically employs the state difference matrix analysis method. Let the adjacency matrix of the current physical network be... The adjacency matrix of the second target topology is Both are A 0-1 matrix; where the adjacency matrix of the current physical network is... Instead of static preset values, the real-time connection state matrix is ​​generated by the decision control terminal reading the latest network-wide switch position telesignal signals stored in the data acquisition terminal and parsing them using a topology coloring algorithm. This ensures that the benchmark for differential calculations is strictly synchronized with the physical reality. The system calculates the differential matrix. Traversing the matrix Identify the differences in the upper triangular elements: If Marked as a closed switch The Add operation command; if Marked as disconnect switch The Cut instruction; The system executes a strict instruction sequencing algorithm to replace general logic library calls: an instruction priority queue is established, all Add operation instructions are forced to be placed at the head of the queue, all Cut operation instructions are placed at the tail of the queue, and a pre-set duration of loop current detection waiting instructions is inserted between the Add and Cut instruction groups; this sequence consists of a set of specific switching action commands, the physical meaning of which is to reconstruct the current complex network topology into a simple structure defined by the second target topology. It is called dimensionality reduction because it significantly reduces the operational freedom and complexity of the system. Specifically, dimensionality reduction here is technically defined as reducing the cyclomatic complexity of the network topology, that is, mapping the state space of the physical network from a high-dimensional coupled state with multiple closed loops, i.e., corresponding to loop currents with multiple degrees of freedom, to a low-dimensional radial state without loops, i.e., corresponding to the current distribution uniquely determined by Kirchhoff's laws, thereby mathematically minimizing the dimension of independent control variables; The decision control terminal sends the instruction sequence to the remote control device in the target area, such as the power distribution automation terminal, through an encrypted communication channel; the remote control device receives the instruction and executes physical actions, thereby performing switching operations on switching equipment such as high-voltage circuit breakers and load switches; as the switching operation is completed, the physical network is reconstructed into the target topology.

[0023] Example 5: For each candidate architecture in the initial topology set, the topology cognitive entropy value, which characterizes the complexity of human-computer interaction, is calculated. This includes: the cognitive entropy evaluation end extracting the node connection relationship features and operation logic depth features of the candidate architecture; and the cognitive entropy evaluation end calculating the topology cognitive entropy value based on the node connection relationship features and operation logic depth features using a preset entropy weight algorithm.

[0024] This embodiment details the specific calculation model for topological cognitive entropy. The cognitive entropy evaluation stage performs a feature extraction step, extracting node connection relationship features reflecting the node degree distribution and interconnection density in the network from the candidate architecture. Specifically, these features include the betweenness centrality index of nodes and the operational logic depth feature reflecting the number of switching operation stages required from the power source point to the load point. Based on this, the topological cognitive entropy value is calculated using a preset entropy weight algorithm. The algorithm is defined as a structure-logic composite weighted calculation process, aiming to integrate the abstract topological disorder degree with the specific operational logic into a single quantitative index through weighted summation. To quantify the complexity of human-computer interaction, this embodiment constructs the following calculation model: , in, Topological cognitive entropy, derived from model calculations, physically measures the amount of information processing required for a human scheduler to understand the topology. :No. The probability distribution of structural importance of each node is used to avoid logarithmic calculation anomalies caused by the zero betweenness centrality of the terminal nodes. , defined here The normalized weights after Laplace smoothing: ,in, The betweenness centrality of this node in the candidate architecture is defined. To clarify the calculation method of this metric and meet the requirement of sufficient disclosure, this embodiment specifically adopts the betweenness centrality algorithm based on the shortest path for calculation: ,in, For any two nodes in the network and The total number of shortest paths between them. For the nodes passed through in these shortest paths The number of paths; if If the value is zero, the item is recorded as 0. This indicator physically quantifies the criticality of a node as a network hub in power flow regulation. The total number of nodes. The preset smoothing factor has a value of [value missing]. ; This metric serves as a static measure of the potential for a node to receive attention in the current topology due to its central position; among which, Representing the The betweenness centrality value of each node in the entire network; the system performs a normalization check before calculation to ensure... This is to satisfy the probability space constraints of information entropy calculation; this probability distribution The introduction of entropy calculation enables the network structure to be effectively distinguished. When the network is highly centralized, such as star-shaped, the probability distribution is extremely uneven, resulting in a low entropy value, which is consistent with the psychological fact that humans have a low cognitive load for such structures. The average number of logical judgment steps under the current candidate architecture is precisely calculated using a graph theory search algorithm: the candidate architecture is modeled with the substation as the root node. Given a directed graph, traverse the set of all end load nodes in the graph. The root node is calculated using breadth-first search (BFS). To each end node Shortest path length ,but ; This indicator physically represents the average number of switch levels that a dispatcher needs to confirm when tracing a fault from the power supply side to the fault point when isolating a fault. The theoretical minimum number of logical decision steps is derived from the baseline data of a standard radial network, typically set as the average depth of an ideal single radial line. To ensure the uniqueness and calculability of the baseline value, this embodiment defines it as the current network node size. The ideal balanced binary tree depth is, i.e. , used to characterize the shortest physical logical depth under the premise of satisfying connectivity; The cognitive weighting coefficient is obtained based on historical cognitive load experimental data: the system pre-constructs a test sample set containing different topologies and logical depths, and records the fault judgment time of multiple senior schedulers under each sample. Establish a multiple linear regression model , Solving the regression coefficients using the least squares method and and the ratio Determined as In this embodiment, the calculated ratio converges to 0.5. Therefore, in practical applications, this coefficient is directly set to a constant of 0.5 to adjust the weight of the impact of logic depth on the overall entropy value. Regarding concerns about the rationality of the model construction, this embodiment specifically points out: [The coefficient is missing from the original text.] It acts as a dimensional mapping operator for cognition and information in the formula; It is the dimensionless logic depth ratio Converted to the equivalent unit of information entropy, namely Shannon entropy, the latter half of the formula is given the same physical meaning as the former half, namely, representing the comprehensive cognitive energy consumed by the human brain when analyzing topology, thus ensuring the physical consistency of linear superposition. The total number of nodes in the candidate architecture, along with the upper limit and probability distribution in the summation formula. The domain of definition must remain consistent to ensure the integrity of information entropy calculation; The natural logarithm operator, denotes the natural constant. Logarithmic operations with base 0; This embodiment quantifies psychological cognitive load into a mathematical index using the above formula. The first half of the formula measures the uncertainty of operation, while the second half introduces a structural complexity penalty term. This calculation model not only considers the physical connection of the topology itself, but also deeply integrates the length of the human thought chain when dealing with faults. The smaller the calculated entropy value, the more human-friendly the topology is and the lower the operational risk, thus providing a quantitative basis for human-machine-friendly power grid design.

[0025] Example 6: The decision control terminal, based on the distribution drift state of real-time load flow data relative to historical unstructured engineering data, includes: calculating the statistical distribution characteristics of real-time load flow data; calculating the divergence value between the statistical distribution characteristics and the baseline distribution characteristics corresponding to historical unstructured engineering data, and determining the divergence value as the distribution drift state.

[0026] This embodiment details the quantification method for distribution drift state; the decision control terminal performs statistical distribution characteristic calculations to construct probability density functions for real-time load flow data and baseline load data derived from historical unstructured engineering data, respectively; to address the mathematical limitation that probability distributions cannot be directly constructed from single-moment data points, the system maintains a length of [missing information] in memory. The system employs a first-in, first-out (FIFO) sliding window queue. Specifically, it calculates the statistical distribution characteristics of real-time load flow data: extracting the total active power time series data from the target area's gate table within the sliding window queue to achieve scalar aggregation of high-dimensional data; and setting the rules for dividing feature intervals: extracting the extreme value range of the load in this area from the historical database. Divide it into Equally spaced discrete intervals; perform boundary clamping operation before statistics: if the power value of a certain sampling point... If, then it is forcibly assigned to the first feature interval; if Then it will be forcibly classified into the first category. A set of characteristic intervals are used to prevent indexing errors caused by data going out of bounds; The frequency of power data falling into each interval within the sliding window is statistically analyzed and normalized to generate a real-time probability distribution function that reflects the data fluctuation pattern within the current time window. To quantify the degree of deviation between the two, the system calculates the divergence value between the statistical distribution characteristics and the baseline distribution characteristics, and determines this divergence value as the distribution drift state; in this step, the baseline distribution characteristics... It is generated in advance by the system based on load records after cleaning massive amounts of historical unstructured data; Specifically, to achieve the transformation from unstructured data to probability distribution vectors, the system incorporates an NLP parsing module: it performs OCR recognition on historical maintenance logs, uses regular expressions to extract numerical entities after keywords such as load peak and current readings, and reconstructs historical time-series data by combining log timestamps; after removing data with deviations exceeding a certain threshold... After processing the noise data, strictly adhere to the same extreme value range as described above. sum of interval numbers Statistical analysis and normalization are performed to generate a static probability vector. The data is stored in the system's resident memory to ensure strict alignment of the physical dimensions during divergence calculation. This embodiment uses the discretized form of Kullback-Leibler divergence for calculation. , in, The data deviation value representing the distribution drift state is derived from real-time calculation and its physical meaning is the degree of deviation of the current data distribution from historical experience. Real-time load flow data in the first The probability of occurrence in each feature interval is derived from real-time monitoring data statistics based on sliding time window data; it should be noted that during the calculation process, if Then, according to the limit property, this term is defined. Set to 0 to avoid numerical calculation errors; Historical benchmark data in the first The probability of occurrence in each feature interval is derived from a static vector pre-set in the historical database; The smoothing coefficient, whose value is determined based on the data acquisition accuracy, aims to eliminate computational singularities; specifically, it is the minimum quantization step size for the system to read real-time load stream data. For example, sensor resolution, setting ; In this embodiment, the value is taken as The physical meaning is to prevent numerical calculations with a denominator of zero from being abnormal, and to avoid interfering with the effective probability distribution. The total number of characteristic intervals; The symbol for the natural logarithm, whose base is a mathematical constant. ; This embodiment calculates the divergence value, enabling the system to keenly capture subtle changes in the real-time data distribution relative to historical experience. When the divergence value increases significantly, it provides a clear mathematical signal indicating that the current power grid is operating in an unknown region lacking historical experience support. This provides a precise and objective triggering condition for the system to switch from pursuing efficiency to pursuing safety.

[0027] Example 7: The method also includes: after the remote control device completes the switching operation, the decision control terminal collects the steady-state feedback data of the target area; the decision control terminal corrects the weight parameters of the graph neural network model based on the steady-state feedback data.

[0028] This embodiment describes the system's self-evolution mechanism; after the remote control device completes the switching operation, the decision control terminal executes the feedback acquisition step to collect the steady-state feedback data of the target area, specifically including the voltage recovery curve and power flow distribution after the switching. The system enters the model correction phase, adjusting the weight parameters of the graph neural network model based on steady-state feedback data. Specifically, the system uses weighted least squares (WLS) to estimate the state of the collected voltage and power flow data, and constructs a truth adjacency matrix reflecting the true connectivity of the current physical network using a topology error identification algorithm. Normalized residuals of measurements after system state estimation ,in, The nominal error standard deviation of the instrument corresponding to the measurement point is taken from the equipment ledger database, such as 0.2%. If the maximum normalized residual exceeds the preset bad data threshold, such as 3.0, the corresponding switch status is determined to be inconsistent with the model, and the switch status is reversed, i.e., 0 is set to 1 or 1 is set to 0. In order to prevent logical dead loops caused by multiple bad data, the system sets a maximum number of correction iterations. In each iteration, only the switch associated with the single measurement point with the largest residual is selected for state reversal. If the residual does not converge to below the threshold after the next iteration, the set of feedback data is marked as invalid and no further model update is performed; until the residual converges, the true physical topology is determined. The system will output the prediction probability matrix from the graph neural network. With the truth matrix By comparison, a binary cross-entropy loss function is constructed: The gradient of the loss function with respect to the model weights is calculated using the backpropagation algorithm, and the parameters of the graph neural network are updated using the Adam optimizer. This embodiment constructs a closed-loop learning system. By continuously absorbing physical feedback from real-time switching operations, the graph neural network model can gradually correct the initial deviation caused by errors in historical drawings. This mechanism enables the generated candidate architecture to increasingly approximate the real physical utility tunnel environment, thereby continuously improving the accuracy of topology generation in subsequent excavation processes and realizing the system's self-improvement and evolution during operation.

[0029] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An industrial data mining method for power engineering design, characterized in that, The data mining system used in the method includes a data acquisition end, a topology generation end, a cognitive entropy evaluation end, and a decision control end. The method includes: the data acquisition end acquiring historical unstructured engineering data and real-time load flow data of the target area; the topology generation end generating an initial topology set containing multiple candidate architectures based on the historical unstructured engineering data and using a preset graph neural network model; the cognitive entropy evaluation end calculating the topology cognitive entropy value, which characterizes the complexity of human-computer interaction, for each candidate architecture in the initial topology set; and the decision control end determining the target topology scheme from the initial topology set based on the distribution drift state of the real-time load flow data relative to the historical unstructured engineering data and the topology cognitive entropy value.

2. The industrial data mining method for power engineering design according to claim 1, characterized in that, The step of determining a target topology scheme from the initial topology set based on the distribution drift state of the real-time load flow data relative to the historical engineering unstructured data, combined with the topology cognitive entropy value, includes: when the data deviation value represented by the distribution drift state is less than a preset drift threshold, the decision control terminal filters a first target topology from the initial topology set based on a preset power supply capacity maximization criterion; the decision control terminal determines the first target topology as the target topology scheme.

3. The industrial data mining method for power engineering design according to claim 2, characterized in that, The step of determining a target topology scheme from the initial topology set based on the distribution drift state of the real-time load flow data relative to the historical engineering unstructured data, combined with the topology cognitive entropy value, includes: when the data deviation value represented by the distribution drift state is greater than or equal to the drift threshold, the decision control terminal filters a second target topology from the initial topology set based on the criterion of minimizing the topology cognitive entropy value; the decision control terminal determines the second target topology as the target topology scheme.

4. The industrial data mining method for power engineering design according to claim 3, characterized in that, The method further includes: when the decision control terminal determines the second target topology, generating a dimensionality reduction control instruction sequence that matches the second target topology; and sending the dimensionality reduction control instruction sequence to a remote control device of the target area.

5. The industrial data mining method for power engineering design according to claim 4, characterized in that, The method further includes: the remote control device receiving the dimensionality reduction control command sequence; the remote control device performing switching operations on the switching equipment in the target area according to the dimensionality reduction control command sequence to construct a physical network.

6. The industrial data mining method for power engineering design according to claim 1, characterized in that, The step of calculating the topological cognitive entropy value, which characterizes the complexity of human-computer interaction, for each candidate architecture in the initial topology set includes: the cognitive entropy evaluation end extracting the node connection relationship features and operation logic depth features of the candidate architecture; and the cognitive entropy evaluation end calculating the topological cognitive entropy value based on the node connection relationship features and the operation logic depth features using a preset entropy weight algorithm.

7. The industrial data mining method for power engineering design according to claim 1, characterized in that, The decision control terminal, based on the distribution drift state of the real-time load flow data relative to the historical unstructured engineering data, includes: the decision control terminal calculating the statistical distribution characteristics of the real-time load flow data; the decision control terminal calculating the divergence value between the statistical distribution characteristics and the baseline distribution characteristics corresponding to the historical unstructured engineering data, and determining the divergence value as the distribution drift state.

8. The industrial data mining method for power engineering design according to claim 5, characterized in that, The method further includes: after the remote control device completes the switching operation, the decision control terminal collects the steady-state feedback data of the target area; the decision control terminal corrects the weight parameters of the graph neural network model based on the steady-state feedback data.