Data association analysis method and platform for thermal power production index management

By collecting equipment status data within thermal power plants, calculating transfer entropy, and constructing equipment relationship maps, the problem of unquantified influence among multiple devices in thermal power production indicator management was solved, enabling precise assessment and effective control among devices.

CN121810087APending Publication Date: 2026-04-07BEIJING HUADIAN TIANREN ELECTRIC POWER CONTROL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional thermal power production indicator management fails to effectively handle the mutual influence between multiple devices, resulting in a lack of data correlation and difficulty in locating the root cause of indicator anomalies.

Method used

By collecting operational status data from multiple devices within a thermal power plant, the transfer entropy is calculated, a transfer entropy distribution network is established, the interaction types and intensity between devices are defined, a dynamic device relationship graph is constructed, and production indicators are regulated and optimized.

Benefits of technology

It has achieved the accuracy of equipment correlation analysis and the high efficiency and reliability of indicator control in the process of thermal power production indicator management, accurately located the root cause of anomalies, and optimized control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121810087A_ABST
    Figure CN121810087A_ABST
Patent Text Reader

Abstract

The invention discloses a data association analysis method and platform for thermal power production index management, and relates to the technical field of industrial data mining, and the method comprises the steps: collecting the operation state data of a plurality of devices in a thermal power plant, calculating the transfer entropy between any two devices according to the state time sequence data of each device, and building a transfer entropy distribution network; according to a transfer entropy value and direction between any two devices in the network, interaction relationship types and interaction strength between the devices are defined, a dynamic device relationship graph is constructed, and the interaction relationship types comprise a symbiotic relationship, a competitive relationship and a parasitic relationship; and carrying out production index regulation and control optimization based on the atlas. According to the method, the technical problems that the obtained data is lack of relevance and the source is difficult to locate when the index is abnormal due to the fact that the mutual influence among multiple devices is not effectively processed in the traditional thermal power production index management are solved, and the technical effects of defining the mutual influence among the devices and realizing accurate evaluation and effective regulation and control of the thermal power production index are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial data mining, in particular to a data correlation analysis method and platform for thermal power production index management. BACKGROUND

[0002] Accurate management of thermal power production indexes is crucial for energy consumption control and benefit improvement of enterprises. The existing technology mainly collects single equipment operation data to monitor indexes, and the data processing process focuses on independent parameter analysis without correlating interactive data among multiple devices. However, thermal power production relies on the cooperation of multiple devices, and there are complex dynamic influences among devices. The traditional data processing method cannot quantify the influence direction and strength among devices, resulting in lack of correlation of the obtained data, insufficient accuracy, difficulty in locating the root cause when the index is abnormal, and difficulty in meeting the needs of accurate evaluation and effective regulation of thermal power production indexes. SUMMARY

[0003] The present application provides a data correlation analysis method and platform for thermal power production index management, which solves the technical problems that the mutual influence among multiple devices is not effectively handled in traditional thermal power production index management, resulting in lack of correlation of the obtained data and difficulty in locating the root cause when the index is abnormal.

[0004] In a first aspect, the present application provides a data correlation analysis method for thermal power production index management, which comprises: collecting operation state data of multiple devices in a thermal power plant, calculating the transfer entropy between any two devices according to the state time series data of each device, and establishing a transfer entropy distribution network; defining the interaction relationship type and interaction strength between devices based on the transfer entropy value and direction between any two devices in the transfer entropy distribution network, constructing a dynamic device relationship graph, and the interaction relationship type includes symbiotic relationship, competitive relationship and parasitic relationship; based on the dynamic device relationship graph, production index regulation and optimization are carried out.

[0005] In a second aspect, the present application provides a data correlation analysis platform for thermal power production index management, which comprises: a transfer entropy distribution network construction module for collecting operation state data of multiple devices in a thermal power plant, calculating the transfer entropy between any two devices according to the state time series data of each device, and establishing a transfer entropy distribution network; a device relationship graph construction module for defining the interaction relationship type and interaction strength between devices based on the transfer entropy value and direction between any two devices in the transfer entropy distribution network, constructing a dynamic device relationship graph, and the interaction relationship type includes symbiotic relationship, competitive relationship and parasitic relationship; a production index optimization execution module for carrying out production index regulation and optimization based on the dynamic device relationship graph.

[0006] One or more technical solutions provided in the present application have at least the following technical effects or advantages: The application collects the running state data of multiple devices in the thermal power plant, calculates and establishes a transfer entropy distribution network, defines the interaction relationship of the devices, constructs a dynamic device relationship graph, generates a benchmark graph combined with working condition clustering, compares the benchmark graph with a real-time graph to obtain relationship abnormal transfer features, and matches and adjusts the control action to adjust the operation parameters, so as to accurately locate the abnormal root of the thermal power production index and optimize the control strategy, make the device correlation analysis in the thermal power production index management process more accurate, the index control more efficient and reliable, and achieve the technical effects of clearly defining the mutual influence between devices and realizing accurate evaluation and effective control of the thermal power production index. BRIEF DESCRIPTION OF DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0008] Figure 1 is a flowchart of the data correlation analysis method for thermal power production index management provided by the embodiments of the present application.

[0009] Figure 2 is a structural schematic diagram of the data correlation analysis platform for thermal power production index management provided by the embodiments of the present application.

[0010] The figure mark explanation: transfer entropy distribution network construction module 1, device relationship graph construction module 2, production index optimization execution module 3. DETAILED DESCRIPTION

[0011] The present application provides a data correlation analysis method and platform for thermal power production index management, which solves the technical problem that the mutual influence between multiple devices is not effectively handled in traditional thermal power production index management, resulting in lack of correlation of the obtained data and difficulty in locating the root when the index is abnormal.

[0012] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0013] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices.

[0014] Embodiment one, as shown, a data correlation analysis method for thermal power production index management, wherein the method comprises: Figure 1 Collecting the running state data of a plurality of devices in a thermal power plant, calculating the transfer entropy between any two devices according to the state time series data of each device, and establishing a transfer entropy distribution network.

[0015] In the embodiments of the present application, the transfer entropy is a quantitative measure of the strength of the directed information flow between devices.

[0016] Specifically, first, determine the key devices in the thermal power plant that affect the production index and the corresponding operating parameters that need to be monitored, including the temperature and pressure of the boiler, the speed of the turbine, the power of the generator, the flow of the feed water pump, etc. Then install appropriate sensors at the designated monitoring positions of these key devices. Temperature parameters correspond to temperature sensors, pressure parameters correspond to pressure sensors, speed parameters correspond to speed sensors, and flow parameters correspond to flow sensors.

[0017] The sensors real-time perceive the physical state of the device during operation, and convert the temperature, pressure, speed, flow, etc. into continuous electrical signals. Then the electrical signals output by the sensors are transmitted to the backend processing device through a special data transmission line. The backend processing device filters the received electrical signals to remove the clutter signals caused by electromagnetic interference in the field, and then converts the filtered electrical signals into digital signals recognizable by a computer.

[0018] After that, the digital signals are transmitted to the central data storage system of the thermal power plant through the communication network such as industrial Ethernet. The central data storage system checks the data at the same time of receiving the digital signals, checks whether the data is missing or has abnormal values beyond the normal operating range of the device, marks the abnormal data and supplements the missing data. Finally, continuous and accurate running state data of a plurality of devices in the thermal power plant are obtained, which provides basic data support for subsequent analysis of the correlation between devices.

[0019] ​Next, the state time series data of each device is preprocessed by normalization. The preprocessed state time series data of any two devices are extracted. First, the target embedding dimension and target time delay parameters required for calculating the transfer entropy are analyzed. Then, based on this, the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device are calculated. Finally, using the devices as nodes, directed edges are constructed based on the transfer entropy values ​​and transfer directions, thereby establishing a transfer entropy distribution network. This step will be explained in detail later.

[0020] Based on the transfer entropy value and direction between any two devices in the transfer entropy distribution network, the interaction relationship type and interaction intensity between devices are defined, and a dynamic device relationship graph is constructed. The interaction relationship types include symbiotic relationship, competitive relationship and parasitic relationship.

[0021] Optionally, firstly, the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device are extracted. If both the first and second transfer entropy values ​​are higher than a first preset threshold and are positively correlated, the interaction relationship between the two devices is determined to be a symbiotic relationship, and the minimum transfer entropy value is taken as the symbiotic interaction strength. If both the first and second transfer entropy values ​​are higher than the first preset threshold and are negatively correlated, the interaction relationship between the two devices is determined to be a competitive relationship, and the ratio of the transfer entropy difference to the sum of the transfer entropy differences is calculated as the competitive interaction strength. If the first transfer entropy value is higher than the first preset threshold and the second transfer entropy value is lower than the second preset threshold, the interaction relationship between the two devices is determined to be a parasitic relationship, and the maximum transfer entropy value is taken as the parasitic interaction strength. Finally, the construction of the dynamic device relationship graph is completed. This step will be explained in detail later.

[0022] Based on the dynamic equipment relationship graph, production indicators are adjusted and optimized.

[0023] In one embodiment of this application, historical operating data is first clustered by operating condition to identify multiple typical operating conditions. Next, a corresponding baseline equipment relationship map is calculated for each typical operating condition. The dynamically constructed equipment relationship map is then compared with the baseline equipment relationship map corresponding to the current operating condition to generate abnormal relationship transfer features. Finally, based on the generated abnormal relationship transfer features, production indicators are adjusted and optimized; this step will be described in detail later.

[0024] Furthermore, the method provided in this application embodiment includes: Normalize and preprocess the state time series data of each device; extract the state time series data of any two devices after preprocessing, analyze the target embedding dimension and target time delay parameter required for the transfer entropy, calculate the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device; construct directed edges based on the transfer entropy value and direction, and establish the transfer entropy distribution network with the device as the node.

[0025] Specifically, the first step is to clarify the type and corresponding numerical range of the time-series data for each device's status. Next, a linear normalization method is selected to process the time-series data for each device, filtering out the maximum and minimum values ​​in the dataset. For each data point in this category, a normalized value is obtained by subtracting the minimum value of that category from the value of a single data point, and then dividing by the difference between the maximum and minimum values. This process is repeated for all types of time-series data for all devices, ultimately resulting in normalized data where all device time-series data falls within the 0-1 range, eliminating dimensional and numerical range differences in the status parameters of different devices.

[0026] Next, target devices are determined for both the first and second transfer directions. With an initial embedding dimension m=1, the nearest neighbor of each data point in the temporal data of the target device's state is identified in the m-dimensional state space. Then, the embedding dimension is increased to m+1, and the rate of change of distance between each data point and its nearest neighbor in the previous m-dimensional space is calculated. If this rate of change exceeds a preset threshold, the corresponding nearest neighbor is determined to be a false nearest neighbor. The embedding dimension is then gradually increased, and the above judgment steps are repeated. Simultaneously, the proportion of false nearest neighbors is counted until the proportion of false nearest neighbors decreases to a preset proportion threshold. The embedding dimension at this point is determined as the target embedding dimension required to calculate the transfer entropy for the corresponding transfer direction. This step will be explained in detail later.

[0027] Then, based on the target embedding dimension and target time delay parameter under the corresponding transfer direction, the preprocessed state time series data of any two devices are reconstructed into a vector set containing the source device state vector, the target device historical state vector, and the target device future state vector. Next, based on this vector set, the transfer entropy value from the source device to the target device is calculated using a probabilistic estimation algorithm, and finally the first transfer entropy value and the second transfer entropy value are generated. This step will be explained in detail later.

[0028] Next, each piece of equipment within the power plant is designated as a node in the network, ensuring that each device has an independent identifier within the network. For each pair of devices, a directed edge is constructed based on the calculated transfer entropy value and direction. The transfer entropy value reflects the strength of directed information flow between devices, while the direction clarifies the path of information transmission. If a transfer entropy value exists between two devices, meaning the information transmission strength is greater than 0, a directed edge is drawn between the source and target device nodes. The direction of the edge is determined by the direction of the transfer entropy, i.e., from the source device to the target device. The weight of the edge is determined by the magnitude of the transfer entropy value; the larger the transfer entropy value, the higher the weight, which directly reflects the strength and direction of information transmission between devices.

[0029] Then, all device pairs are traversed. For each pair of devices with a transfer entropy value, directed edges are drawn according to the above rules. For example, if there is a transfer entropy value A between the first device and the second device, and a transfer entropy value B between the second device and the first device, and both A and B are greater than 0, then a directed edge with weight A is drawn between the first device node and the second device node, and a directed edge with weight B is drawn between the second device node and the first device node. If only a unidirectional transfer entropy value exists, then only the directed edge in the corresponding direction is drawn.

[0030] Finally, all equipment nodes and their corresponding directed edges are integrated to form a complete transfer entropy distribution network. This network, through nodes and directed weighted edges, clearly presents the strength and path of directed information associations between equipment in a thermal power plant, providing intuitive network structure support for subsequent equipment association analysis and production indicator management.

[0031] Furthermore, the method provided in this application embodiment includes: For the first and second transfer directions, target devices are determined respectively. Under the initial embedding dimension m, the nearest neighbor of each data point in the state time series data of the target device is identified in the m-dimensional state space, where m=1. The initial embedding dimension is increased to m+1, and the rate of change of the distance between each data point and the nearest neighbor in the m-dimensional space is calculated. If the rate of change exceeds a preset rate of change threshold, the nearest neighbor is determined to be a false nearest neighbor. The embedding dimension is gradually increased and judgment is performed. The proportion of false nearest neighbors is counted until the proportion of false nearest neighbors drops to a preset proportion threshold. The embedding dimension at this time is determined as the target embedding dimension required to calculate the transfer entropy under the corresponding transfer direction.

[0032] Optionally, when analyzing the target embedding dimension required for the transfer entropy analysis, the target device is first determined for each of the two transfer directions. This is a preliminary operation for analyzing bidirectional information transmission. Specifically, the first transfer direction is from the first device to the second device, where the second device is the target device; the second transfer direction is from the second device to the first device, where the first device is the target device. After determining the target device, the initial embedding dimension m is set to 1, where 1 is the lowest dimension, meaning the analysis starts from the most basic state space. In the 1-dimensional state space of m=1, the temporal data of the target device's state is processed, and the Euclidean distance method is used to identify the nearest neighbor of each data point: the Euclidean distance between each data point and all other data points is calculated, and the data point with the smallest distance value is determined as the nearest neighbor of that data point.

[0033] Next, the initial embedding dimension is increased from m to m+1, that is, from 1 to 2. Simultaneously, the state time-series data of the target device is reconstructed dimensionally. During reconstruction, data is combined in the manner of "current time data + historical time data" to form data points in an m+1 dimensional state space. For example, in 2 dimensions, each data point is... τ is the target time delay parameter, which will be further determined later. Here, we will first complete the dimensional reconstruction. After the reconstruction is completed, we will use Euclidean distance again to find the nearest neighbor of each data point in the m+1 dimensional space. Then, we will use the rate of change formula: (m+1 dimensional distance - m dimensional distance) / m dimensional distance to calculate the rate of change of the distance between the m+1 dimensional nearest neighbor and the nearest neighbor in the m dimensional state space. This will intuitively reflect the degree of change in the distance between the nearest neighbors after the increase in dimensionality.

[0034] Next, a preset change rate threshold is set, such as 0.1 commonly used in industrial applications. If the distance change rate of a data point exceeds this threshold, the nearest neighbor of that data point in m-dimensional space is determined to be a false nearest neighbor. An excessively high change rate means that the previous nearest neighbor in m-dimensional space was a false proximity due to insufficient dimension, and does not truly reflect the proximity relationship of the system state. Then, the proportion of false nearest neighbors to all data points of the target device is calculated. If the proportion is higher than a preset proportion threshold, such as 5% commonly used in industrial data processing, the embedding dimension is increased to m+2, and the steps of reconstructing the dimension, finding the nearest neighbor, calculating the change rate, judging false nearest neighbors, and calculating the proportion are repeated. This cycle continues until the proportion of false nearest neighbors drops below the preset proportion threshold. The embedding dimension at this point is the target embedding dimension for the corresponding transition direction. This dimension defines the length of the historical window, i.e., how many past data points are needed to uniquely determine the current state of the system, accurately describing the historical memory length of the device's operating state, which meets the core requirement that the embedding dimension must match the complexity of the system state.

[0035] Through the step-by-step operation based on the pseudo-nearest neighbor method, combined with the targeted setting of the transfer direction of thermal power equipment and the target equipment, the target embedding dimension required for calculating the transfer entropy was determined. This dimension can accurately match the historical memory requirements of the equipment's operating status, providing key parameter support for the subsequent accurate quantification of the transfer entropy of the directed information flow intensity between equipment.

[0036] Furthermore, the method provided in this application embodiment includes: The first transfer direction is from the first device to the second device; the second transfer direction is from the second device to the first device.

[0037] Specifically, in the process of analyzing the target embedding dimension required for the transfer entropy in the aforementioned steps, it is necessary to first clarify the specific directions of the first and second transfer directions. The transfer entropy needs to quantify the information flow intensity in a specific direction. Ambiguity in the direction will lead to confusion in the corresponding devices for subsequent dimension calculations, affecting the accuracy of the analysis.

[0038] Firstly, regarding the first transfer direction, based on the correspondence logic of source device-target device in directed association analysis, it is clearly defined as the direction from the first device to the second device. That is, in this direction, the first device is determined to be the source device for information output, and the second device is the target device for receiving information. This defines the starting and ending devices of information transmission in this direction, providing clear objects for subsequent screening of target devices and calculation of embedding dimensions in this direction.

[0039] Next, for the second transfer direction, according to the symmetric definition rules of bidirectional directed correlation analysis, it is defined as the direction from the second device to the first device. That is, at this time, the second device is the source device for information output, and the first device is the target device for receiving information, forming a complete bidirectional coverage with the first transfer direction, ensuring that all possible directed information transmission paths between the two devices are clearly defined.

[0040] Furthermore, the method provided in this application embodiment includes: Calculate the mutual information between the state time-series data of the target device and its own sequence after different time delays, and plot the function curve of mutual information relative to time delay; determine the time delay corresponding to the first local minimum of the function curve as the target time delay parameter under the corresponding transfer direction.

[0041] In this embodiment of the application, mutual information is a tool for quantifying how much information is contained between two variables or data sequences, and can be used to determine the degree of correlation between the two.

[0042] Specifically, the mutual information between the target equipment's state time series data and its own sequences with different time delays is first calculated. Histograms are a specific means of quantifying the degree of information correlation between two sequences. The specific operation is as follows: First, based on the sampling frequency of the thermal power equipment's state time series data, a reasonable time delay range is set, typically 1 to 30 sampling intervals, covering possible dynamic correlation cycles. For each set time delay parameter τ, the original state time series data of the target equipment is denoted as x(t), and a sequence x(t-τ) with a time delay of τ sampling intervals is generated, so that the original sequence and the time-delayed sequence form a one-to-one data pair (x(t), x(t-τ)).

[0043] Next, the numerical ranges of both the original sequence and the time-delayed sequence are divided into several equal intervals, i.e., the "intervals" used to divide the data range in the histogram. The number of data points in each interval is counted, and the marginal probability distribution p(x) of the original sequence, the marginal probability distribution p(xτ) of the time-delayed sequence, and the joint probability distribution p(x,xτ) of the two are calculated. Finally, the mutual information is substituted into the mutual information calculation formula to obtain the mutual information value corresponding to the time delay τ. The magnitude of the mutual information value reflects the correlation strength between the original sequence and the time-delayed sequence. The higher the value, the stronger the correlation between the two. However, if the time delay is too small, adjacent state vectors may be too close and contain a lot of redundant information. If the value is too low, the adjacent state vectors may lose their dynamic correlation and become discontinuous due to the excessive time delay.

[0044] After calculating the mutual information for all time delays, the next step is to plot the function curve. Using the set time delay value τ as the horizontal axis and the calculated mutual information value as the vertical axis, each corresponding data point for each time delay and mutual information is marked on a Cartesian coordinate system. Then, a smooth curve is used to connect these data points in ascending order of time delay on the horizontal axis, forming a function curve of mutual information relative to time delay. This curve visually demonstrates the impact of time delay changes on the correlation strength between the original and lagged sequences. Peak regions in the curve correspond to high mutual information values, indicating a large amount of redundant information in the sequence at that time delay; valley regions correspond to lower mutual information values. From these, time delay points that reduce redundancy while maintaining dynamic correlation are selected.

[0045] Next, we filter by observing the changing trend of the function curve. Starting from the end with the smallest time delay on the left side of the horizontal axis, we trace the curve's trend along the direction of increasing time delay: when the curve changes from a continuously declining trend to its first rise, this turning point is the first local minimum of the curve. The time delay corresponding to this local minimum is determined as the target time delay parameter in the corresponding transition direction. If the time delay is less than this value, the curve is still in a high mutual information region, which will lead to a large amount of redundant information when constructing the state vector later. If a subsequent local minimum is selected, the corresponding time delay will be too large, causing adjacent state vectors to lose their dynamic correlation and become discontinuous. The first local minimum balances the requirements of redundant information and the continuity of dynamic correlation.

[0046] By determining the target time delay parameters through the above steps, the problems of redundant information due to too small a time delay and the breakage of dynamic correlation due to too large a time delay are effectively avoided when constructing the state vector. This provides a suitable time sampling interval for subsequent accurate reconstruction of the state vector and calculation of the transfer entropy.

[0047] Furthermore, the method provided in this application embodiment includes: Based on the target embedding dimension and target time delay parameter under the corresponding transfer direction, the preprocessed state time series data of any two devices are reconstructed into a vector set of source device state vector, target device historical state vector, and target device future state vector; based on the vector set, the transfer entropy value from the source device to the target device is calculated by a probability estimation algorithm to generate the first transfer entropy value and the second transfer entropy value.

[0048] Specifically, firstly, based on the target embedding dimension and target time delay parameter under the corresponding transfer direction, vector reconstruction is performed on the preprocessed state time series data of any two devices. Specifically, the source device state vector is constructed by selecting consecutive historical data points from the source device with the same number of target embedding dimensions and combining them into a vector in chronological order; the target device historical state vector is constructed by similarly selecting consecutive historical data points from the target device with the same number of target embedding dimensions; and the target device future state vector is constructed by selecting the next state data of the target device at the corresponding time as the vector. This forms a vector set containing the source device state vector, the target device historical state vector, and the target device future state vector.

[0049] The transition entropy value is then calculated using a probabilistic estimation algorithm. This algorithm estimates the probability by statistically analyzing the frequency of different state combinations in the time series. First, it statistically analyzes the joint occurrence frequency of the source device state vector and the target device's historical state vector in the vector set. Then, it statistically analyzes the frequency of the source device state vector occurring alone, as well as the joint occurrence frequency of the source device state vector, the target device's historical state vector, and the target device's future state vector. Based on these frequencies, conditional probabilities are calculated: the conditional probability of the target device's future state vector under the combined influence of the source device state vector and the target device's historical state vector, and the conditional probability of the target device's future state vector under the influence of only the target device's historical state vector. Finally, according to the transition entropy formula, these conditional probabilities are substituted into the calculation to obtain the transition entropy value from the source device to the target device, thus generating the first and second transition entropy values.

[0050] Furthermore, the method provided in this application embodiment includes: Extract the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device; when both the first and second transfer entropy values ​​are higher than a first preset threshold and are positively correlated, the interaction relationship between the first and second devices is determined to be a symbiotic relationship, and the minimum transfer entropy is taken as the symbiotic interaction strength; when both the first and second transfer entropy values ​​are higher than the first preset threshold and are negatively correlated, the interaction relationship between the first and second devices is determined to be a competitive relationship, and the ratio of the difference in transfer entropy to the sum of transfer entropy is calculated as the competitive interaction strength; when the first transfer entropy value is higher than the first preset threshold and the second transfer entropy value is lower than the second preset threshold, the interaction relationship between the first and second devices is determined to be a parasitic relationship, and the maximum transfer entropy is taken as the parasitic interaction strength.

[0051] In one embodiment, the bidirectional transfer entropy value between any two devices is first extracted. The transfer entropy distribution network constructed in the preceding steps stores the directed transfer entropy data of each device pair. By matching and retrieving the unique device identifier, such as the device number, the first transfer entropy value from the first device to the second device, and the second transfer entropy value from the second device to the first device, can be extracted, providing basic data support for subsequent determination of interaction relationships.

[0052] When determining a symbiotic relationship, a first preset threshold is set. This threshold is determined by statistically analyzing the distribution characteristics of the transfer entropy values ​​between equipment during the historical normal operation of a thermal power plant, taking the mean plus one standard deviation as the first preset threshold. This ensures that the threshold can distinguish between significant and weak interactions. Next, it is determined whether both the first and second transfer entropy values ​​are higher than this threshold. The Pearson correlation coefficient method is used to calculate the correlation between the two transfer entropy value sequences: if the correlation coefficient is greater than 0.3, a critical value is set at a significance level of 0.05, indicating a positive correlation. When both the first and second transfer entropy values ​​are higher than the first preset threshold and are positively correlated, the two equipment are determined to have a symbiotic relationship. In this case, the minimum transfer entropy between the two is taken as the symbiotic interaction strength, i.e., the minimum value between the first and second transfer entropy values. In a symbiotic relationship, the two equipment are interdependent, and the minimum transfer entropy reflects the lower limit of their interaction, avoiding an overestimation of the actual synergistic effect by a single high value.

[0053] When determining a competitive relationship, first confirm that both the first and second transfer entropy values ​​are higher than a first preset threshold. Then, use the Pearson correlation coefficient method to determine the correlation: if the correlation coefficient is less than -0.3, it is considered a negative correlation. When both conditions are met, the two devices are in a competitive relationship. To calculate the intensity of the competitive interaction, first calculate the difference between the first and second transfer entropy values ​​and take their absolute values. Then calculate the sum of the first and second transfer entropy values. Finally, divide the difference by the sum to obtain the ratio. This ratio quantifies the degree of imbalance in competition; a larger ratio indicates a more significant inhibitory effect of one device on the other during competition.

[0054] When determining a parasitic relationship, a second preset threshold is first set, typically the maximum transfer entropy value in historical data when there is no significant unidirectional interaction between devices. This ensures that values ​​below this threshold can be considered as having no effective reverse interaction. When the first transfer entropy value is higher than the first preset threshold and the second transfer entropy value is lower than the second preset threshold, it indicates that only a strong unidirectional interaction exists, and a parasitic relationship is determined. In this case, the maximum transfer entropy value between the first and second transfer entropy values ​​is taken as the parasitic interaction strength. The core of the parasitic relationship is the unidirectional action of the source device on the target device, and the transfer entropy value in the direction of strong action can accurately reflect the strength of this unidirectional dependence.

[0055] Finally, a dynamic device relationship graph is constructed: each device is treated as a node in the graph, and node attributes are labeled with information such as device number and device type. Based on the interaction relationship type determined in the above steps (symbiotic, competitive, parasitic), different line types such as solid lines, dashed lines, dotted lines, or colors are used to distinguish the edge types. The calculated interaction strength is used as the edge weight and labeled in the edge attributes. At the same time, based on the real-time updated transition entropy value, the edge type and weight between nodes are dynamically adjusted. If a change in the transition entropy value causes a change in the interaction relationship type, such as from symbiotic to competitive, the edge line type or color is updated synchronously; if the interaction strength changes, the edge weight value is updated, ultimately forming a dynamic device relationship graph that can reflect the interaction status between devices in real time.

[0056] By extracting bidirectional transfer entropy values, using threshold judgment and correlation analysis to determine the type of interaction relationship, calculating the corresponding interaction intensity, and combining graph structure modeling to construct a dynamic graph, the interaction characteristics between thermal power plant equipment are clearly quantified, forming an intuitive and real-time updated visualization carrier of equipment relationships, providing a precise basis for equipment interaction status for subsequent production indicator control and optimization.

[0057] Furthermore, the method provided in this application embodiment includes: Historical operating data is clustered to identify multiple typical operating conditions. A baseline equipment relationship map is calculated for each typical operating condition. The dynamically constructed equipment relationship map is compared with the baseline equipment relationship map corresponding to the current operating condition to generate abnormal relationship transfer features. Production indicators are adjusted and optimized based on the abnormal relationship transfer features.

[0058] Optionally, the K-means clustering algorithm is used to process the historical operating data. First, key feature parameters are screened from the historical operating data of the thermal power plant. These parameters need to reflect the core operating conditions and typically include multiple core indicators such as power generation, main steam temperature, main steam pressure, feedwater pump flow rate, and induced draft fan current. This avoids redundancy due to too many features or unclear distinctions between operating conditions due to too few features. Next, the screened feature parameters are preprocessed by normalization, linearly normalizing them to the 0-1 range to eliminate the influence of differences in the dimensions of different parameters on the clustering results.

[0059] Then, the number of clusters K is determined using the elbow rule: First, the range of K is set, usually 2-5, corresponding to typical operating conditions of thermal power plants such as full load, 75% rated load, 50% peak load, and start-up / shutdown load. The sum of squared errors (SSE) within clusters under different K values ​​is calculated. When the rate of decrease of SSE with increasing K value suddenly slows down, the corresponding K value is the optimal number of clusters. After determining the K value, K cluster centers are initialized, that is, K sets of historical data are randomly selected as initial centers. Then, the Euclidean distance from each data sample to each cluster center is calculated iteratively, and the sample is assigned to the nearest cluster. The center of each cluster is then recalculated, that is, the average value of each feature of all samples in the cluster is taken. The above steps of allocating samples and updating centers are repeated until the position of the cluster centers no longer changes significantly, such as the maximum distance between the centers of two adjacent iterations being less than 0.01. Finally, K typical operating conditions are obtained, and each operating condition corresponds to a subset of historical data with similar operating characteristics.

[0060] Next, a subset of historical data for the corresponding operating condition is extracted. All historical operational data belonging to this condition are selected from the clustering results to ensure that the data subset represents the normal operating status of the equipment under this condition. Then, following the aforementioned process of constructing a dynamic equipment relationship graph—calculating the bidirectional transfer entropy between equipment, determining the interaction relationship type, and determining the interaction strength—the data subset for this operating condition is processed: For each segment of continuous operational data in the data subset, such as every 24 hours as a data segment, the transfer entropy between equipment is calculated to determine the interaction relationship type of each equipment pair—symbiotic, competitive, or parasitic—and the interaction strength. Then, the calculation results for all data segments are statistically analyzed. The interaction relationship type is selected based on the highest frequency; for example, if a equipment pair is symbiotic in 80% of the data segments, then the baseline relationship type is symbiotic. The interaction strength is taken as the average of the corresponding strengths of all data segments. If outliers exist, they can be removed using the 3σ criterion before averaging, thereby eliminating the impact of single-segment data fluctuations on the baseline graph. Through the above steps, a baseline device relationship graph is finally generated for each typical working condition, which reflects the normal interaction status of the devices under that working condition. The nodes of the graph represent devices, and the type and weight of the edges correspond to the normal interaction relationship and intensity, respectively.

[0061] Subsequently, when comparing the real-time constructed dynamic equipment relationship map with the benchmark equipment relationship map corresponding to the current operating condition, it is first necessary to determine the current operating condition: real-time collection of core characteristic parameters such as the current power generation and main steam pressure of the thermal power plant, after normalization, and calculation of the Euclidean distance between it and the cluster centers of each typical operating condition. The operating condition corresponding to the cluster with the closest distance is determined as the current operating condition to ensure the accuracy of the comparison object. Then, map comparison is carried out, focusing on two aspects: first, the consistency of the interaction relationship type, checking whether the relationship type of each equipment pair in the real-time map is consistent with the benchmark map, for example, if the benchmark is a symbiotic relationship, is it still a symbiotic relationship in the real-time map? Second, the degree of difference in interaction intensity, calculating the relative error between the real-time intensity and the benchmark intensity, relative error = |real-time intensity - benchmark intensity| / benchmark intensity, and setting a relative error threshold. The commonly used threshold in the industrial field is 20%, which can be adjusted by those skilled in the art according to the actual operating accuracy requirements of the power plant. If a device pair exhibits "inconsistent relationship type" or "relative error exceeding the threshold," it is determined that the device pair has a relationship anomaly. The device pair's identifier, anomaly type (type anomaly or intensity anomaly), and anomaly degree are recorded. This information is then integrated to generate a relationship anomaly transfer feature, which can accurately locate the deviation between the device interaction state and the normal state under the current operating conditions.

[0062] Finally, preset key production indicators are constructed, and an action mapping knowledge base for abnormal relationship transfer scenarios is established based on historical data analysis. The generated abnormal relationship transfer features are then input into this action mapping knowledge base to match the corresponding execution actions. By adjusting relevant execution parameters, the production indicators are ultimately regulated and optimized. This step will be explained in detail later.

[0063] By employing K-means clustering to classify typical operating conditions, calculating the equipment relationship map based on historical data, and comparing the real-time map with the benchmark to generate abnormal features, the system achieves accurate identification of abnormal interaction states of thermal power plant equipment, providing a clear basis for subsequent targeted production indicator control and optimization.

[0064] Furthermore, the method provided in this application embodiment includes: Construct preset key production indicators and analyze an action mapping knowledge base based on historical data under abnormal relationship transfers; input the abnormal relationship transfer features into the action mapping knowledge base to match running actions, and complete the production indicator regulation by adjusting running parameters.

[0065] In one embodiment, when constructing preset key production indicators, reference is made to industry operating standards for thermal power plants, such as the "Calculation Method for Technical and Economic Indicators of Thermal Power Plants," and the core key production indicators are determined in conjunction with the actual operational needs of the power plant. These typically include power generation, coal consumption rate for power supply, main steam temperature qualification rate, and boiler thermal efficiency. These indicators directly reflect the quality and efficiency of production operation and are the core objectives of regulation and optimization. Next, an action mapping knowledge base under abnormal relationship transitions is constructed. First, all records containing abnormal relationship transition characteristics are selected from historical operating data. Each record must cover information such as the type of abnormal characteristic, the corresponding regulation action, the range of action execution parameters, and the effect of indicator changes after regulation. Then, a classification statistical method is used to group these records according to the type of abnormal characteristic. Within each group, based on the experience of power plant operation experts, the 1-2 regulation actions with the best regulation effect—that is, the largest improvement in indicators and the best operational stability—and their corresponding parameter ranges are selected and organized into a structured action mapping knowledge base. Each entry in the base represents the correspondence between abnormal characteristic type → optimal regulation action → parameter adjustment range.

[0066] Next, after inputting the real-time generated relationship anomaly transfer features into the action mapping knowledge base, a rule-based matching method is used to match operational actions: core attributes of the anomaly features are extracted, such as the equipment pairs involved in the anomaly, the anomaly type, and the anomaly severity, and compared with the core attributes of each entry's anomaly feature type in the action mapping knowledge base. Entries with completely identical attributes or the highest similarity are found, where the anomaly severity deviation does not exceed 10%. Based on the matched entries, the corresponding operational actions and parameter adjustment ranges are determined. Subsequently, through the power plant's distributed control system (DCS), the corresponding operational parameters are adjusted in real time according to the determined parameter adjustment range. During the adjustment process, changes in key production indicators are continuously monitored. If the indicators do not achieve the expected improvement effect, the process returns to the action mapping knowledge base to rematch suboptimal control actions and readjust until the key production indicators reach the expected optimization target.

[0067] In summary, the data correlation analysis method for thermal power production indicator management provided in this application has the following technical effects: This application collects time-series data on the status of equipment in thermal power plants, and obtains equipment interaction data through normalization, transfer entropy calculation, and operating condition clustering. It calculates the interaction relationships and abnormal transfer characteristics between equipment, and combines them with an action mapping knowledge base to match and control actions, thereby optimizing production indicators, making the control of thermal power plant production indicators more precise, improving the reliability and efficiency of thermal power production management, and achieving the technical effect of clarifying the mutual influence between equipment and realizing accurate evaluation and effective control of thermal power production indicators.

[0068] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a data correlation analysis platform for thermal power production indicator management, the platform comprising: The transfer entropy distribution network construction module 1 is used to collect the operating status data of multiple devices in a thermal power plant, calculate the transfer entropy between any two devices based on the time series data of the status of each device, and establish a transfer entropy distribution network.

[0069] Device relationship graph construction module 2, based on the transfer entropy value and direction between any two devices in the transfer entropy distribution network, defines the interaction relationship type and interaction intensity between devices, and constructs a dynamic device relationship graph. The interaction relationship type includes symbiotic relationship, competitive relationship and parasitic relationship.

[0070] The production indicator optimization execution module 3 optimizes and controls production indicators based on the dynamic equipment relationship graph.

[0071] Furthermore, the transfer entropy distribution network construction module 1 is used to perform the following steps: Normalize and preprocess the state time series data of each device; extract the state time series data of any two devices after preprocessing, analyze the target embedding dimension and target time delay parameter required for the transfer entropy, calculate the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device; construct directed edges based on the transfer entropy value and direction, and establish the transfer entropy distribution network with the device as the node.

[0072] Furthermore, the transfer entropy distribution network construction module 1 is used to perform the following steps: For the first and second transfer directions, target devices are determined respectively. Under the initial embedding dimension m, the nearest neighbor of each data point in the state time series data of the target device is identified in the m-dimensional state space, where m=1. The initial embedding dimension is increased to m+1, and the rate of change of the distance between each data point and the nearest neighbor in the m-dimensional space is calculated. If the rate of change exceeds a preset rate of change threshold, the nearest neighbor is determined to be a false nearest neighbor. The embedding dimension is gradually increased and judgment is performed. The proportion of false nearest neighbors is counted until the proportion of false nearest neighbors drops to a preset proportion threshold. The embedding dimension at this time is determined as the target embedding dimension required to calculate the transfer entropy under the corresponding transfer direction.

[0073] Furthermore, the transfer entropy distribution network construction module 1 is used to perform the following steps: The first transfer direction is from the first device to the second device; the second transfer direction is from the second device to the first device.

[0074] Furthermore, the transfer entropy distribution network construction module 1 is used to perform the following steps: Calculate the mutual information between the state time-series data of the target device and its own sequence after different time delays, and plot the function curve of mutual information relative to time delay; determine the time delay corresponding to the first local minimum of the function curve as the target time delay parameter under the corresponding transfer direction.

[0075] Furthermore, the transfer entropy distribution network construction module 1 is used to perform the following steps: Based on the target embedding dimension and target time delay parameter under the corresponding transfer direction, the preprocessed state time series data of any two devices are reconstructed into a vector set of source device state vector, target device historical state vector, and target device future state vector; based on the vector set, the transfer entropy value from the source device to the target device is calculated by a probability estimation algorithm to generate the first transfer entropy value and the second transfer entropy value.

[0076] Furthermore, the device relationship graph construction module 2 is used to perform the following steps: Extract the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device; when both the first and second transfer entropy values ​​are higher than a first preset threshold and are positively correlated, the interaction relationship between the first and second devices is determined to be a symbiotic relationship, and the minimum transfer entropy is taken as the symbiotic interaction strength; when both the first and second transfer entropy values ​​are higher than the first preset threshold and are negatively correlated, the interaction relationship between the first and second devices is determined to be a competitive relationship, and the ratio of the difference in transfer entropy to the sum of transfer entropy is calculated as the competitive interaction strength; when the first transfer entropy value is higher than the first preset threshold and the second transfer entropy value is lower than the second preset threshold, the interaction relationship between the first and second devices is determined to be a parasitic relationship, and the maximum transfer entropy is taken as the parasitic interaction strength.

[0077] Furthermore, the production indicator optimization execution module 3 is used to perform the following steps: Historical operating data is clustered to identify multiple typical operating conditions. A baseline equipment relationship map is calculated for each typical operating condition. The dynamically constructed equipment relationship map is compared with the baseline equipment relationship map corresponding to the current operating condition to generate abnormal relationship transfer features. Production indicators are adjusted and optimized based on the abnormal relationship transfer features.

[0078] Furthermore, the production indicator optimization execution module 3 is used to perform the following steps: Construct preset key production indicators and analyze an action mapping knowledge base based on historical data under abnormal relationship transfers; input the abnormal relationship transfer features into the action mapping knowledge base to match running actions, and complete the production indicator regulation by adjusting running parameters.

[0079] The data association analysis platform for thermal power production indicator management provided in this embodiment of the invention can execute the data association analysis method for thermal power production indicator management provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0080] Although this application makes various references to certain modules in the platform according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0081] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A data correlation analysis method for managing thermal power production indicators, characterized in that, include: Collect operating status data of multiple devices in a thermal power plant, calculate the transfer entropy between any two devices based on the time series data of each device, and establish a transfer entropy distribution network. Based on the transfer entropy value and direction between any two devices in the transfer entropy distribution network, the interaction relationship type and interaction intensity between devices are defined, and a dynamic device relationship graph is constructed. The interaction relationship types include symbiotic relationship, competitive relationship and parasitic relationship. Based on the dynamic equipment relationship graph, production indicators are adjusted and optimized.

2. The data correlation analysis method for thermal power production indicator management as described in claim 1, characterized in that, Based on the state time series data of each device, calculate the transfer entropy between any two devices and establish a transfer entropy distribution network, including: The status timing data of each device are preprocessed by normalization. Extract the state time series data of any two devices after preprocessing, analyze the target embedding dimension and target time delay parameter required for the transfer entropy, and calculate the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device. Using devices as nodes, directed edges are constructed based on the transfer entropy value and direction to establish the transfer entropy distribution network.

3. The data correlation analysis method for thermal power production indicator management as described in claim 2, characterized in that, Extract the state time series data of any two preprocessed devices, and analyze the target embedding dimension and target time delay parameters required for the transfer entropy analysis, including: For the first transfer direction and the second transfer direction, the target device is determined respectively. Under the initial embedding dimension m, the nearest neighbor of each data point in the state time series data of the target device is identified in the m-dimensional state space, where m=1; Increase the initial embedding dimension to m+1 and calculate the rate of change of the distance between each data point and its nearest neighbor in the m-dimensional space. If the rate of change exceeds a preset rate of change threshold, the nearest neighbor is determined to be a false nearest neighbor. The embedding dimension is then gradually increased and judged, and the proportion of false nearest neighbors is counted until the proportion of false nearest neighbors drops to a preset proportion threshold. The embedding dimension at this time is then determined as the target embedding dimension required to calculate the transfer entropy in the corresponding transfer direction.

4. The data correlation analysis method for thermal power production indicator management as described in claim 3, characterized in that, The first transfer direction is from the first device to the second device; the second transfer direction is from the second device to the first device.

5. The data correlation analysis method for thermal power production indicator management as described in claim 3, characterized in that, The steps for generating the target time delay parameters include: Calculate the mutual information between the state time-series data of the target device and its own sequence after different time delays, and plot the function curve of mutual information relative to time delay; The time delay corresponding to the first local minimum of the function curve is determined as the target time delay parameter in the corresponding transfer direction.

6. The data correlation analysis method for thermal power production indicator management as described in claim 5, characterized in that, Calculating the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device includes: Based on the target embedding dimension and target time delay parameter under the corresponding transfer direction, the preprocessed state time series data of any two devices are reconstructed into a vector set of source device state vector, target device historical state vector and target device future state vector; Based on the vector set, the transfer entropy value from the source device to the target device is calculated using a probability estimation algorithm, generating the first transfer entropy value and the second transfer entropy value.

7. The data correlation analysis method for thermal power production indicator management as described in claim 1, characterized in that, Based on the transfer entropy value and direction between any two devices in the transfer entropy distribution network, the interaction relationship types and interaction strengths between devices are defined, and a dynamic device relationship graph is constructed. The interaction relationship types include symbiotic relationships, competitive relationships, and parasitic relationships, including: Extract the first transfer entropy value from the first device to the second device and the second transfer entropy value from the second device to the first device; When both the first transfer entropy value and the second transfer entropy value are higher than the first preset threshold and are positively correlated, the interaction relationship between the first device and the second device is determined to be a symbiotic relationship, and the minimum transfer entropy is taken as the symbiotic interaction strength. When both the first transfer entropy value and the second transfer entropy value are higher than the first preset threshold and are negatively correlated, the interaction relationship between the first device and the second device is determined to be a competitive relationship, and the ratio of the difference in transfer entropy to the sum of transfer entropy is calculated as the competitive interaction intensity. When the first transfer entropy value is higher than the first preset threshold and the second transfer entropy value is lower than the second preset threshold, the interaction relationship between the first device and the second device is determined to be a parasitic relationship, and the maximum transfer entropy is taken as the parasitic interaction intensity.

8. The data correlation analysis method for thermal power production indicator management as described in claim 1, characterized in that, Based on the dynamic equipment relationship graph, production indicator control and optimization are performed, including: Historical operating data is clustered according to operating conditions to divide it into several typical operating conditions; A baseline equipment relationship map is calculated for each typical operating condition. The dynamically constructed equipment relationship map in real time is compared with the baseline equipment relationship map corresponding to the current operating condition to generate relationship anomaly transfer features. Production indicators are adjusted and optimized based on the aforementioned abnormal transfer characteristics of relationships.

9. The data correlation analysis method for thermal power production indicator management as described in claim 8, characterized in that, Production indicator regulation and optimization based on the aforementioned abnormal relationship transfer characteristics includes: Construct a pre-defined key production indicator and an action mapping knowledge base based on historical data analysis of abnormal relationship transfers; The abnormal transfer features of the relationship are input into the action mapping knowledge base to match the running actions, and the production indicators are controlled by adjusting the running parameters.

10. A data correlation analysis platform for thermal power production indicator management, characterized in that, The platform is used to implement the data correlation analysis method for thermal power production indicator management according to any one of claims 1-9, and the platform includes: The transfer entropy distribution network construction module is used to collect the operating status data of multiple devices in a thermal power plant, calculate the transfer entropy between any two devices based on the time series data of the status of each device, and establish a transfer entropy distribution network. The device relationship graph construction module defines the interaction relationship type and interaction intensity between devices based on the transfer entropy value and direction between any two devices in the transfer entropy distribution network, and constructs a dynamic device relationship graph. The interaction relationship types include symbiotic relationship, competitive relationship and parasitic relationship. The production indicator optimization and execution module optimizes production indicators based on the dynamic equipment relationship graph.