Multi-source heterogeneous data processing method for fault diagnosis and related equipment
By periodically processing and feature screening of multi-source heterogeneous data of the power station, a fault diagnosis model with a dual-network structure is constructed, which solves the periodic problem in the existing technology that is difficult to deal with multi-source heterogeneous data, and improves the accuracy and reliability of power station fault diagnosis.
Patent Information
- Application Number
- CN202510525234.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to process periodically in multi-source heterogeneous data, resulting in fault diagnosis results that are interfered with by periodic non-periodic data in power station fault diagnosis.
By obtaining multi-source heterogeneous data of the power station, it is divided into fault data, diagnostic data and impact data, characterize and extract the diagnostic data, and periodically divide the influencing data based on the extracted characteristics, calculate the mutual information index between the diagnostic data, periodic impact data set and fault data, perform feature screening, and obtain multi-source features. Then, a power station fault diagnosis model with dual network structure is built, including time-series extraction network and fault diagnosis network to perform fault diagnosis.
Effectively process the periodicity in timing data, improves the accuracy and reliability of power plant fault diagnosis, and can diagnose short-term failures caused by fluctuations in the cycle and long-term failures caused by long-term trends across cycles.
Smart Images

Figure CN120045984A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-source heterogeneous data processing, and in particular to a multi-source heterogeneous data processing method and related equipment for fault diagnosis. Background Art
[0002] With the rapid development of the industry, the stable operation of power stations is crucial to energy supply. However, power stations are easily affected by various factors during operation, resulting in various faults, reducing power generation efficiency and stability, and even threatening power grid security. In order to ensure the reliable operation of power stations, timely and accurate fault diagnosis technology has become a key requirement.
[0003] Traditional power station fault diagnosis methods, including photovoltaic power stations, mostly rely on a single data source, such as monitoring only conventional operating parameters such as component current and voltage, which limits the diagnostic effect. In the face of complex faults, a single data set is difficult to provide comprehensive information, which can easily lead to misjudgment and missed judgment. In recent years, multi-source data fusion technology has brought new opportunities for fault diagnosis. By integrating multi-source heterogeneous data such as power station environmental monitoring, component status, and inverter control, it can theoretically more comprehensively reflect the operating status of the power station and improve the accuracy and reliability of diagnosis.
[0004] However, in practical applications, the use of multi-source data for power plant fault diagnosis faces many technical problems and obstacles. When diagnosing power plant faults, data is usually obtained from multiple sources such as the power plant system, meteorological monitoring system, dual-light imaging system, telesignaling, and remote control system to form multi-source heterogeneous data, and then the multi-source heterogeneous data is tested to achieve power plant fault diagnosis. However, in many practical situations, time series data is affected by a variety of periodic factors, and the data is also closely related to many factors such as weather, season, and environment. This relationship of multi-factor influence leads to extremely complex changes in time series data. It is difficult for existing technologies to process the periodicity in time series data, resulting in fault diagnosis results in power plant fault diagnosis that are still interfered by periodic and non-periodic data.
[0005] In terms of data preprocessing and feature extraction, feature extraction is quite difficult. Data is affected by many factors such as weather, season, and environment, and its time series data changes are extremely complex. It is very difficult to accurately extract key features that can effectively characterize fault characteristics, and it is very easy to have inaccurate and incomplete feature extraction, thereby missing key fault information.
[0006] In terms of data preprocessing and feature extraction, data is affected by multiple factors such as weather, season, and environment. The changes in time series data are complex, and it is difficult to accurately extract key features that can effectively characterize fault characteristics. It is easy for feature extraction to be inaccurate and incomplete, and key fault information to be missed.
[0007] At the same time, when constructing a neural network model for fault diagnosis, it is necessary to comprehensively consider the heterogeneous characteristics, periodicity and other data characteristics of multi-source heterogeneous data, and how to meet the requirements of fault diagnosis accuracy and efficiency according to the data characteristics is an urgent problem to be solved at present. Summary of the Invention
[0008] Based on the problems proposed in the above background technology, the purpose of the present invention is to provide a method and related equipment for processing multi-source heterogeneous data for fault diagnosis, which solves the problem that the prior art is difficult to process the periodicity in multi-source heterogeneous data, resulting in a fault diagnosis result still being interfered by periodic and non-periodic data in power station fault diagnosis.
[0009] The present invention is realized through the following technical solutions: The first aspect of the present invention provides a method for processing multi-source heterogeneous data for fault diagnosis, including the following steps: Obtain multi-source heterogeneous data of a power station, and divide the multi-source heterogeneous data into fault data, diagnostic data and influencing data; Extract characterization features from the diagnostic data, and perform periodic division on the influencing data according to the extracted characterization features to obtain a periodic influence data set; Calculate the mutual information index between the diagnostic data, the periodic influence data set and the fault data, and perform feature screening on the periodic influence data set and the diagnostic data according to the mutual information index to obtain multi-source features; Construct a power station fault diagnosis model, wherein the fault diagnosis model includes a fault diagnosis network and a time series period extraction network; Input the multi-source features into the time series period extraction network for time series period detection, and input the detected time series period into the fault diagnosis network for fault diagnosis to obtain a power station fault result.
[0010] In the above technical solution, first, multi-source heterogeneous data of a power station is obtained. This multi-source heterogeneous data comes from multiple data sources such as a power station system and a meteorological monitoring system. Then, the multi-source heterogeneous data is divided into fault data, diagnostic data and influencing data according to types. Fault data refers to the faults and fault types existing in the power station, such as dirt and cracks on the surface of components; diagnostic data refers to the data used to judge whether there are faults in the power station, such as the voltage, current, power, etc. of components and inverters; influencing data refers to the data that has a relevant impact on power station faults, such as data on environmental factors such as light intensity, temperature, and humidity. Therefore, the multi-source heterogeneous data is divided, and then the divided data is processed by type to determine the correlation relationship between the multi-source heterogeneous data.
[0011] Among them, characteristic features including power amplitude in the diagnostic data are extracted, and the characteristic features are used to display the operation status of the power station. Then, based on the extracted characteristic features, the influencing data is periodically divided. The purpose of this step is to explore the periodicity of the influence of the influencing data on the diagnostic data. Compared with the prior art that only detects time series data, this step can capture short-term fluctuations within the period and also detect long-term trends across periods.
[0012] Since most of the data used to diagnose power station faults show non-linear relationships, after forming the periodic influence data set, the mutual information index between the diagnostic data and the periodic influence data set, and between the diagnostic data and the fault data is calculated respectively. The mutual information index is an index used to obtain the correlation between two data sets. Through the mutual information index, the linear and non-linear correlations between two data sets can be obtained. Feature screening is performed on the periodic influence data set and the diagnostic data according to the mutual information index, and multi-source features for fault diagnosis can be obtained.
[0013] During the process of diagnosing power station faults, the periodicity in the time series data is processed. On the one hand, the diagnostic data needs to be processed periodically, and on the other hand, the power station fault diagnosis model needs to be improved to meet the processing of periodic data. Therefore, the power station fault diagnosis model constructed by this method has a dual-network structure. One network is a time series period extraction network for extracting time series periodic data, and the other network is a fault diagnosis network for diagnosing faults in the extracted time series periodic data. Thus, the power station fault results obtained can diagnose short-term faults caused by fluctuations within the period and also long-term faults caused by long-term trends across periods.
[0014] In an optional embodiment, the extraction of characteristic features from the diagnostic data includes the following steps: Use the Tsfresh algorithm to extract power features from the diagnostic data to obtain the kurtosis value of the daily power of the power station; Perform fluctuation analysis on the kurtosis value of the daily power of the power station, and divide the diagnostic data into power station fluctuation days and power station stable days based on the fluctuation analysis results.
[0015] In an optional embodiment, the periodic division of the influencing data according to the extracted characteristic features includes the following steps: Align the diagnostic data and the influencing data in time based on the power station fluctuation days and the power station stable days, and divide the influencing data into fluctuation day influencing data and stable day influencing data according to the time alignment result; Merge the fluctuation day influencing data with the diagnostic data corresponding to the power station fluctuation days to generate fluctuation detection data; Perform drift detection on the fluctuation detection data to obtain mutation data points; Based on the mutation data points, perform situation change analysis on the diagnostic data, and determine the period division points according to the results of the situation change analysis; According to the period division points, perform periodic division on the influence data to obtain several segments of periodic influence data, and integrate the several segments of periodic influence data to generate a periodic influence data set.
[0016] In an alternative embodiment, calculating the mutual information index between the diagnostic data, the periodic influence data set, and the fault data includes the following steps: Extract any one diagnostic parameter from the diagnostic data, and determine the periodic influence parameter from the periodic influence data set based on the diagnostic parameter; Calculate the mutual information index between the diagnostic parameter and the periodic influence parameter to obtain the first mutual information index value; Based on the first mutual information index value, obtain the second mutual information index value of the mutual information index between the diagnostic parameter and the fault data.
[0017] In an alternative embodiment, the time series period extraction network includes: A one-dimensional period extraction layer for converting the multi-source features into two-dimensional time series features; A two-dimensional period extraction layer for extracting the intra-period change features and the inter-period change features from the two-dimensional time series features; A fusion period layer for fusing the intra-period change features and the inter-period change features.
[0018] In an alternative embodiment, the fault diagnosis network includes: A one-dimensional fault diagnosis layer for converting the multi-source features into frequency features and period features; A two-dimensional fault diagnosis layer for performing periodic diagnosis on the frequency features and the period features to obtain intra-period features and inter-period features; A fusion diagnosis layer for fusing the intra-period features and the inter-period features, and performing fault diagnosis according to the fused intra-period features and inter-period features.
[0019] In an alternative embodiment, the two-dimensional fault diagnosis layer includes an intra-period module and an inter-period module; Perform period feature distillation on the intra-period module and the inter-period module through the output data of the two-dimensional period extraction layer, where the period feature distillation includes the following steps: Perform fast Fourier transform on the multi-source features to obtain period data; Based on the period data, reconstruct the multi-source features to obtain a two-dimensional feature tensor; Calculate the absolute value of the longitudinal tensor and the absolute value of the transverse tensor of the two-dimensional feature tensor respectively, use the absolute value of the longitudinal tensor as the frequency distillation knowledge to adjust the parameters of the period module, and use the absolute value of the transverse tensor as the period distillation knowledge to adjust the parameters of the period module.
[0020] The second aspect of the present invention provides a multi-source heterogeneous data processing system for fault diagnosis, including: A data acquisition module, configured to acquire multi-source heterogeneous data of a power station, and divide the multi-source heterogeneous data into fault data, diagnostic data, and influence data; A characterization extraction module, configured to extract characterization features from the diagnostic data, and perform periodic division on the influence data according to the extracted characterization features to obtain a periodic influence data set; A feature screening module, configured to calculate the mutual information index between the diagnostic data, the periodic influence data set and the fault data, and perform feature screening on the periodic influence data set and the diagnostic data according to the mutual information index to obtain multi-source features; A model construction module, configured to construct a power station fault diagnosis model, where the fault diagnosis model includes a fault diagnosis network and a time series period extraction network; A fault diagnosis module, configured to input the multi-source features into the time series period extraction network for time series period detection, and input the detected time series period into the fault diagnosis network for fault diagnosis to obtain a power station fault result.
[0021] The third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a multi-source heterogeneous data processing method for fault diagnosis.
[0022] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a multi-source heterogeneous data processing method for fault diagnosis.
[0023] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. Based on the extracted characterization features, the influence data is periodically divided to explore the periodicity of the influence of the influence data on the diagnostic data. Compared with the prior art that only explores the time series data, the present invention can capture the short-term fluctuations within the period and also explore the long-term trends across periods; 2. Construct a power station fault diagnosis model with a dual-network structure. One network is a time-series cycle extraction network for extracting time-series periodic data, and the other network is a fault diagnosis network for diagnosing faults in the extracted time-series periodic data. Thus, the obtained power station fault results can diagnose short-term faults caused by fluctuations within the diagnosis cycle and long-term faults caused by cross-cycle long-term trends. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 FIG. is a schematic flowchart of a multi-source heterogeneous data processing method for fault diagnosis provided in Embodiment 1 of the present invention; Figure 2 FIG. is a schematic structural diagram of a power station fault diagnosis model provided in Embodiment 1 of the present invention; Figure 3 FIG. is a schematic structural diagram of a time-series cycle extraction network provided in Embodiment 1 of the present invention; Figure 4 FIG. is a schematic structural diagram of a multi-source heterogeneous data processing system for fault diagnosis provided in Embodiment 2 of the present invention; Figure 5 FIG. is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the embodiments and the drawings. The illustrative embodiments and descriptions thereof of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0026] Figure 1 FIG. is a schematic flowchart of a multi-source heterogeneous data processing method for fault diagnosis provided in Embodiment 1 of the present invention. As Figure 1 shown, the multi-source heterogeneous data processing method for fault diagnosis includes the following steps: Obtain multi-source heterogeneous data of the power station, and divide the multi-source heterogeneous data into fault data, diagnosis data, and influence data; Extract characterization features from the diagnosis data, and perform periodic division on the influence data according to the extracted characterization features to obtain a periodic influence data set; Calculate the mutual information index between the diagnostic data, the periodic influence data set, and the fault data, and perform feature screening on the periodic influence data set and the diagnostic data according to the mutual information index to obtain multi-source features; Construct a power plant fault diagnosis model, where the fault diagnosis model includes a fault diagnosis network and a time series period extraction network; Input the multi-source features into the time series period extraction network for time series period detection, and input the detected time series period into the fault diagnosis network for fault diagnosis to obtain the power plant fault result.
[0027] It should be noted that when performing power plant fault diagnosis, data is usually obtained from multiple sources such as the power plant system, meteorological monitoring system, dual-light imaging system, telemetry, and remote control system to form multi-source heterogeneous data, and then the multi-source heterogeneous data is detected to achieve power plant fault diagnosis. However, in many actual situations, time series data is affected by various periodic factors, and the data is also closely related to many factors such as weather, season, and environment. This relationship of multi-factor influence makes the change of time series data extremely complex. It is difficult for the existing technology to process the periodicity in time series data, resulting in fault diagnosis results still being interfered by periodic and aperiodic data in power plant fault diagnosis.
[0028] In view of the technical problems existing in the existing technology, this method proposes a multi-source heterogeneous data processing method for fault diagnosis. It first obtains the multi-source heterogeneous data of the power plant, which comes from multiple data sources such as the power plant system and meteorological monitoring system, and then divides the multi-source heterogeneous data into fault data, diagnostic data, and influence data according to the type. Fault data refers to the faults and fault types existing in the power plant, such as dirt and cracks on the surface of components; diagnostic data refers to the data used to judge whether there are faults in the power plant, such as the voltage, current, and power of components and inverters; influence data refers to the data that has a related influence on power plant faults, such as data of environmental factors such as light intensity, temperature, and humidity. Therefore, the multi-source heterogeneous data is divided, and then the divided data is processed by type to determine the correlation relationship between the multi-source heterogeneous data.
[0029] Among them, extract the characterization features including power amplitude, etc. in the diagnostic data, and this characterization feature is used to show the operation status of the power plant. Then, based on the extracted characterization features, the influence data is divided periodically. The purpose of this step is to explore the periodicity of the influence of the influence data on the diagnostic data. Compared with the existing technology that only explores time series data, this step can capture short-term fluctuations within the period and also explore long-term trends across periods.
[0030] Since most of the data used for diagnosing power plant faults show non - linear relationships, after forming the periodic influence data set, the mutual information indices between the diagnostic data and the periodic influence data set, and between the diagnostic data and the fault data are calculated respectively. The mutual information index is an index used to obtain the correlation between two data sets. Through the mutual information index, the linear and non - linear correlations between two data sets can be obtained. Feature screening is performed on the periodic influence data set and the diagnostic data according to the mutual information index, and multi - source features for fault diagnosis can be obtained.
[0031] In the process of diagnosing power plant faults, when dealing with the periodicity in time - series data, on the one hand, it is necessary to perform periodic processing on the diagnostic data, and on the other hand, it is necessary to improve the power plant fault diagnosis model to meet the processing of periodic data. Therefore, the power plant fault diagnosis model constructed by this method has a dual - network structure. One network is a time - series period extraction network for extracting time - series periodic data, and the other network is a fault diagnosis network for diagnosing faults in the extracted time - series periodic data. Thus, the power plant fault results obtained can diagnose short - term faults caused by fluctuations within the diagnosis period and long - term faults caused by long - term trends across periods.
[0032] Furthermore, error analysis is carried out using the power plant fault results and the fault data, and the error analysis results are then used to optimize the power plant fault diagnosis model.
[0033] In an alternative embodiment, the extraction of characterization features from the diagnostic data includes the following steps: Use the Tsfresh algorithm to extract power features from the diagnostic data to obtain the kurtosis value of the power of the power plant on a daily basis; Perform fluctuation analysis on the kurtosis value of the power of the power plant on a daily basis, and divide the diagnostic data into power plant fluctuation days and power plant stable days based on the results of the fluctuation analysis.
[0034] It should be noted that the Tsfresh algorithm is an algorithm for analyzing time - series and extracting available features. In this embodiment, the Tsfresh algorithm is used to extract power features in the diagnostic data at daily intervals to obtain the kurtosis value of the power of the power plant on a daily basis. This kurtosis value can reflect the change of the power generation power of the power plant every day. Perform fluctuation analysis on the change of the power generation power of the power plant every day, find the peaks and valleys, and then find the kurtosis values around the peaks and valleys. By judging the kurtosis values of the peaks, valleys and their surroundings, the time periods of the peaks and valleys can be divided into fluctuation time periods and stable time periods. Based on this, the diagnostic data can be divided into power plant fluctuation days and power plant stable days, and thus the fluctuation period and the stable period can be obtained.
[0035] In an alternative embodiment, the periodic division of the influence data according to the extracted characterization features includes the following steps: Align the diagnostic data with the impact data based on the power station's fluctuating days and stable days, and divide the impact data into fluctuating-day impact data and stable-day impact data according to the result of the time alignment. Merge the fluctuating-day impact data with the diagnostic data corresponding to the power station's fluctuating days to generate fluctuation detection data. Perform drift detection on the fluctuation detection data to obtain mutation data points. Based on the mutation data points, conduct a situation change analysis on the diagnostic data, and determine the cycle division points according to the result of the situation change analysis. According to the cycle division points, perform a periodic division on the impact data to obtain several segments of periodic impact data, and integrate the several segments of periodic impact data to generate a periodic impact data set.
[0036] It should be noted that aligning the power station's fluctuating days and stable days with the impact data correlates the diagnostic data and the impact data on the time line through time alignment to determine the impact relationship of the events in the impact data on the diagnostic data. The impact data after time alignment can be divided into fluctuating-day impact data and stable-day impact data according to the fluctuation situation of the power station.
[0037] Since it is necessary to analyze the situation of the fluctuating days, the fluctuating-day impact data is merged with the power station's fluctuating-day data to generate fluctuation detection data. Here, the fluctuating-day impact data and the diagnostic data corresponding to the power station's fluctuating days are time-aligned.
[0038] Then, use the drift algorithm to perform drift detection on the fluctuation detection data. The drift algorithm is a non-parametric clustering algorithm based on density estimation. Its core idea is to continuously adjust the positions of data points to make them "drift" towards the region with the highest density, thereby finding the local maximum of the probability density function of the data and realizing clustering. In this embodiment, the drift algorithm is used to detect the fluctuation detection data, and the detection data corresponding to the large fluctuations of the power station is called mutation data.
[0039] The mutation data points detected by the drift algorithm are fluctuation detections from the data level, which is the extracted fluctuation process. Extract all the mutation data points and conduct a situation change analysis on the diagnostic data corresponding to all the mutation data points. This situation change analysis is a fluctuation detection from the characteristic level to detect whether the mutation data points conform to the characteristics to exclude the interference of abnormal data. Use the mutation data points that conform to the characteristics as the cycle division points for periodic division.
[0040] Further, the periodic impact data set is expressed as follows: ; In the above formula, represents the th segment and the th periodic influence data in the periodic influence data set; is the number of periodic influence data in each segment.
[0041] It should be noted that , , are mutation data points, which are used for subsequent reconstruction of multi-source features based on periodic data.
[0042] In an alternative embodiment, calculating the mutual information index between the diagnostic data, the periodic influence data set, and the fault data includes the following steps: Extract any one diagnostic parameter from the diagnostic data, and determine the periodic influence parameter from the periodic influence data set based on the diagnostic parameter; Calculate the mutual information index between the diagnostic parameter and the periodic influence parameter to obtain the first mutual information index value; Based on the first mutual information index value, calculate the mutual information index between the diagnostic parameter and the fault data to obtain the second mutual information index value.
[0043] In this embodiment, the MIC algorithm (Maximal Information Coefficient) is used to calculate the mutual information index between the diagnostic parameter and the periodic influence parameter as the first mutual information index value.
[0044] Similarly, use the MIC algorithm to calculate the mutual information index between the diagnostic parameter and the fault data, and multiply this mutual information index by the first mutual information index value to obtain the second mutual information index value.
[0045] Among them, the first mutual information index value represents the mutual influence relationship between the periodic influence parameter and the diagnostic parameter.
[0046] The second mutual information index value represents the mutual influence relationship between the periodic influence parameter and the fault data.
[0047] Furthermore, according to the mutual information index, perform associated feature screening on the diagnostic data and the periodic influence data, including: screening the periodic influence parameters with high second mutual information index as the multi-source features of the influence data source. Screening the diagnostic parameters with high mutual information index between the diagnostic parameter and the fault data as the multi-source features of the diagnostic data source.
[0048] In an alternative embodiment, the time series period extraction network includes: A one-dimensional period extraction layer for converting the multi-source features into two-dimensional time series features; A two-dimensional period extraction layer for extracting intra-period change features and inter-period change features from the two-dimensional time-series features; A fusion period layer for fusing the intra-period change features and the inter-period change features.
[0049] In an alternative embodiment, the fault diagnosis network includes: A one-dimensional fault diagnosis layer for transforming the multi-source features into frequency features and period features; A two-dimensional fault diagnosis layer for performing periodic diagnosis on the frequency features and the period features to obtain intra-period features and inter-period features; A fusion diagnosis layer for fusing the intra-period features and the inter-period features and performing fault diagnosis based on the fused intra-period features and inter-period features.
[0050] It should be noted that the power station fault diagnosis model is a two-layer network model. As Figure 2 shown, the power station fault diagnosis model includes a time-series period extraction network and a fault diagnosis network.
[0051] The time-series period extraction network is used to extract the change features within the period and the change features between periods in the multi-source features. Among them, the change features within the period represent, in fault diagnosis: the change of the fault data with the diagnostic data and the influence data within a short period; the change features between periods represent, in fault diagnosis: the change of the fault data with the diagnostic data and the influence data between each change period.
[0052] Specifically, the time-series period extraction network is as Figure 3 shown. The time-series period extraction network includes a one-dimensional period extraction layer, a two-dimensional period extraction layer, and a fusion period layer. Taking the power-time feature in the multi-source features as an example, the one-dimensional period extraction layer is used to transform the power-time feature into two-dimensional time-series data and input it into the two-dimensional period extraction layer. The two-dimensional time-series data includes time-series data in the frequency dimension and time-series data in the period dimension. Among them, the time-series data in the frequency dimension contains the change of the fault data within the period, and the time-series data in the period dimension contains the change of the fault data between periods. Therefore, the two-dimensional period extraction layer is used to extract the intra-period change features and the inter-period change features from the two-dimensional time-series data. Among them, the two-dimensional period extraction layer extracts the intra-period change features from the time-series data in the frequency dimension and extracts the inter-period change features from the time-series data in the period dimension. Finally, the fusion period layer fuses the inter-period change features and the intra-period change features, transforming the two-dimensional data into one-dimensional data for information aggregation.
[0053] The multi-scale pattern learning in the time domain and the multi-period pattern learning in the frequency domain of multi-source features are carried out through the time series period extraction network, thus solving the problem that the existing technology is difficult to process the periodicity in time series data, resulting in the problem that the fault diagnosis results in power plant fault diagnosis are still interfered by periodic and aperiodic data.
[0054] Furthermore, although the time series period extraction network solves the problem of periodic feature extraction and processing of multi-source features by the existing network models, the time series period extraction network established based on the Transformer architecture is difficult to meet the requirements of actual deployment in terms of computational overhead, especially for the processing of multi-source data, making the time series period extraction network more massive and more complex.
[0055] Based on the above problems, this embodiment further proposes a fault diagnosis network, and the fault diagnosis network is constructed by using a lightweight network architecture. In this embodiment, the lightweight network architecture is a multi-layer perceptron network architecture, including multiple network architectures that can capture time series features, such as the TinyTime Mixers model (a lightweight time series basic model). The lightweight network architecture reduces the complexity and the number of parameters of the model by methods such as reducing the number of network layers, optimizing the activation function, and parameter sharing, thereby improving the computational efficiency and reducing the memory occupation, and is suitable for processing multi-source data to reduce the overhead and improve the computational efficiency. However, although the lightweight network architecture has the ability to process data quickly, its fault diagnosis ability is weak and it is difficult to obtain accurate fault diagnosis results. Therefore, in this embodiment, the knowledge in the time series period extraction network is migrated to the fault diagnosis network by means of knowledge distillation to improve the fault diagnosis ability of the fault diagnosis network.
[0056] Specifically, the fault diagnosis network includes a one-dimensional fault diagnosis layer, a two-dimensional fault diagnosis layer, and a fusion diagnosis layer. The one-dimensional fault diagnosis layer is used to transform the multi-source features into frequency features including the changing features within the fault period and periodic features including the changing features between fault periods.
[0057] The two-dimensional fault diagnosis layer is used to perform periodic diagnosis of faults on frequency characteristics and period characteristics to obtain in-period characteristics and during-period characteristics. Among them, the two-dimensional fault diagnosis layer includes an in-period module and a during-period module. The in-period module is used to diagnose the in-period change characteristics of frequency characteristics, and the during-period module is used to diagnose the during-period change characteristics of period characteristics. Since relying solely on the in-period module and the during-period module for fault diagnosis is difficult to obtain accurate fault diagnosis results. Therefore, the input end of the in-period module is connected to the output end of the two-dimensional period extraction layer. The two-dimensional period extraction layer transfers frequency-related knowledge to the in-period module through knowledge distillation to adjust the parameters of the in-period module to improve the diagnostic accuracy of the in-period module for frequency characteristics; similarly, the input end of the during-period module is connected to the output end of the two-dimensional period extraction layer. The two-dimensional period extraction layer transfers period-related knowledge to the during-period module through knowledge distillation to adjust the parameters of the during-period module to improve the diagnostic accuracy of the during-period module for period characteristics.
[0058] Finally, the fusion diagnosis layer fuses the in-period characteristics and the during-period characteristics to generate fusion characteristics, and the fusion diagnosis layer performs fault diagnosis on the fusion characteristics to obtain the fault diagnosis result.
[0059] Furthermore, the output parameters of the fusion period layer are used to adjust the parameters of the fault diagnosis network.
[0060] In an alternative embodiment, the two-dimensional fault diagnosis layer includes an in-period module and a during-period module; Perform period feature distillation on the in-period module and the during-period module through the output data of the two-dimensional period extraction layer. Among them, the period feature distillation includes the following steps: Perform a fast Fourier transform on the multi-source features to obtain period data; Reconstruct the multi-source features based on the period data to obtain a two-dimensional feature tensor; Calculate the absolute value of the longitudinal tensor and the absolute value of the transverse tensor of the two-dimensional feature tensor respectively. Use the absolute value of the longitudinal tensor as frequency distillation knowledge to adjust the parameters of the in-period module, and use the absolute value of the transverse tensor as period distillation knowledge to adjust the parameters of the during-period module.
[0061] It should be noted that in the one-dimensional period extraction layer, a fast Fourier transform is performed on the multi-source features to obtain period data, and the multi-source features are reconstructed based on the period data to reconstruct the one-dimensional multi-source features into two-dimensional feature tensors. Among them, the column tensor of the two-dimensional feature tensor is the frequency, which is used to represent the in-period change; the row tensor of the two-dimensional feature tensor is the period, which is used to represent the during-period change. Calculate the absolute value of the column tensor and the absolute value of the row tensor of the two-dimensional feature tensor respectively to obtain the absolute value of the longitudinal tensor and the absolute value of the transverse tensor.
[0062] The absolute value of the longitudinal tensor is used as frequency distillation knowledge to adjust the weights of each hidden layer in the period module, so as to achieve in-period feature transfer; similarly, the absolute value of the transverse tensor is used as cycle distillation knowledge to adjust the weights of each hidden layer in the period module, so as to achieve during-period feature transfer.
[0063] Furthermore, reconstruct the multi-source features based on the periodic data, including: Perform a fast Fourier transform on the multi-source features to obtain the intensity of the frequency components. Among them, the multi-source features are expressed as follows: ; In the above formula, represents the th segment and the th feature in the multi-source features.
[0064] In this embodiment, to show the role of the mutation data points in this embodiment, assume that the mutation data points are not screened out. At this time, the multi-source features correspond one-to-one with the periodic influence data set .
[0065] Select the frequency components with the largest intensity of the frequency components in each segment of the multi-source features: ; In the above formula, represents the th segment and the th frequency component.
[0066] At this time, calculate the period length corresponding to each segment of the multi-source features: ; In the above formula, represents the th segment and the th period length.
[0067] Among them, the intensity of the frequency components, the frequency components, and the period length constitute the periodic data. During the reconstruction process, the frequency components are used as the number of columns of the two-dimensional feature tensor, and the period length is used as the number of rows of the two-dimensional feature tensor.
[0068] The purpose of calculating the original one-dimensional data in segments based on the mutation data points is to retain the influence of the influence data and the diagnostic data on the periodicity of the fault data. The one-dimensional event data without segmentation cannot reflect the influence of the influence data and the diagnostic data on the periodicity of the fault data.
[0069] Figure 4 This is the structural schematic diagram of the multi-source heterogeneous data processing system for fault diagnosis provided in Embodiment 2 of the present invention, asFigure 4 As shown in the figure, the system includes: A data acquisition module, configured to acquire multi-source heterogeneous data of a power station, and divide the multi-source heterogeneous data into fault data, diagnostic data, and impact data; A characterization extraction module, configured to extract characterization features from the diagnostic data, and perform periodic division on the impact data according to the extracted characterization features to obtain a periodic impact data set; A feature screening module, configured to calculate the mutual information index between the diagnostic data, the periodic impact data set, and the fault data, and perform feature screening on the periodic impact data set and the diagnostic data according to the mutual information index to obtain multi-source features; A model construction module, configured to construct a power station fault diagnosis model, where the fault diagnosis model includes a fault diagnosis network and a timing cycle extraction network; A fault diagnosis module, configured to input the multi-source features into the timing cycle extraction network for timing cycle detection, and input the detected timing cycle into the fault diagnosis network for fault diagnosis to obtain a power station fault result.
[0070] Figure 5 The figure is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. As Figure 5 shown, the electronic device includes a processor 21, a memory 22, an input device 23, and an output device 24; the number of processors 21 in the computer device may be one or more. Figure 5 Taking one processor 21 as an example; the processor 21, the memory 22, the input device 23, and the output device 24 in the electronic device may be connected through a bus or other means. Figure 5 Taking the connection through the bus as an example.
[0071] The memory 22, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules. The processor 21 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 22, that is, implements the multi-source heterogeneous data processing method for fault diagnosis in Embodiment 1.
[0072] The memory 22 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 22 may further include a memory remotely disposed relative to the processor 21, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0073] The input device 23 may be used to receive user input such as an id and a password. The output device 24 is used to output a network configuration page.
[0074] Embodiment 4 of the present invention further provides a computer-readable storage medium, and the computer-executable instructions are used to implement the multi-source heterogeneous data processing method for fault diagnosis provided in Embodiment 1 when executed by a computer processor.
[0075] A storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the method operations provided in Embodiment 1, and may also execute related operations in the multi-source heterogeneous data processing method for fault diagnosis provided in any embodiment of the present invention.
[0076] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-source heterogeneous data processing method for fault diagnosis, characterized in that: The steps include: Acquire multi-source heterogeneous data of a power station, and divide the multi-source heterogeneous data into fault data, diagnosis data, and impact data; Extracting characterization features from the diagnostic data, and periodically dividing the impact data according to the extracted characterization features to obtain a periodic impact data set; Calculating the mutual information index between the diagnostic data, the periodic impact data set and the fault data, and performing feature screening on the periodic impact data set and the diagnostic data according to the mutual information index to obtain multi-source features; Constructing a power plant fault diagnosis model, wherein the fault diagnosis model includes a fault diagnosis network and a timing cycle extraction network; The multi-source features are input into the timing cycle extraction network to perform timing cycle detection, and the detected timing cycle is input into the fault diagnosis network to perform fault diagnosis to obtain a power plant fault result.
2. The multi-source heterogeneous data processing method for fault diagnosis according to claim 1, characterized in that: Extracting the characteristic features of the diagnostic data comprises the following steps: Using the Tsfresh algorithm to extract power features from the diagnostic data, and obtain the peak value of the power station's daily power; A fluctuation analysis is performed on the peak value of the power plant's daily power, and based on the fluctuation analysis result, the diagnostic data is divided into power plant fluctuation days and power plant stable days.
3. The multi-source heterogeneous data processing method for fault diagnosis according to claim 2, characterized in that: The impact data is periodically divided according to the extracted characterization features, including the following steps: Based on the power station fluctuation day and the power station stable day, the diagnostic data and the impact data are time-aligned, and the impact data is divided into fluctuation day impact data and stable day impact data according to the result of the time alignment; Merging the fluctuation day impact data with the diagnostic data corresponding to the power station fluctuation day to generate fluctuation detection data; Performing drift detection on the fluctuation detection data to obtain mutation data points; Performing situation change analysis on the diagnostic data based on the mutation data points, and determining a period division point according to the situation change analysis result; The impact data is periodically divided according to the periodic division points to obtain a number of segments of periodic impact data, and the several segments of periodic impact data are integrated to generate a periodic impact data set.
4. The multi-source heterogeneous data processing method for fault diagnosis according to claim 3 is characterized in that: Calculating the mutual information index between the diagnostic data, the periodic impact data set and the fault data comprises the following steps: Extracting any one diagnostic parameter from the diagnostic data, and determining a periodic influence parameter from the periodic influence data set based on the diagnostic parameter; Calculating the mutual information index between the diagnostic parameter and the periodic influencing parameter to obtain a first mutual information index value; A second mutual information index value is obtained based on the mutual information index between the diagnostic parameter and the fault data according to the first mutual information index value.
5. The multi-source heterogeneous data processing method for fault diagnosis according to claim 1, characterized in that: The timing cycle extraction network comprises: A one-dimensional period extraction layer, used for converting the multi-source features into two-dimensional time series features; A two-dimensional period extraction layer, used to extract intra-period change features and period change features from the two-dimensional time series features; The fusion period layer is used to fuse the intra-period change feature and the period change feature.
6. The multi-source heterogeneous data processing method for fault diagnosis according to claim 5, characterized in that: The fault diagnosis network comprises: A one-dimensional fault diagnosis layer, used for converting the multi-source features into frequency features and period features; A two-dimensional fault diagnosis layer, used for performing periodic diagnosis on the frequency characteristics and the period characteristics to obtain intra-period characteristics and period characteristics; The fusion diagnosis layer is used to fuse the intra-period features with the period features, and perform fault diagnosis based on the fused intra-period features and period features.
7. The multi-source heterogeneous data processing method for fault diagnosis according to claim 6, characterized in that: The two-dimensional fault diagnosis layer includes an intra-period module and a period module; The period feature distillation is performed on the intra-period module and the period module through the output data of the two-dimensional period extraction layer, wherein the period feature distillation comprises the following steps: Perform fast Fourier transform on multi-source features to obtain periodic data; Reconstructing the multi-source features based on the periodic data to obtain a two-dimensional feature tensor; The absolute value of the longitudinal tensor and the absolute value of the transverse tensor of the two-dimensional feature tensor are calculated respectively, and the absolute value of the longitudinal tensor is used as the frequency distillation knowledge to adjust the parameters of the intra-period module, and the absolute value of the transverse tensor is used as the period distillation knowledge to adjust the parameters of the period module.
8. A multi-source heterogeneous data processing system for fault diagnosis, characterized in that: include: A data acquisition module, used to acquire multi-source heterogeneous data of the power station, and divide the multi-source heterogeneous data into fault data, diagnosis data and impact data; A characterization extraction module, used to extract characterization features from the diagnostic data, and periodically divide the impact data according to the extracted characterization features to obtain a periodic impact data set; A feature screening module, used to calculate the mutual information index between the diagnostic data, the periodic impact data set and the fault data, and perform feature screening on the periodic impact data set and the diagnostic data according to the mutual information index to obtain multi-source features; A model building module, used to build a power plant fault diagnosis model, wherein the fault diagnosis model includes a fault diagnosis network and a timing cycle extraction network; The fault diagnosis module is used to input the multi-source features into the timing cycle extraction network to perform timing cycle detection, and input the detected timing cycle into the fault diagnosis network to perform fault diagnosis to obtain a power station fault result.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for processing multi-source heterogeneous data for fault diagnosis as claimed in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-source heterogeneous data processing method for fault diagnosis as described in any one of claims 1 to 7 is implemented.