Electric power engineering safety quality intelligent diagnosis system based on AI
By analyzing and filtering the validity and invalidity of the augmented data generated by the adversarial network, a training set was constructed, which solved the accuracy problem of the transformer fault prediction model and achieved more accurate fault prediction.
Patent Information
- Application Number
- CN202510885200.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing transformer fault prediction models have poor fault prediction accuracy, mainly because the sample datasets extended by GAN generative adversarial networks lack diversity and heterogeneity, resulting in large generalization errors and affecting the stability of transformer operation and maintenance.
By acquiring normal data, fault data, and generative adversarial network augmentation data of transformers, analyzing the valid and invalid values of each augmentation data, determining its retention value, eliminating redundant or heterogeneous data, constructing a training set, and training a fault prediction model.
This improved the accuracy of the transformer fault prediction model, ensured the diversity and representativeness of the training data, and enhanced the accuracy and stability of fault prediction.
Smart Images

Figure CN120833085A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of transformer fault prediction, and in particular to an AI-based power engineering safety quality intelligent diagnosis system. BACKGROUND
[0002] Power engineering is related to the production, transmission and distribution of electric energy, mainly including the installation, commissioning, operation and maintenance of power equipment. The transformer is an important power equipment in power engineering. After the installation of the transformer in the power engineering is completed, the safety quality of the transformer during operation needs to be diagnosed in a timely manner for maintenance.
[0003] The fault prediction model based on deep learning, machine learning and other AI technologies has significant advantages in the safety quality intelligent diagnosis of power engineering compared to traditional manual inspection and diagnosis methods. It can predict the occurrence of equipment failure and maintain the equipment in a timely manner by monitoring and analyzing the operation data of the installed power equipment in the power engineering.
[0004] However, the scale of the transformer fault sample data is usually very small, which makes it difficult for the AI-based fault prediction model to obtain enough training data. Therefore, the existing method usually expands the transformer fault sample data, such as using the sample data set expanded by the GAN generative adversarial network to train the transformer fault prediction model. However, the GAN generative adversarial network is prone to mode collapse, which makes the expanded sample data set still lack of diversity and cannot effectively cover the fault data distribution of different categories of transformers. In addition, the GAN generative adversarial network may generate heterogeneous sample data (due to the use of too much noise data), therefore, the transformer fault prediction model trained by the sample data set expanded by the GAN generative adversarial network directly is prone to large generalization error, which further leads to misdiagnosis when diagnosing the safety quality of the transformer during operation after the installation in the power engineering is completed, thereby affecting the operation and maintenance of the transformer and the stability of the power transmission and transformation system.
[0005] That is, the fault prediction accuracy of the transformer fault prediction model provided by the prior art is poor. SUMMARY
[0006] In order to solve the technical problem of poor fault prediction accuracy of the transformer fault prediction model provided by the prior art, the purpose of the present application is to provide an AI-based power engineering safety quality intelligent diagnosis system, and the technical solution adopted is as follows: In a first aspect, an embodiment of the present application provides an AI-based power engineering safety quality intelligent diagnosis system, which comprises: The acquisition module is configured to acquire a plurality of normal data, a plurality of fault data, and a plurality of augmented data, wherein the normal data is operation data of the transformer in a normal working state, the fault data is operation data of the transformer in a fault state, and the augmented data is data obtained by inputting the plurality of fault data into a generative adversarial network; The analysis module is configured to acquire an effective value and an invalid value of each augmented data, wherein the effective value is used to represent a data perfection degree of the corresponding augmented data to a fault type, and the invalid value is used to represent a data distribution deviation degree of the corresponding augmented data from the associated plurality of fault data, and the augmented data and the associated plurality of fault data belong to the same fault type. The determination module is configured to determine a reserved value of each augmented data according to the effective value and the invalid value of each augmented data. The training module is configured to construct a training set based on the plurality of normal data, the plurality of fault data, and the plurality of augmented data after the elimination, and train a fault prediction model according to the training set, wherein the reserved value and a probability of elimination of the corresponding augmented data are in a negative correlation relationship.
[0007] In one embodiment, the analysis module includes an effective analysis submodule, and the effective analysis submodule is configured to: acquire a feature difference value and a range feature value corresponding to each augmented data, wherein the feature difference value is used to represent a difference degree between the corresponding augmented data and the associated plurality of fault data, and the range feature value is used to represent an increase amplitude of a data coverage range of the associated plurality of fault data in a corresponding feature space after adding the corresponding augmented data. obtain the effective value of each augmented data according to the feature difference value and the range feature value corresponding to each augmented data.
[0008] In one embodiment, the effective analysis submodule includes a feature difference analysis unit, and the feature difference analysis unit is configured to: acquire a first feature vector of each augmented data after feature extraction, and a second feature vector of each fault data after feature extraction; analyze a difference degree between the first feature vector of each augmented data and the plurality of second feature vectors associated therewith, to obtain a plurality of vector difference degrees corresponding to each augmented data; determine a mean value of the plurality of vector difference degrees corresponding to each augmented data as the feature difference value corresponding to each augmented data.
[0009] In one embodiment, the acquisition of the first feature vector of each augmented data after feature extraction, and the second feature vector of each fault data after feature extraction includes: obtaining a plurality of first feature values of each augmented data after feature extraction, and a plurality of second feature values of each fault data after feature extraction, wherein the plurality of first feature values and a plurality of preset data dimensions correspond one by one, and the plurality of second feature values also correspond one by one with the plurality of data dimensions; splicing the plurality of first feature values of each augmented data after feature extraction to obtain a first spliced vector of each augmented data, and splicing the plurality of second feature values of each fault data after feature extraction to obtain a second spliced vector of each fault data; performing normalization processing on the first spliced vector of each augmented data to obtain a first feature vector corresponding to each augmented data, and performing normalization processing on the second spliced vector of each fault data to obtain a second feature vector of each fault data.
[0010] In an embodiment, the effective analysis submodule includes a range analysis unit, and the range analysis unit is configured to: obtaining a plurality of first feature values of each augmented data after feature extraction, and a plurality of second feature values of each fault data after feature extraction, wherein the plurality of first feature values and a plurality of preset data dimensions correspond one by one, and the plurality of second feature values also correspond one by one with the plurality of data dimensions; aggregating the first feature value of each augmented data at each data dimension and the plurality of second feature values associated with the data dimension to obtain a first feature set corresponding to each augmented data at each data dimension; and aggregating the plurality of second feature values associated with each augmented data at each data dimension to obtain a second feature set corresponding to each augmented data at each data dimension; analyzing the difference between the first feature set and the second feature set corresponding to each augmented data at each data dimension to obtain a range feature value corresponding to each augmented data.
[0011] In an embodiment, the step of analyzing the difference between the first feature set and the second feature set corresponding to each augmented data at each data dimension to obtain a range feature value corresponding to each augmented data includes: determining a dimension range value corresponding to each augmented data at each data dimension according to the ratio of the range of the first feature set and the second feature set corresponding to each augmented data at each data dimension; determining the mean value of a plurality of dimension range values corresponding to each augmented data as the range feature value corresponding to each augmented data, wherein the plurality of dimension range values correspond one by one with the plurality of data dimensions.
[0012] In one embodiment, the effective analysis submodule comprises an effective value calculation unit, which is configured to: normalizing the feature difference value corresponding to each augmented data to obtain a first normalized value corresponding to each augmented data, and normalizing the range feature value corresponding to each augmented data to obtain a second normalized value corresponding to each augmented data; multiplying the first normalized value and the second normalized value corresponding to each augmented data to determine the effective value of each augmented data.
[0013] In one embodiment, the analysis module comprises an invalid analysis submodule, which is configured to: performing difference analysis on each first time series data of each augmented data and a plurality of second time series data associated with the corresponding data category to obtain a plurality of data difference values corresponding to each augmented data under each data category, wherein the first time series data is the time series data of the corresponding augmented data under the corresponding data category, and the second time series data is the time series data of the fault data associated with the corresponding augmented data under the corresponding data category; determining the smallest data difference value from the plurality of data difference values corresponding to each augmented data under each data category as the difference index value of the augmented data under the data category; calculating the mean of the plurality of difference index values of each augmented data under the plurality of data categories to obtain the invalid value of each augmented data.
[0014] In one embodiment, the determination module is configured to: normalizing the effective value and the invalid value of each augmented data respectively to obtain a normalized effective value and a normalized invalid value of each augmented data; determining the ratio of the normalized effective value and the normalized invalid value of each augmented data as the retention value of each augmented data.
[0015] In one embodiment, the training module comprises a training set construction unit, which is configured to: sorting the plurality of augmented data in ascending order according to the retention value of each augmented data to obtain an augmented data sequence; sequentially eliminating augmented data from the first position of the augmented data sequence, so that the number ratio between the first sample number and the second sample number is a preset ratio, wherein the first sample number is the total number of the plurality of fault data and the remaining augmented data after elimination, and the second sample number is the number of the plurality of normal data; The training set is constructed by taking the plurality of fault data and the remaining augmented data after elimination as positive samples and the plurality of normal data as negative samples.
[0016] In a second aspect, another embodiment of the present application provides an AI-based power engineering safety quality intelligent diagnosis method, which comprises the following steps: Obtaining a plurality of normal data, a plurality of fault data and a plurality of augmented data, wherein the normal data are operation data of a transformer in a normal working state, the fault data are operation data of the transformer in a fault state, and the augmented data are data obtained by inputting the plurality of fault data into a generative adversarial network; Obtaining an effective value and an invalid value of each augmented data, wherein the effective value is used to represent the data perfection degree of the corresponding augmented data to the fault type, and the invalid value is used to represent the data distribution deviation degree of the corresponding augmented data from the plurality of fault data, and the augmented data and the plurality of fault data belong to the same fault type; Determining a reserved value of each augmented data according to the effective value and the invalid value of each augmented data; Constructing a training set based on the plurality of normal data, the plurality of fault data and the plurality of augmented data after elimination, and training a fault prediction model according to the training set, wherein the reserved value and the probability of elimination of the corresponding augmented data are in a negative correlation relationship.
[0017] In a third aspect, another embodiment of the present application further provides an electronic device, which comprises a processor, a memory and a computer program stored on the memory and executable on the processor, and the computer program is executed by the processor to implement the steps of the method of the second aspect.
[0018] In a fourth aspect, another embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the method of the second aspect.
[0019] The present application has the following advantages: After the multiple expansion data for supplementing and perfecting the fault data generated by the generative adversarial network output are obtained, the retention value of each expansion data is comprehensively determined by further obtaining the data perfection degree of each expansion data to the fault type to which the expansion data belongs and the degree of each expansion data deviating from the data distribution of multiple fault data of the same fault type, that is, accurately quantifying the necessary degree of retention of each expansion data, so as to identify redundant data or heterogeneous data that cannot effectively supplement or perfect the multiple fault data in the multiple expansion data, and then reasonably eliminate the multiple expansion data according to the retention value, and construct a training set according to the multiple expansion data after elimination, the multiple normal data and the multiple fault data, so as to train a fault prediction model capable of accurately predicting the fault of the transformer. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0021] Figure 1 A structure schematic diagram of an AI-based power engineering safety quality intelligent diagnosis system provided by an embodiment of the present application; Figure 2 A flowchart of an AI-based power engineering safety quality intelligent diagnosis method provided by an embodiment of the present application; Figure 3 A structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined object, the specific implementation, structure, features and effects of an AI-based power engineering safety quality intelligent diagnosis system according to the present application are described in detail as follows by combining with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0024] The specific scheme of an AI-based power engineering safety quality intelligent diagnosis system provided by the present application will be specifically described below by combining with the drawings.
[0025] The application provides an AI-based power engineering safety quality intelligent diagnosis system, please refer to Figure 1 which shows a structure schematic diagram of an AI-based power engineering safety quality intelligent diagnosis system 100 provided by an embodiment of the application, the system 100 comprises the following modules: The acquisition module 101 is configured to acquire a plurality of normal data, a plurality of fault data and a plurality of expansion data.
[0026] The normal data is the running data of the transformer in the normal working state, the fault data is the running data of the transformer in the fault state, and the expansion data is the data obtained by inputting the plurality of fault data into a generative adversarial network.
[0027] The running data includes data corresponding to multiple categories (such as current, voltage, sound, vibration, etc.), for example: voltage data (collected by a voltage sensor) of the transformer during operation, current data (collected by a current sensor), vibration data (collected by a vibration sensor (such as an accelerometer)), and sound data (collected by an audio sensor (such as a microphone)).
[0028] It should be understood that the voltage data and the current data are both time series data, i.e., representing the voltage signal and the current signal of the transformer within a period of time; similarly, the vibration data and the sound data are also time series data, representing the vibration signal and the sound signal of the transformer within a period of time.
[0029] In addition, the plurality of fault data collectively corresponds to a plurality of fault types, for example, the plurality of fault types can include phase-to-phase short circuit, winding inter-turn short circuit, single-phase grounding short circuit, etc. In application, the number of fault data corresponding to different fault types can be set to be consistent to avoid differences in prediction accuracy of different fault types due to differences in the number of fault data, so that the finally trained fault prediction model can provide more accurate fault prediction results for various fault conditions, thereby enhancing the general applicability of the finally trained fault prediction model in complex scenarios.
[0030] In some embodiments, in the case that the working environment of the transformer is easily affected by factors such as temperature and humidity in the environment and electromagnetic interference, interpolation method can be used to supplement the missing values in the collected running data (through current sensor / voltage sensor), and Savitzky-Golay filtering algorithm can be used to smooth the noise data in the collected running data to remove noise and outliers in the running data, so as to ensure the accuracy and consistency of the data.
[0031] The obtaining process of the plurality of augmented data is as follows: taking a plurality of fault data belonging to the same fault type as network training samples, performing multiple alternating iterative training on the generator and the discriminator in the generative adversarial network (first training the generator / discriminator, then training the discriminator / generator, that is, completing one alternating iterative training), and outputting, through the generator after alternating iterative training (input is random noise), a plurality of augmented data belonging to the same fault type as the input plurality of fault data.
[0032] The analysis module 102 is configured to obtain an effective value and an invalid value of each augmented data.
[0033] The effective value is used to represent the data perfection degree of the corresponding augmented data to the fault type to which the corresponding augmented data belongs, and the invalid value is used to represent the degree of deviation of the corresponding augmented data from the data distribution of the associated plurality of fault data, and the augmented data and the associated plurality of fault data belong to the same fault type.
[0034] The data perfection degree of the corresponding augmented data to the fault type to which the corresponding augmented data belongs can also be understood as: compared with a scheme of using a plurality of fault data of the fault type to which the corresponding augmented data belongs to perform model training, the rise amplitude of the prediction accuracy of the fault prediction model to the fault type to which the corresponding augmented data belongs when using the corresponding augmented data and the plurality of fault data of the fault type to which the corresponding augmented data belongs to perform model training.
[0035] The higher the effective value is, the more accurate and comprehensive the corresponding augmented data can be considered to perform the data distribution and fault characteristics of the fault type to which the corresponding augmented data belongs after the corresponding augmented data is added to the plurality of fault data of the fault type to which the corresponding augmented data belongs, and the prediction accuracy of the fault prediction model to the fault type can also be higher.
[0036] Based on the setting of the effective value, among the plurality of augmented data, augmented data that is too similar to the original fault data (such augmented data can be approximately understood as duplicate data of the corresponding fault data) can be identified, so as to reduce the probability of participation of such redundant augmented data in model training, so that each fault type can fully represent the data distribution characteristics thereof through the corresponding plurality of fault data and augmented data, so as to facilitate the fault prediction model to more accurately complete the prediction work of each fault type.
[0037] The degree of deviation of the corresponding augmented data from the data distribution of the associated plurality of fault data can also be understood as: compared with a scheme of using a plurality of fault data of the fault type to which the corresponding augmented data belongs to perform model training, the decrease amplitude of the prediction accuracy of the fault prediction model to the fault type to which the corresponding augmented data belongs when using the corresponding augmented data and the plurality of fault data of the fault type to which the corresponding augmented data belongs to perform model training.
[0038] The higher the invalid value is, the more incorrect the corresponding expansion data is to the data distribution and fault characteristics of the fault type to which the expansion data belongs after the expansion data is added to the multiple fault data of the fault type, and the lower the prediction accuracy of the fault prediction model for the fault type is.
[0039] Based on the setting of the invalid value, heterogeneous data relative to the multiple fault data can be identified among the multiple expansion data, so as to reduce the probability of the participation of the expansion data with such heterogeneity in model training, so that the addition of the expansion data can help each fault type to more fully represent the data distribution characteristics thereof, so as to enable the fault prediction model to more accurately complete the prediction work for each fault type.
[0040] Further, the analysis module 102 comprises an effective analysis submodule 1021, which is configured to: obtain a feature difference value and a range feature value corresponding to each expansion data, wherein the feature difference value is used to represent the difference degree between the corresponding expansion data and the associated multiple fault data, and the range feature value is used to represent the increase amplitude of the data coverage range of the associated multiple fault data in the corresponding feature space after the corresponding expansion data is added; obtain an effective value of each expansion data according to the feature difference value and the range feature value corresponding to each expansion data.
[0041] The setting of the feature difference value aims to quantitatively represent the low-dimensional difference between the corresponding expansion data and the associated multiple fault data at the data level, and the setting of the range feature value aims to quantitatively represent the high-dimensional difference between the corresponding expansion data and the associated multiple fault data in the feature space, so as to comprehensively consider the data improvement degree of the corresponding expansion data for the fault type to which the expansion data belongs from multiple aspects, so that the determined effective value is more accurate and reliable.
[0042] The feature difference value and the effective value are positively correlated, and the range feature value and the effective value are also positively correlated.
[0043] It should be understood that the larger the feature difference value is, the larger the difference between the position of the corresponding expansion data in the feature space (the space formed by the multiple fault data) and the multiple positions of the associated multiple fault data in the feature space is, and therefore, the corresponding expansion data can fill the regions in the feature space which are not covered by the associated multiple fault data, so as to enrich the data distribution of the fault type to which the expansion data belongs in the feature space, and further increase the probability of the fault prediction model accurately learning the fault characteristics of the fault type to which the expansion data belongs, so that the trained fault prediction model can have a higher prediction accuracy for the fault type to which the expansion data belongs.
[0044] The greater the range characteristic value is, the stronger the increase of the data coverage range of the plurality of fault data associated with the corresponding expansion data is, and thus the plurality of fault data added with the corresponding expansion data can better cover the potential distribution range of the fault type in the feature space.
[0045] Further, the effective analysis submodule 1021 comprises a feature difference analysis unit 10211, configured to: obtain a first feature vector of each expansion data after feature extraction and a second feature vector of each fault data after feature extraction; analyze the difference degree between the first feature vector of each expansion data and the plurality of second feature vectors associated therewith, to obtain a plurality of vector difference degrees corresponding to each expansion data; determine the mean of the plurality of vector difference degrees corresponding to each expansion data as the feature difference value corresponding to each expansion data.
[0046] Exemplarily, the obtaining of the first feature vector and the second feature vector can be completed by a feature extraction tool. For example, the feature extraction tool can be a tool constructed based on a convolutional autoencoder.
[0047] In an example, the Euclidean distance between the first feature vector of each expansion data and the second feature vector associated therewith can be determined as the vector difference degree corresponding to each expansion data. Accordingly, the Euclidean distances between the first feature vector of each expansion data and the plurality of second feature vectors associated therewith are calculated respectively, to obtain the plurality of vector difference degrees corresponding to each expansion data.
[0048] In some embodiments, the obtaining of the first feature vector of each expansion data after feature extraction and the second feature vector of each fault data after feature extraction comprises: obtaining a plurality of first feature values of each expansion data after feature extraction and a plurality of second feature values of each fault data after feature extraction, wherein the plurality of first feature values and a plurality of data dimensions are one-to-one corresponding, and the plurality of second feature values are also one-to-one corresponding to the plurality of data dimensions; splicing the plurality of first feature values of each expansion data after feature extraction to obtain a first splicing vector of each expansion data, and splicing the plurality of second feature values of each fault data after feature extraction to obtain a second splicing vector of each fault data; The first spliced vector of each augmented data is normalized to obtain a first feature vector corresponding to each augmented data, and the second spliced vector of each fault data is normalized to obtain a second feature vector of each fault data.
[0049] The data dimension is specifically a composite dimension composed of a type corresponding to the data source and a monitored index. The type corresponding to the data source includes a current type, a voltage type, a sound type, a vibration type, etc. The monitored index includes a statistical index (such as a mean value, a variance, a standard deviation, and an extreme value), a time domain index (such as a waveform factor and a peak factor), a frequency domain index (such as a spectral peak, an average frequency, a spectral centroid, and a power spectral density), etc.
[0050] In this embodiment, the feature extraction operation should be understood as a process of using a processing rule matched with a detected index of a corresponding data dimension to analyze data for a data part of the augmented data / fault data corresponding type (such as a current type, a voltage type, a sound type, and a vibration type).
[0051] For example, a plurality of first feature values of each augmented data after feature extraction can be sequentially connected end to end to obtain a first feature vector (one-dimensional vector) corresponding to each augmented data. The plurality of first feature values of each augmented data after feature extraction can also be up-down spliced and left-right spliced according to the data dimension to obtain a first feature vector (m x n two-dimensional vector, m is the total number of types corresponding to the data source, and n is the total number of monitored indexes) corresponding to each augmented data.
[0052] The second feature vector is obtained in the same way as the first feature vector, and details are not repeated here to avoid repetition.
[0053] The above normalization processing can be understood as normalizing (scaling the value to the (0, 1) interval) each feature element in the first feature vector and the second feature vector. For example, the above normalization processing can use Min-Max normalization processing or Z-Score standardization.
[0054] It should be noted that, compared with a method of directly extracting a feature vector through a complex operation such as convolution, a method of obtaining a feature value based on simple operations such as mean value and variance, and splicing each feature value to form a feature vector (such as a first feature vector and a second feature vector) can support more rich and flexible index setting and type setting, reduce the calculation complexity of the feature vector, and improve the efficiency of obtaining the feature vector.
[0055] Further, the effective analysis submodule 1021 includes a range analysis unit 10212, which is configured to: obtaining a plurality of first feature values of each augmented data after feature extraction, and a plurality of second feature values of each fault data after feature extraction, wherein the plurality of first feature values and a plurality of preset data dimensions correspond one by one, and the plurality of second feature values also correspond one by one with the plurality of data dimensions; aggregating the first feature value of each augmented data at each data dimension and the plurality of second feature values associated with the data dimension to obtain a corresponding first feature set of each augmented data at each data dimension; and aggregating the plurality of second feature values associated with each augmented data at each data dimension to obtain a corresponding second feature set of each augmented data at each data dimension; analyzing the difference between the corresponding first feature set and the second feature set of each augmented data at each data dimension to obtain a range feature value corresponding to each augmented data.
[0056] The above range feature value is used to quantitatively represent the difference degree of the data coverage range of the plurality of fault data associated with the corresponding augmented data before and after adding the augmented data.
[0057] For example, the intersection-union ratio between the first feature set and the second feature set of each augmented data at each data dimension can be determined as the range feature value corresponding to each augmented data.
[0058] In some embodiments, the step of analyzing the difference between the corresponding first feature set and the second feature set of each augmented data at each data dimension to obtain a range feature value corresponding to each augmented data comprises: determining a dimension range value corresponding to each augmented data at each data dimension according to the ratio of the range of the first feature set and the second feature set corresponding to each augmented data at each data dimension; determining the mean value of the plurality of dimension range values corresponding to each augmented data as the range feature value corresponding to each augmented data, wherein the plurality of dimension range values correspond one by one with the plurality of data dimensions.
[0059] In this embodiment, the data coverage range of the corresponding set (such as the first feature set and the second feature set) is represented by the range (the difference between the maximum value and the minimum value) of the set, and the corresponding dimension range value of the corresponding augmented data at the corresponding data dimension is determined by calculating the ratio of the range, so as to accurately represent the difference degree of the data coverage range of the plurality of fault data associated with the corresponding augmented data before and after adding the augmented data at the corresponding data dimension. Finally, the range feature value of the corresponding augmented data is determined by mean value calculation, so as to reduce the influence of extreme values as much as possible, and make the determined dimension range value more accurate and reliable.
[0060] Further, the effective analysis submodule 1021 comprises an effective value calculation unit 10213, configured to: normalize the feature difference value corresponding to each augmented data to obtain a first normalized value corresponding to each augmented data, and normalize the range feature value corresponding to each augmented data to obtain a second normalized value corresponding to each augmented data; the product of the first normalized value and the second normalized value corresponding to each augmented data is determined as the effective value of each augmented data.
[0061] In this way, by means of normalization processing, the numerical difference between the feature difference value and the range feature value is balanced, so as to avoid the case that the determined effective value is excessively affected by the feature difference value or the range feature value due to the numerical difference between the two, and the accuracy of the finally determined effective value is guaranteed. Herein, using the product of the two normalized values as the effective value can simplify the calculation of the effective value as much as possible under the premise of guaranteeing the accuracy of the effective value, improve the calculation efficiency of the effective value, and further reduce the preparation time of the training set of the fault prediction model.
[0062] Herein, the normalization processing on the feature difference value and the range feature value can be completed by means of Min-Max normalization method.
[0063] Further, the analysis module 102 comprises an invalid analysis submodule 1022, configured to: perform difference analysis on each first time series data of each augmented data and a plurality of second time series data associated with the augmented data under the corresponding data category, to obtain a plurality of data difference values corresponding to each augmented data under each data category, wherein the first time series data is the time series data of the corresponding augmented data under the corresponding data category, and the second time series data is the time series data of the fault data associated with the corresponding augmented data under the corresponding data category; from the plurality of data difference values corresponding to each augmented data under each data category, determine the smallest data difference value as the difference index value of the augmented data under the data category; calculate the mean value of the plurality of difference index values of each augmented data under the plurality of data categories to obtain the invalid value of each augmented data.
[0064] Exemplarily, the plurality of data categories corresponding to the fault data and the augmented data can be current, voltage, sound, vibration, etc.
[0065] Specifically, dynamic time warping (DTW) can be performed on the first time series data of each extended data under each data category and the multiple second time series data associated with it under the data category to calculate multiple DTW distances between the first time series data of each extended data under each data category and the multiple second time series data associated with it under the data category, and use them as multiple data difference values of each extended data under each data dimension.
[0066] Selecting the smallest data difference value among multiple data difference values of each extended data under each data category as the difference index value of the corresponding extended data under the corresponding data category can accurately quantify the degree of difference between the corresponding extended data under the corresponding data category and multiple other associated fault data.
[0067] Calculating the mean of multiple difference index values can reduce the impact of extreme values, so that invalid values can accurately reflect the overall deviation of the data distribution of the corresponding expanded data relative to its associated multiple fault data in multiple data categories.
[0068] It should be understood that the larger the invalid value is, the greater the degree to which the corresponding expanded data deviates from the real data distribution of the fault type to which it belongs, that is, the more likely the corresponding expanded data is to be heterogeneous data that does not have the real fault characteristics of the transformer. Therefore, the corresponding expanded data should be eliminated to ensure the authenticity of the training samples used to train the fault prediction model.
[0069] The determination module 103 is configured to determine a reserved value of each extended data according to the valid value and invalid value of each extended data.
[0070] The above-mentioned retention value is used to indicate the necessary degree to which the corresponding extended data is retained.
[0071] Specifically, the determining module 103 is configured to: Normalize the valid value and invalid value of each extended data respectively to obtain the normalized valid value and normalized invalid value of each extended data; The ratio of the normalized valid value to the normalized invalid value of each extended data is determined as the retained value of each extended data.
[0072] It should be understood that the valid values and invalid values will be normalized to the same interval (such as (0, 1)) to avoid large numerical differences between the two, thereby ensuring the accuracy of the subsequently determined retention values.
[0073] The method used for normalization processing can be found in the above description, and will not be described again to avoid repetition.
[0074] The training module 104 is configured to construct a training set based on the plurality of normal data, the plurality of fault data and the plurality of expanded data after the elimination, and train a fault prediction model according to the training set, wherein the reserved value of each expanded data is negatively correlated with the probability of being eliminated.
[0075] In one example, a reserved threshold (e.g., 2) can be preset, and the expanded data lower than the reserved threshold in the plurality of expanded data can be eliminated to obtain the plurality of expanded data after the elimination.
[0076] In an example, the fault prediction model can be a 1D-CNN neural network or a DBN neural network.
[0077] In an application, the operation data of a transformer installed in a power engineering can be used as an input of the fault prediction model, so as to perform transformer fault prediction by the fault prediction model and output a fault prediction result, and then corresponding treatment is performed based on the obtained fault prediction result.
[0078] For example, when the fault prediction result indicates that the transformer is in normal operation, the input operation data can be labeled in combination with the actual situation (e.g., when the transformer is in normal operation, the input operation data is labeled as normal data; and when the transformer has a fault, the input operation data is labeled as fault data of false recognition), so as to be used for subsequent iterative updating of the fault prediction model. Alternatively, when the fault prediction result indicates that the transformer has a fault risk, a worker is reminded to perform corresponding detection and maintenance on the target transformer.
[0079] After the plurality of expanded data used for supplementing and perfecting the fault data is obtained by the generative adversarial network, the reserved value of each expanded data is determined by further obtaining the data perfection degree of each expanded data to the fault type to which the expanded data belongs and the degree of each expanded data deviating from the data distribution of the plurality of fault data of the same fault type, that is, the necessary degree of each expanded data being reserved is accurately quantified, so as to identify redundant data or heterogeneous data that cannot effectively supplement or perfect the plurality of fault data in the plurality of expanded data, and then the plurality of expanded data is reasonably eliminated according to the reserved value, and the training set is constructed according to the plurality of expanded data after the elimination, the plurality of normal data and the plurality of fault data, so as to train the fault prediction model capable of accurately predicting the fault of the transformer.
[0080] In one embodiment, the training module comprises a training set construction unit configured to: According to the reserved value of each expanded data, the plurality of expanded data is sorted in descending order to obtain an expanded data sequence. remove the augmented data one by one from the beginning of the augmented data sequence, so that the quantity ratio between the first sample quantity and the second sample quantity is a preset ratio, wherein the first sample quantity is the total quantity of the plurality of fault data and the remaining augmented data after the removal, and the second sample quantity is the quantity of the plurality of normal data; construct a training set by taking the plurality of fault data and the remaining augmented data after the removal as positive samples and the plurality of normal data as negative samples.
[0081] Exemplarily, the preset ratio can be 1:1.5 or 1:2.
[0082] For example, if the total quantity of the augmented data in the augmented data sequence is 1000, and the quantity of the augmented data becomes 500 when the quantity ratio between the first sample quantity and the second sample quantity becomes the preset ratio, the removal operation is that the augmented data is removed one by one from the beginning of the augmented data sequence until the 500th position of the augmented data sequence (if the first position is numbered as 0, then the sequence number of the 500th position is 499).
[0083] It should be noted that the system provided in the above embodiments is only exemplified by the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the above described functions.
[0084] The present application provides an AI-based power engineering safety quality intelligent diagnosis method, please refer to Figure 2 which shows a flowchart of an AI-based power engineering safety quality intelligent diagnosis method according to an embodiment of the present application, the method comprises: In a second aspect, another embodiment of the present application provides an AI-based power engineering safety quality intelligent diagnosis method, which comprises: Step S1, obtaining a plurality of normal data, a plurality of fault data and a plurality of augmented data.
[0085] The normal data is the running data of the transformer in the normal working state, the fault data is the running data of the transformer in the fault state, and the augmented data is the data obtained by inputting the plurality of fault data into a generative adversarial network.
[0086] Step S2, obtaining the valid value and the invalid value of each augmented data.
[0087] The effective value is used to represent the data perfection degree of the corresponding expansion data to the fault type, and the ineffective value is used to represent the data distribution degree of the corresponding expansion data deviating from the associated multiple fault data, and the expansion data and the associated multiple fault data belong to the same fault type.
[0088] Step S3, determining the reserved value of each expansion data according to the effective value and the ineffective value of each expansion data.
[0089] Step S4, constructing a training set based on the multiple normal data, the multiple fault data and the multiple expansion data after elimination, and training a fault prediction model according to the training set.
[0090] The reserved value and the probability of the corresponding expansion data being eliminated are in a negative correlation relationship.
[0091] The AI-based power engineering safety quality intelligent diagnosis method and the AI-based power engineering safety quality intelligent diagnosis system provided by the above embodiment belong to the same concept, and the specific implementation process is detailed in the system embodiment, which will not be repeated here.
[0092] The embodiment of the application further provides an electronic device. Figure 3 The electronic device can include a processor 301, a memory 302, and a program 3021 stored in the memory 302 and executable on the processor 301.
[0093] The program 3021 is executed by the processor 301 to implement Figure 2 Any step in the corresponding method embodiment and the same beneficial effects can be achieved, which will not be repeated here.
[0094] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiment method can be completed by program instructions related to hardware, and the program can be stored in a readable medium.
[0095] The embodiment of the application further provides a readable storage medium, and the readable storage medium stores a computer program. Figure 2 Any step in the corresponding method embodiment and the same technical effects can be achieved, to avoid repetition, which will not be repeated here.
[0096] The computer readable storage medium of the embodiments of the present application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.
[0097] The computer readable signal medium can include a computer readable program code in a baseband or propagated as a carrier wave in a propagation medium. Such a propagated signal can take a wide variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that can be used to carry or store computer readable program code for use by or in connection with an instruction execution system, apparatus or device.
[0098] The program code contained on the storage medium can be transmitted by any suitable medium, including but not limited to wireless, wire line, optical fiber, RF, etc., or any suitable combination of the above.
[0099] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In an embodiment of the application, the remote computer can be a server or another desktop computer.
[0100] The embodiment of the present application further provides a computer program product, which, when running on a computer, causes the computer to execute the related steps to realize the AI-based power engineering safety quality intelligent diagnosis method provided by the above embodiment.
[0101] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0102] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment mainly explains the difference from other embodiments.
Claims
1. An AI-based power engineering safety quality intelligent diagnosis system, characterized in that, The system comprises: an acquisition module, configured to acquire a plurality of normal data, a plurality of fault data, and a plurality of augmented data, wherein the normal data are operation data of a transformer in a normal working state, the fault data are operation data of the transformer in a fault state, and the augmented data are data obtained after the plurality of fault data are input into a generative adversarial network; an analysis module, configured to acquire an effective value and an invalid value of each augmented data, wherein the effective value is used to represent a data perfection degree of the corresponding augmented data on a fault type, and the invalid value is used to represent a data distribution deviation degree of the corresponding augmented data from a plurality of fault data associated with the corresponding augmented data, and the augmented data and the plurality of fault data associated with the augmented data belong to the same fault type; a determination module, configured to determine a reserved value of each augmented data according to the effective value and the invalid value of each augmented data; a training module, configured to construct a training set based on the plurality of normal data, the plurality of fault data, and the plurality of augmented data after elimination, and to obtain a fault prediction model according to the training set, wherein the reserved value and a probability of elimination of the corresponding augmented data are in a negative correlation relationship. 2.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 1, characterized in that, The analysis module comprises an effective analysis submodule, and the effective analysis submodule is configured to: acquire a feature difference value and a range feature value corresponding to each augmented data, wherein the feature difference value is used to represent a difference degree between the corresponding augmented data and a plurality of fault data associated with the corresponding augmented data, and the range feature value is used to represent an increase amplitude of a data coverage range of the plurality of fault data associated with the corresponding augmented data in a corresponding feature space after the corresponding augmented data are added; acquire an effective value of each augmented data according to the feature difference value and the range feature value corresponding to each augmented data. 3.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 2, characterized in that, The effective analysis submodule comprises a feature difference analysis unit, and the feature difference analysis unit is configured to: acquire a first feature vector of each augmented data after feature extraction, and a second feature vector of each fault data after feature extraction; analyze a difference degree between the first feature vector of each augmented data and a plurality of second feature vectors associated with the augmented data, to obtain a plurality of vector difference degrees corresponding to each augmented data; determine a mean value of the plurality of vector difference degrees corresponding to each augmented data as a feature difference value corresponding to the augmented data. 4.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 3, characterized in that, The acquisition of the first feature vector of each augmented data after feature extraction, and the second feature vector of each fault data after feature extraction comprises: acquiring a plurality of first feature values of each augmented data after feature extraction, and a plurality of second feature values of each fault data after feature extraction, wherein the plurality of first feature values and a plurality of preset data dimensions are in one-to-one correspondence, and the plurality of second feature values are also in one-to-one correspondence with the plurality of data dimensions; splicing the plurality of first feature values of each augmented data after feature extraction to obtain a first splicing vector of each augmented data, and splicing the plurality of second feature values of each fault data after feature extraction to obtain a second splicing vector of each fault data; The first spliced vector of each augmented data is normalized to obtain a first feature vector corresponding to each augmented data, and the second spliced vector of each fault data is normalized to obtain a second feature vector of each fault data. 5.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 2, characterized in that, The effective analysis submodule includes a range analysis unit, which is configured to: Obtain a plurality of first feature values of each augmented data after feature extraction, and a plurality of second feature values of each fault data after feature extraction, wherein the plurality of first feature values and a plurality of preset data dimensions correspond one by one, and the plurality of second feature values also correspond one by one with the plurality of data dimensions; Aggregate the first feature value of each augmented data at each data dimension and the plurality of second feature values associated with the data dimension to obtain a first feature set corresponding to each augmented data at each data dimension; and aggregate the plurality of second feature values associated with each augmented data at each data dimension to obtain a second feature set corresponding to each augmented data at each data dimension; Analyze the difference between the first feature set and the second feature set corresponding to each augmented data at each data dimension to obtain a range feature value corresponding to each augmented data. 6.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 5, characterized in that, The step of analyzing the difference between the first feature set and the second feature set corresponding to each augmented data at each data dimension to obtain a range feature value corresponding to each augmented data includes: Determine a dimension range value corresponding to each augmented data at each data dimension according to the ratio of the range of the first feature set and the second feature set corresponding to each augmented data at each data dimension; Determine the mean value of the plurality of dimension range values corresponding to each augmented data as the range feature value corresponding to each augmented data, wherein the plurality of dimension range values correspond one by one with the plurality of data dimensions. 7.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 2, characterized in that, The effective analysis submodule includes an effective value calculation unit, which is configured to: Normalize the feature difference value corresponding to each augmented data to obtain a first normalized value corresponding to each augmented data, and normalize the range feature value corresponding to each augmented data to obtain a second normalized value corresponding to each augmented data; The product of the first normalized value and the second normalized value corresponding to each augmented data is determined as the effective value of each augmented data. 8.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 1, characterized in that, The analysis module includes an invalid analysis submodule, which is configured to: Perform difference analysis on each first time series data of each augmented data and the plurality of second time series data associated with the corresponding data category to obtain a plurality of data difference values corresponding to each augmented data at each data category, wherein the first time series data is the time series data of the corresponding augmented data at the corresponding data category, and the second time series data is the time series data of the fault data associated with the corresponding augmented data at the corresponding data category; From the plurality of data difference values corresponding to each augmented data at each data category, determine the smallest data difference value as the difference index value of each augmented data at the data category; An average of the plurality of difference indicator values of each augmented data in the plurality of data categories is calculated to obtain an invalid value of each augmented data. 9.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 1, characterized in that, The determining module is configured to: The valid value and the invalid value of each augmented data are normalized respectively to obtain a normalized valid value and a normalized invalid value of each augmented data. A ratio of the normalized valid value and the normalized invalid value of each augmented data is determined as a retention value of each augmented data. 10.The AI-based electric power engineering safety quality intelligent diagnosis system according to claim 1, characterized in that, The training module comprises a training set construction unit, which is configured to: According to the retention value of each augmented data, the plurality of augmented data are sorted in descending order to obtain an augmented data sequence; Starting from the first position of the augmented data sequence, augmented data are sequentially removed so that the quantity ratio between a first sample quantity and a second sample quantity is a preset ratio, wherein the first sample quantity is the total quantity of the plurality of fault data and the augmented data remaining after removal, and the second sample quantity is the quantity of the plurality of normal data; The plurality of fault data and the augmented data remaining after removal are taken as positive samples, and the plurality of normal data are taken as negative samples to construct a training set.
Citation Information
Patent Citations
High-voltage circuit breaker fault diagnosis method based on generative adversarial network
CN113505876A
Electric power new energy equipment fault diagnosis method and device based on small samples
CN117992899A
BCTGAN data expansion method for extreme unbalanced data fault diagnosis
CN118277839A
Data expansion and optimization method and equipment for transformer fault diagnosis
CN119179899A
Fault diagnosis method and device for equipment lubricating system, equipment and medium
CN119903454A