Transformer fault early warning method based on multi-modal data fusion and dynamic weight
By collecting multimodal data and performing differentiated preprocessing and feature extraction, a dynamic weight allocation module is constructed. Combined with an optimized extreme learning machine model, the problems of single data source and insufficient weight adaptability in transformer fault early warning methods are solved, realizing accurate early warning and intelligent operation and maintenance of transformer faults.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA THREE GORGES UNIV
- Filing Date
- 2026-03-04
- Publication Date
- 2026-07-10
AI Technical Summary
Existing transformer fault early warning methods rely on single data sources, and the fusion of multi-source data has poor generalization and insufficient weight adaptability, resulting in early warning accuracy, timeliness and adaptability that are difficult to meet the operation and maintenance needs of modern power systems.
Multimodal heterogeneous data is collected, and a dynamic weight allocation module is constructed through differentiated preprocessing and multi-branch feature extraction. Combined with an optimized extreme learning machine model, fault early warning is achieved, and the model parameters are dynamically updated to adapt to changes in operating conditions.
It enables accurate detection and timely early warning of transformer faults, improves the reliability and adaptability of early warning, reduces false alarm rate and missed alarm rate, and supports intelligent management of transformer operation and maintenance.
Smart Images

Figure CN122365323A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of transformer fault early warning technology, and in particular to a transformer fault early warning method based on multimodal data fusion and dynamic weighting. Background Technology
[0002] As a core hub in the power system's generation, transmission, distribution, and consumption processes, the safe and stable operation of transformers directly determines the reliability of power supply and is crucial for ensuring the overall operational efficiency of the power system and preventing power accidents. During long-term continuous operation, transformers are susceptible to various faults due to internal factors such as insulation aging, partial discharge, abnormal core magnetic circuits, and oil deterioration. They also face significant fluctuations in external environmental temperature and humidity, electromagnetic interference and dust corrosion around substations, and changes in operating conditions such as sudden load changes and voltage level fluctuations. If transformer faults are not detected and addressed in a timely manner, they can lead to equipment damage and partial power outages, or even trigger cascading grid failures, causing widespread power outages and resulting in huge economic losses. Furthermore, they can have severe negative impacts on industrial production, residential lives, and other aspects of society. Therefore, accurate and timely fault warnings for transformers are of paramount importance.
[0003] In existing technologies, transformer fault early warning methods often rely on analysis of single-type monitoring data. For example, they may rely solely on oil chromatography to detect dissolved gas content in the oil to determine faults, on electrical quantity monitoring to analyze voltage and current anomalies, or on vibration signals to identify mechanical faults. For instance, while CN119312124A, a transformer fault early warning method and system, employs the concept of multi-source data fusion, in practical applications, insufficient data sources can still lead to incomplete information and thus, early warning bias. While CN120448913A, a transformer fault prediction method and related device based on multi-source data fusion, aims for multi-source data integration, it still suffers from legacy issues of traditional single-data monitoring models. For example, relying solely on specific types of data for preliminary analysis can still affect the subsequent fusion effect to some extent. Such single-data monitoring models struggle to comprehensively and accurately reflect the true operating state of the transformer, fail to capture fault precursors under the coupled effects of different factors, and are prone to early warning bias due to incomplete information, making it difficult to meet the comprehensive requirements of modern power systems for transformer state perception.
[0004] Meanwhile, existing early warning methods that attempt to integrate multi-source data generally employ a fixed-weight approach. This means that the weight coefficients for each type of data are pre-set and remain unchanged, failing to dynamically adjust the importance of each data point for fault diagnosis based on the actual operating conditions of the transformer. For example, while CN119312124A and CN120448913A mention multi-source data integration, they do not explicitly mention a technical solution for dynamically adjusting weights based on operating conditions. However, transformer operating conditions are constantly changing, and the contribution of different modal data to fault diagnosis varies significantly under different conditions. For instance, real-time operating data under high load conditions has a much higher reference value than historical maintenance data, and environmental data under high humidity and high temperature environments has a significantly increased impact on fault diagnosis. Fixed-weight integration methods cannot adapt to these complex and variable operating conditions, resulting in poor generalization and insufficient weight adaptation of multi-source data integration. In practical applications, this significantly reduces the accuracy and timeliness of fault early warning.
[0005] Furthermore, existing transformer fault early warning technologies have several shortcomings in data processing and feature extraction: denoising and completion methods for time-series operational data are relatively crude, easily losing subtle fault features; effective structuring conversion of unstructured historical maintenance text data is difficult, preventing the full exploitation of fault patterns contained within the text; environmental data processing fails to consider its dynamic trends, making it impossible to accurately capture the impact of environmental factors on equipment performance degradation; and feature extraction for various data types lacks a targeted architecture, resulting in insufficient mining of fault-related features for different data types, further reducing the reliability of fault early warning. For example, while CN119312124A and CN120448913A mention data preprocessing and feature extraction, they do not elaborate on how to adopt optimal processing and extraction methods for different types of data.
[0006] With the continuous advancement of smart grid construction, the power system's requirements for intelligent and refined transformer operation and maintenance are increasing. The aforementioned technical defects of traditional fault early warning methods have become a key bottleneck restricting accurate early warning of transformer faults. There is an urgent need to develop a transformer fault early warning method that can achieve efficient fusion of multimodal data and dynamic adjustment of weights according to operating conditions, so as to improve the accuracy, reliability and adaptability of fault early warning and provide strong protection for the safe and stable operation of transformers.
[0007] In power systems, transformers, as core equipment for energy conversion and transmission, directly impact the reliability and economy of the entire power grid through their safe and stable operation. However, with the expansion of power grid scale and the increasing complexity of the operating environment, the risk of transformer failures has significantly increased. Internal factors such as insulation aging, partial discharge, and core failures, as well as external factors such as temperature, humidity, and electromagnetic interference, can all be significant causes of transformer failures. Once a failure occurs, it can not only damage the equipment but also trigger a chain reaction, causing widespread power outages and other serious consequences, resulting in enormous losses to social and economic development.
[0008] Traditional transformer fault early warning methods mainly rely on single-type data monitoring and analysis, such as oil chromatography analysis and electrical quantity monitoring. While these methods can reflect the transformer's operating status to some extent, their reliance on a single data source makes it difficult to comprehensively and accurately capture the complex characteristics of transformer faults. For example, although oil chromatography analysis can detect changes in the composition and concentration of dissolved gases in transformer oil, it may not be effective in identifying certain early faults or non-oil circuit faults; and single-factor electrical quantity monitoring may be affected by factors such as load changes and voltage fluctuations, leading to false alarms or missed alarms.
[0009] Furthermore, traditional methods for multi-source data fusion often employ fixed weights, failing to dynamically adjust the importance of each data source based on the actual operating conditions of the transformer. In complex and ever-changing operating environments, this fixed-weight fusion approach struggles to adapt to variations in fault characteristics under different conditions, significantly reducing the accuracy and timeliness of fault warnings. For instance, in high-temperature and high-humidity environments, the impact of environmental factors on transformer performance may be more pronounced. If the original fixed-weight data fusion method is still used in such cases, it may fail to accurately reflect the degree of environmental influence on faults.
[0010] In summary, existing transformer fault early warning methods suffer from problems such as single data sources, poor generalization of multi-source data fusion, and insufficient weight adaptability, making it difficult to meet the high requirements of modern power systems for transformer operation and maintenance. Therefore, developing a transformer fault early warning method that can comprehensively consider multimodal data and achieve dynamic weight adjustment has significant practical significance and application value. Summary of the Invention
[0011] The technical problem to be solved by this invention is to provide a transformer fault early warning method based on multimodal data fusion and dynamic weighting. This method addresses the technical issues of traditional transformer fault early warning methods relying on single data sources, poor generalization of multi-source data fusion, and insufficient weight adaptability, which make it difficult to meet the accuracy, timeliness, and adaptability of early warnings to meet the needs of modern power system operation and maintenance. The invention enables accurate perception and timely early warning of faults under complex transformer operating conditions, improves the reliability and adaptability of fault early warning, and provides strong support for transformer operation and maintenance decision-making.
[0012] To achieve the above technical objectives, the present invention adopts the following technical solution: This invention comprehensively collects multimodal heterogeneous data of transformers and performs differentiated and refined preprocessing. It extracts fault-related features of various types of data through a multi-branch architecture, and builds a dynamic weight allocation module that adapts to operating conditions based on an improved intelligent algorithm to achieve accurate fusion of multimodal features. Finally, it completes accurate prediction of transformer fault probability and type based on an optimized extreme learning machine model, and dynamically updates model parameters according to changes in operating conditions to ensure the real-time performance and accuracy of early warning results.
[0013] Specifically, the transformer fault early warning method of the present invention includes the following steps: Step 1: Multimodal data acquisition: Collect real-time operating data, environmental data, and historical maintenance data of the transformer. Real-time operating data is acquired by a sensor array, historical maintenance data is exported from the operation and maintenance system, and environmental data comes from meteorological stations and electromagnetic interference monitoring equipment, so as to achieve full-dimensional data coverage of transformer operating status, external influencing factors, and historical fault patterns.
[0014] Step 2, Differentiated Data Preprocessing: Preprocessing is performed separately for the three types of data based on their heterogeneous characteristics. For real-time running data, particle filtering is used to denoise and cubic spline interpolation is used to fill in missing values to ensure the continuity and integrity of time-series data. For historical maintenance data, the conditional random field algorithm is used to extract knowledge triples and clean the structured tables to achieve the structured transformation of unstructured text data. For environmental data, the exponential weighted moving average method is used to suppress fluctuations and accurately capture the changing trends of environmental data, laying a high-quality data foundation for subsequent feature extraction.
[0015] Step 3, Multi-branch Feature Extraction: Construct a multi-branch feature extraction architecture that includes a time-series branch, a text branch, and an environment branch. The time-series branch uses an improved convolutional neural network to extract time-series dependency features and subtle fault features from real-time running data. The text branch uses a topic model optimized by an improved word frequency-inverse document frequency algorithm to extract fault theme features from historical maintenance data. The environment branch uses a gradient boosting decision tree to extract the attenuation features of environmental data on transformer equipment performance. Each branch extracts features in parallel and outputs corresponding feature vectors to comprehensively mine fault-related information in various types of data.
[0016] Step 4, Dynamic Weight Feature Fusion: Based on the improved sparrow search algorithm, a dynamic weight allocation module is constructed. First, the three types of feature vectors are reduced to the same dimension through principal component analysis. Then, with the contribution of each modal feature to the fault as the optimization objective, the optimal dynamic weight coefficient of each modal feature is calculated iteratively. Through weighted fusion, a multimodal fusion feature matrix is obtained, which solves the problem that traditional fixed weights cannot adapt to the complex working conditions of transformers and realizes the working condition adaptation of feature fusion.
[0017] Step 5: Optimize the model for fault early warning: Construct an optimized extreme learning machine fault early warning model, using a multimodal fusion feature matrix as input, and outputting the transformer's fault probability and specific fault type; during model training, historical fault data and simulated fault data are used to optimize parameters, and a regularization term is introduced to improve the model's generalization ability and operational stability; when the change in transformer operating parameters exceeds the threshold, the dynamic weight allocation module is triggered to recalculate the weights and update the model, and finally output and visualize the early warning results, generate a report for storage, and provide maintenance personnel with direct and actionable early warning information.
[0018] The present invention provides a transformer fault early warning method based on multimodal data fusion and dynamic weighting, which has the following beneficial effects: 1. This invention breaks through the limitations of traditional transformer fault early warning relying on single data. It comprehensively collects multimodal heterogeneous data from real-time operation, external environment, and historical maintenance, achieving full-dimensional coverage of equipment operating status, external influencing factors, and historical fault patterns. It fully reflects the true health status of the transformer from the data source, avoiding early warning deviations caused by incomplete data.
[0019] 2. This invention designs differentiated preprocessing schemes for the heterogeneous characteristics of three types of data: time series, text, and environment. It overcomes the adaptability defects of traditional unified preprocessing, realizes refined processing of different types of data, greatly improves the effectiveness and usability of the original data, and lays a solid foundation of high-quality data for subsequent feature extraction.
[0020] 3. This invention uses particle filtering to denoise real-time runtime sequence data. By constructing a probabilistic model of system state and noise, it accurately estimates and eliminates noise. Compared with traditional filtering methods, it has a better effect on removing non-Gaussian noise and can effectively preserve subtle fault features in the data, avoiding the loss of fault precursor information.
[0021] 4. This invention combines cubic spline interpolation to fill missing values in the running data. By constructing a smooth spline function, it achieves continuous filling of missing points. Compared with traditional methods such as mean and linear interpolation, it is closer to the actual change trend of the data, ensuring the continuity and integrity of time series data and avoiding the impact of missing data on feature extraction results.
[0022] 5. This invention uses the Conditional Random Field algorithm to perform entity recognition and extract knowledge triples from maintenance text data, realizing the structured transformation of unstructured text data, enabling text information to express the transformer fault state in mathematical relationships, and solving the technical problem that traditional text data is difficult to participate in fault feature analysis.
[0023] 6. This invention uses the exponentially weighted moving average method to process environmental data, assigning higher weights to recent data according to time sequence. This not only accurately captures the dynamic changing trends of environmental temperature, humidity, electromagnetic interference, and other data, but also effectively suppresses random abnormal fluctuations, making the environmental data more consistent with the environmental impact patterns of actual transformer operation.
[0024] 7. This invention constructs a multi-branch feature extraction architecture, which allows time-series, text, and environment branches to extract features in parallel. Each branch has its own extraction logic designed for the characteristics of the corresponding data type, enabling targeted mining of fault-related features and significantly improving the comprehensiveness and accuracy of feature extraction.
[0025] 8. The temporal branch of this invention adopts an improved multi-scale convolutional neural network. Through the design of multi-size convolutional kernels and multi-layer network structure, it enhances the ability to extract temporal dependencies and subtle fault features in the running data, and can effectively capture fault precursor information that is easily ignored, such as continuous changes in oil temperature and small fluctuations in current.
[0026] 9. The text branch of this invention optimizes the topic model based on the improved word frequency-inverse document frequency algorithm, introduces a fault-related vocabulary weight enhancement factor, highlights the importance of core fault words such as insulation aging and partial discharge, and enables the topic model to more accurately mine fault topic features in maintenance data, thereby improving the correlation between text features and faults.
[0027] 10. The environmental branch of this invention uses a gradient boosting decision tree to extract environmental degradation features, which can accurately fit the complex nonlinear relationship between environmental data and transformer equipment performance degradation, clearly capture the influence of environmental factors such as high temperature and humidity and electromagnetic interference on equipment, and fill the gap in the traditional method's insufficient mining of environmental features.
[0028] 11. This invention designs a dynamic weight allocation module based on an improved sparrow search algorithm. It iteratively calculates weights with the fault contribution of each modal feature as the optimization target. Compared with the traditional fixed weight fusion method, it realizes the adaptive adjustment of weights according to the operating conditions and solves the problem that fixed weights cannot adapt to the complex and variable operating conditions of transformers.
[0029] 12. Before fusing features, this invention unifies the dimensions of each modality feature vector through principal component analysis, effectively reducing feature dimension redundancy and improving the computational efficiency of feature fusion. At the same time, it retains the core information of each modality feature to the greatest extent, ensuring the effectiveness and representativeness of the fused feature matrix.
[0030] 13. This invention constructs an optimized extreme learning machine fault early warning model, introduces a regularization term to optimize model parameters, effectively improves the model's generalization ability and operational stability, avoids overfitting problems, and enables the model to maintain good early warning performance in different substations and different types of transformers.
[0031] 14. The early warning model of this invention can simultaneously output the probability of transformer failure and the specific type of failure, enabling quantitative judgment of equipment health status and precise fault location. Compared with traditional methods that can only determine whether a fault exists, it provides maintenance personnel with more specific and actionable early warning information, greatly improving the efficiency of maintenance decision-making.
[0032] 15. This invention sets a dynamic update threshold for changes in operating parameters, and triggers weight recalculation and model update through real-time operating condition monitoring, so that the early warning model can adapt to changes in operating conditions such as transformer load rate and voltage level in real time, ensuring the real-time performance and accuracy of fault early warning, and solving the problem that traditional models are fixed and unchanging once trained and have poor adaptability to operating conditions.
[0033] 16. This invention combines multimodal data fusion with dynamic weight adjustment, which greatly improves the reliability of fault early warning, effectively reduces the false alarm rate and false negative rate of fault early warning, can accurately detect fault signs in the early stage of fault, realize early warning of transformer faults, and reserve sufficient fault handling time for maintenance teams.
[0034] 17. The early warning method of the present invention realizes full-process automation of data collection, preprocessing, feature extraction and fault early warning, replacing the traditional fault diagnosis method of manual analysis and experience judgment, greatly reducing the workload of operation and maintenance personnel, improving the intelligent level of transformer operation and maintenance, and effectively reducing labor cost investment.
[0035] 18. This invention can provide accurate early warning for various common transformer faults such as insulation aging, partial discharge, and core failure. It is applicable to a wide range of fault types and, compared with traditional single-fault-type early warning methods, fully meets the actual needs of modern power systems for early warning of all types of transformer faults.
[0036] 19. The method of the present invention is based on a mature algorithm framework and general hardware devices. It does not require additional customized special hardware, has low deployment cost and strong equipment adaptability, and can be quickly applied to transformers of different voltage levels and different models, which facilitates large-scale promotion and application in power systems.
[0037] 20. By enabling accurate and early warning of transformer faults, this invention can effectively avoid power accidents caused by transformer faults, reduce power outage losses and equipment maintenance and replacement costs in the power system, and at the same time improve the stability and reliability of power supply, resulting in significant economic benefits.
[0038] 21. This invention provides scientific and powerful support for transformer operation and maintenance decisions through accurate fault early warning, promotes the transformation of transformer operation and maintenance from traditional passive fault handling to proactive health management, provides core technical support for the construction and development of smart substations, and helps the overall intelligent upgrade of the power system.
[0039] 22. The multimodal data processing and dynamic weight fusion technology of this invention provides a novel solution for the status perception and fault early warning of other power equipment, expands the technical path of intelligent operation and maintenance of power equipment, and has good technical promotion and radiation value.
[0040] 23. The early warning method of the present invention can output and visualize the early warning results in real time, and generate and store early warning reports, realizing standardized and traceable management of fault early warning information, which facilitates maintenance personnel to review and analyze fault patterns and further optimize transformer maintenance strategies. Attached Figure Description
[0041] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart illustrating the transformer fault early warning method of the present invention; Figure 2 This is a schematic diagram illustrating the multi-source data preprocessing and feature extraction of transformers according to the present invention; Figure 3 This is a diagram showing the effect of time-series data preprocessing in this invention; Figure 4 This is a diagram of the improved multi-scale convolutional neural network structure of the present invention; Figure 5 This is the trimodal dynamic weight change curve based on the improved Sparrow Search Algorithm ISSA real-time optimization of the present invention. Detailed Implementation
[0042] The technical solutions of the present invention will be further described below with reference to the embodiments and accompanying drawings: Example 1 This embodiment provides a transformer fault early warning method based on multimodal data fusion and dynamic weighting, such as... Figure 1 As shown, the specific steps include the following: Step 1: Collect real-time transformer operation data: environmental data and historical maintenance data. Real-time operation data is acquired by a sensor array, environmental data comes from a weather station and electromagnetic interference monitoring equipment, and historical maintenance data is exported from the operation and maintenance system. Step 2: Preprocess the three types of data respectively: the running data is denoised by particle filtering and missing values are filled by cubic spline interpolation; the maintenance data is extracted with knowledge triples by the conditional random field algorithm, and the structured table is cleaned at the same time; the environmental data is suppressed by exponential weighted moving average. Step 3: Construct a multi-branch feature extraction architecture: The temporal branch uses an improved convolutional neural network to extract temporal features from the running data; the text feature extraction branch optimizes the topic model based on an improved term frequency-inverse document frequency algorithm, calculates the keyword weights in the fault report, and then inputs the fault features from the maintenance data; the environmental branch uses a gradient boosting decision tree to extract environmental impact features. Step 4: Design dynamic weight allocation module to fuse features: The working condition monitoring unit obtains parameters such as load rate, inputs the improved sparrow search algorithm to calculate the fault contribution of each modal feature, generates dynamic weight coefficients, and weights and fuses the three types of features to obtain the fused feature matrix; Step 5: Construct an extreme learning machine early warning model: The model takes a fusion feature matrix as input and outputs the fault probability and type. During training, historical and simulated fault data are used to optimize the parameters. When the change in operating parameters exceeds the threshold (set at 15%), the weights are recalculated to update the model. Finally, the early warning results are output and visualized, and a report is generated and stored. The goal is to be able to output the corresponding fault status when a single parameter or mixed parameter data is input.
[0043] The specific implementation process for each step is as follows: Step 1: Collect real-time transformer operating data, environmental data, and historical maintenance data. Real-time operating data is acquired by a sensor array, environmental data comes from weather stations and electromagnetic interference monitoring equipment, and historical maintenance data is exported from the operation and maintenance system. For the real-time operating data of the transformer itself, high-precision sensor arrays are installed on key parts of the transformer and its lines. These include, but are not limited to, current transformers, voltage transformers, platinum resistance temperature sensors, and piezoelectric vibration sensors, used to continuously collect the raw waveforms and RMS values of voltage, current, oil temperature, and vibration signals. These sensors are all connected to the data acquisition unit located in the main control room via shielded signal cables. The sampling frequency is configured differently according to the signal characteristics to ensure the accuracy and real-time performance of the data.
[0044] Historical maintenance data is exported from the power grid company's asset management system and production management system via a data interface. This data includes structured periodic test records and defect record forms, as well as unstructured text-based fault analysis reports and maintenance work logs. The exported data packets are transmitted to the local analysis server via a secure isolation device.
[0045] The acquisition of external environmental data relies on miniature weather stations and ultra-high frequency electromagnetic interference (UHF) monitoring probes deployed within the substation area. The weather stations are responsible for collecting environmental temperature, humidity, and atmospheric pressure data, while the electromagnetic probes are used to capture the intensity of electromagnetic interference in specific frequency bands around the substation. All collected multimodal data are tagged with a unified time stamp and temporarily stored on edge computing nodes, laying the foundation for subsequent preprocessing and fusion analysis.
[0046] Step 2: Preprocess the collected multi-source data, such as... Figure 2 As shown, this invention performs preprocessing on three types of data, categorizing them accordingly. Specifically, this includes: Step 2.1, Time Series Data Processing: For the collected real-time operating time series data of the transformer, including the time-varying sequence data of voltage, current, oil temperature, vibration, etc., this data is used as input for particle filtering denoising. During the particle filtering denoising process, a probabilistic model of the system state and noise is constructed. The system state transition model is expressed as: (6); in, , They are respectively , Real-time transformer operating system status; This is the system state transition function; for Time-based system process noise.
[0047] This probabilistic model is used to estimate and eliminate noise in the running data, resulting in denoised time-series data. If missing values exist in the denoised time-series data, cubic spline interpolation is used to fill in the missing parts. Let the interval containing the missing data points be denoted as . Construct a cubic spline function ,satisfy = , = ,and = , = ,in , The second derivative is obtained by solving the system of equations: (7); in, , Derivatives , The coefficient; , , These are the second derivative values of the previous node, the current node, and the next node of the missing data interval, respectively. Given constants, a cubic spline function is obtained after solving. The missing running data is filled in to ensure the continuity and integrity of the data, and finally the preprocessed time series data is obtained.
[0048] Step 2.2: When the Conditional Random Field algorithm extracts knowledge triples from text data, it first performs word segmentation and part-of-speech tagging on the unstructured fault reports to obtain a sequence. Then, the labeled sequence is calculated based on the conditional random field model. conditional probability As shown below: (8); in, Normalization factor; The length of the text sequence; This represents the total number of transition characteristic functions; For the first The feature weights of each transition feature function; For the first One transition characteristic function; , The first , The tag value at each position; The total number of state characteristic functions; For the first Feature weights of each state feature function; For the first Each state feature function is used. This model identifies transformer-related entities and the relationships between them, forming knowledge triples. This achieves the effect of processing transformer text data, allowing the expression of transformer fault states using mathematical relationships during feature fusion.
[0049] Step 2.3: When processing environmental data using the index-weighted moving average, let... Time-based environmental data is Its weighted moving average The calculation is as follows: (1); in, for Raw environmental data collected in real time. for Weighted moving average of environmental data over time. The range of values for the weighting coefficients is: , for Time-weighted moving average. Different weights are assigned to data based on their chronological order, with more recent data receiving greater weight. Higher weighting allows for better capture of changing trends in environmental data while suppressing abnormal fluctuations. During data preprocessing, steps 2.1, 2.2, and 2.3 are performed simultaneously, with the following processing results: Figure 3 As shown.
[0050] Step 3: The present invention constructs a multi-branch feature extraction method, the specific method of which is described as follows: Step 3.1: Extract the temporal features of the running data using an improved convolutional neural network. The network structure is as follows: Figure 4 As shown, the improved convolutional neural network includes an input layer, convolutional layers, pooling layers, and fully connected layers. Let the timing of the input processing data be... ,in The temporal length is specified. Convolutional layers use kernels of various sizes for convolution operations. One convolutional kernel (size is...) ×1) The formula for convolving the input time series is: (9); in, For the first The convolutional kernel at the _th ... The convolution output values at each position; For the first The size of each convolutional kernel; For the first The nth convolutional kernel The weight value of each position; For input timing data in the first... The value at each position; For the first The bias values of each convolution kernel; This is the activation function for the neural network. The size of the convolution kernel is adjusted. The number of convolutional kernels of different sizes and the number of network layers increase the adjustable dimension of the network, enhancing its ability to extract temporal dependencies in operational data, such as the continuous trend of oil temperature changes and the pulse interval pattern of partial discharge signals, as well as subtle features, such as minute fluctuations in current and slight changes in gas concentration. Max pooling is used in the pooling layer to downsample the output of the convolutional layer, using the following formula: (10); in, For pooling step size, For the first The feature map corresponding to the nth convolutional kernel is at the nth convolutional kernel. The output value after pooling at each position; For the first Each convolutional kernel outputs all values within the pooling window of the feature map. Finally, a fully connected layer maps the extracted features to a fixed-dimensional feature vector. ( (For feature dimensions).
[0051] Step 3.2: Optimize the topic model using the improved term frequency-inverse document frequency algorithm to extract fault topic features from maintenance data. First, analyze the fault report set... ( The vocabulary set is obtained by segmenting the data into words based on the number of reports. ( (Total number of words). Improved term frequency - inverse document frequency (...) In the Improved Term Frequency-Inverse Document Frequency (IF) algorithm, for vocabulary... In the document word frequency in The calculation is as follows: (11); in, For vocabulary In the document The number of times it appears in; For document Total number of words in; For document The Middle The number of times each word appears. Inverse document frequency. The calculation is as follows: (12); in, For vocabulary Inverse document frequency; This represents the total number of documents in the fault report set, i.e., the number of reports. For words containing The number of documents is increased by 1 to avoid a denominator of 0. Additionally, a weighting enhancement factor for fault-related terms is introduced. Based on the experience of experts in the field, the core fault terms such as "insulation aging" and "partial discharge" are... If we set the value to 1.5~2.0, and set the value of common vocabulary to 1.0, then the improved version... The value is: (2); in, For vocabulary In the document Improved value; For vocabulary In the document Word frequency in; For vocabulary Inverse document frequency; For vocabulary The fault-related weight enhancement factor. The improved... Matrix Input Topic Model (LDA model, Latent Dirichlet Allocation) assumes that each document... Multiple themes ( The document-topic distribution is a mixture of the number of topics. Model parameters are learned through Gibbs sampling to obtain the document-topic distribution. Thematic-vocabulary distribution For the document Its fault theme feature vector ( (As a feature dimension) derived from document-topic distribution The importance of fault-related terms is highlighted to enable the topic model to more accurately mine fault theme features in maintenance data, and to identify typical fault description themes corresponding to different fault types.
[0052] Step 3.3: Extract environmental impact features using a gradient boosting decision tree. Let the environmental data be... ( (These are environmental variables, including temperature, humidity, wind speed, etc.), and the transformer equipment performance degradation index is... (Including insulation resistance attenuation rate, oil gas growth rate, etc.). Gradient Boosting Decision Tree (GBDT) fits the data by constructing multiple decision trees and progressively improving their performance. and The relationship between them. First, initialize a constant value. ,in The loss function is then generated iteratively. The first decision tree, the Predictions from decision trees Fit the negative gradient of the current loss function: (13); in, For the first The sample at the th The residual of each iteration, one tree corresponds to one iteration; Indicates the reciprocal; No. The true value of the performance degradation index of transformer equipment for each sample; For the first During the nth iteration, the model... The predicted values for each sample are calculated. The decision tree is trained by minimizing the sum of squared residuals. To obtain the leaf node region of the tree ( , For the first (the number of leaf nodes in the tree), and for each leaf node Calculate the optimal output value: (14); in, For the first The first decision tree The optimal output value for each leaf node; To minimize the value of the subsequent summation value; The output values to be optimized are for the leaf nodes.
[0053] The final prediction model is: (15); in, for The final predicted value output by the model composed of decision trees; These are the initial constant values predicted for the model; For the first The learning rate of each decision tree; For the first The predicted values of multiple decision trees are obtained. By stepwise fitting of multiple decision trees, the complex nonlinear relationship between environmental data and transformer equipment performance degradation is captured, such as the accelerated degradation of insulation performance under high temperature and humidity conditions, and the degree of influence of electromagnetic interference on equipment operation stability. This yields a feature vector of environmental factors affecting equipment performance degradation. ( (as dimensional features).
[0054] Step 4: Fuse the multi-source data after feature extraction. Design a dynamic weight allocation module based on the improved Sparrow Search Algorithm (ISSA) to achieve accurate fusion of multi-modal features by iteratively calculating the fault contribution of each modality feature under the current working condition, and finally forming a multi-modal fusion feature matrix.
[0055] After feature extraction in step 3, the temporal feature vector is obtained. Text feature vectors Environmental characteristics ,in These represent the dimensions of each modal feature. The Improved Sparrow Search Algorithm (ISSA) simulates the foraging and vigilance behavior of a sparrow population. Using the contribution of each modal feature to the fault as the optimization objective, iteratively updates the population position to find the optimal weight allocation. The changing trend of each modal weight over monitoring time is shown below. Figure 5 As shown. Initialize the sparrow population location matrix. ,in For population size, The search space dimension (corresponding to the feature weights and related parameters of each modality).
[0056] Step 4.1, Update producer location: Producers, representing approximately 20% of the population and considered the best individuals, are responsible for exploring the globally optimal region. Their position update formula is: (16); In the formula, , For the first , During the nth iteration The first individual sparrow in the... The position value of the dimension; This represents the current iteration number. The maximum number of iterations, For control parameters, The warning value, As a safety threshold, For random numbers that follow a normal distribution, It is a vector consisting entirely of 1s. At that time, producers expand their search area within the safe zone; when At that time, producers move closer to dangerous areas and adjust their search strategies.
[0057] Step 4.2, Follower position update: Followers constitute the majority of the population and adjust their positions based on the positions of producers and the worst-performing individual, using the following formula: (17); in, For the first Iteration number The worst position of the dimension; For the first The optimal position of the producer in the next iteration. The numerical value of the dimension; The total number of sparrows planted with trees; These are random numbers that follow a standard normal distribution. If an individual If the population is in the bottom 50%, learn from the worst individual and try to escape the local optimum; if it is in the top 50%, learn from the producers and refine the local search.
[0058] Step 4.3, Update the location of the vigilant: The vigilant population comprises approximately 10% and performs local perturbations around the globally optimal position, as shown in the formula: (18); In the formula, For the first The second iteration The optimal position of the dimension. The step size control parameter is to follow a uniform distribution U(0,1) to avoid the algorithm getting trapped in local optima.
[0059] The Pearson correlation coefficient was used, with the correlation between each modal feature and the fault label serving as the fitness function. Represented as: (4); In the formula, Based on feature weights The fitness function value of the variable; It is a multimodal fusion feature matrix With fault labels covariance; Multimodal fusion feature matrix Standard deviation; Fault Label The standard deviation.
[0060] Temporal feature weights are obtained through ISSA iterative optimization. Text feature weights Environmental feature weights ,satisfy To unify feature dimensions, dimensionality matching was performed on the features of each modality, and principal component analysis was used to... Dimensionality reduction to the same dimension ,get The final multimodal fusion feature matrix The calculation is as follows: (3); In the formula, This is a multimodal fusion feature matrix; This is the time-series feature vector of real-time running data after dimensionality reduction by principal component analysis; This represents the fault theme feature vector of historical maintenance data after dimensionality reduction by principal component analysis. This is the environmental data decay feature vector after dimensionality reduction by principal component analysis.
[0061] Step 5: Construct a fault early warning model based on extreme learning machine optimization, combine it with real-time operating condition monitoring feedback to achieve dynamic weight updates, and finally generate accurate fault early warning results. The specific process is as follows: The fault warning model adopts a single-hidden-layer feedforward neural network structure, consisting of an input layer, a hidden layer, and an output layer. The input to the model is the multimodal fusion feature matrix obtained in step 4. ,in Represents the number of samples. Representing the dimensions of the fused features, this matrix encompasses comprehensive information from three categories of features: temporal, textual, and environmental; hidden layer settings. Each node uses the Sigmoid function as the activation function to perform nonlinear mapping of the input features. The final output of the output layer consists of two types of core information: first, the probability of the transformer's current fault occurrence, ranging from 0 to 1, with the value closer to 1 indicating a higher fault risk; and second, the specific fault type, such as insulation aging, partial discharge, and core fault, which corresponds to the preset fault categories, thereby achieving quantitative judgment and precise positioning of the equipment's health status.
[0062] During the model training phase, the hidden layer input weights and bias parameters are first randomly generated, and the hidden layer output matrix is calculated based on these parameters. Specifically, for the first... The sample, which is in the th , The output of each hidden layer node is calculated by multiplying and adding the weights corresponding to that node to the feature values of each dimension of the sample element by element, adding the bias term, and then applying the Sigmoid activation function. The outputs of all nodes together constitute the hidden layer output matrix.
[0063] Subsequently, a fault label matrix using one-hot encoding was used. For training purposes, this invention uses the labels "1, 0, 0" to correspond to "insulation aging" and "0, 1, 0" to correspond to "partial discharge". The output weights are key parameters connecting the hidden layer and the output layer, and must satisfy the following... The relationship between the hidden layer output matrix and the final fault result can be obtained by solving the Moore-Penrose generalized inverse. The output weights are the core parameters inside the model and are used to establish the mapping relationship between the intermediate features and the final fault result.
[0064] To further improve the model's generalization ability and operational stability, a regularization term is introduced to optimize the model. A loss function is constructed with the output weights as variables. This function considers both the norm of the output weights and the prediction error of all samples, and is optimized using regularization parameters. Balancing the relationship between the two, the optimized hidden layer output weights are finally obtained. The weight calculation formula is as follows: (5); In the formula, It is the output weight matrix from the hidden layer to the output layer of the Extreme Learning Machine model; This represents the output matrix of the hidden layer in the Extreme Learning Machine model. Output matrix for hidden layer The transpose of the matrix; This is the model regularization parameter, used to balance prediction error and weight norm; For the identity matrix, the dimensions are... Consistent; This is a fault label matrix.
[0065] During the actual operation of the model, key operating parameters such as transformer load rate and voltage level are acquired in real time through the operating condition monitoring unit. When the monitored change in operating parameters exceeds the threshold (set at 15%), the dynamic weight allocation module in step 4 is immediately triggered to recalculate the weight coefficients of each modal feature. The trend of weight change is as follows: Figure 5 As shown, the multimodal fusion feature matrix is updated simultaneously, and the updated matrix is input into the optimized extreme learning machine model.
[0066] Based on new input data, the model uses a complete reasoning process involving the input layer, hidden layer, and output layer to ultimately output the probability of transformer failure and the specific type of failure under the current operating conditions. For example, a failure probability of 0.92 indicates a high failure risk, and a diagnostic conclusion of partial discharge failure is given, thus providing direct and actionable early warning information for maintenance team personnel.
[0067] This invention presents a novel transformer fault early warning method based on multimodal data fusion and dynamic weighting. This method overcomes the shortcomings of traditional methods, such as poor generalization of multi-source heterogeneous data fusion, insufficient adaptability to operating conditions due to fixed weights, and low utilization of single-modal information. It significantly improves early warning accuracy, robustness, real-time performance, and interpretability. By implementing this invention, accurate probability prediction and type identification of various faults, including overheating, discharge, moisture, and core faults, can be achieved. False alarm and false negative rates are significantly reduced, and maintenance shifts from passive to proactive early warning, resulting in substantial economic and safety benefits. This invention provides a core technological engine for intelligent substation status perception and proactive health management.
[0068] This invention demonstrates significant advantages in key aspects such as multimodal preprocessing, heterogeneous feature extraction, and adaptive dynamic weight allocation under operating conditions, providing a novel, efficient, and practical solution for intelligent early warning of multiple fault types in power transformers with high reliability.
[0069] Example 2 In another preferred embodiment, based on Embodiment 1, this embodiment provides a transformer fault early warning method based on multimodal data fusion and dynamic weighting, the specific steps of which are as follows: This embodiment uses a 110kV substation main transformer as the application object to describe in detail the transformer fault early warning method based on multimodal data fusion and dynamic weighting of the present invention, and to verify the feasibility and effectiveness of the method in a real engineering scenario. The algorithms used in this embodiment are all implemented using Python language combined with frameworks such as PyTorch and Scikit-learn. The hardware environment is an Intel Xeon E5 server with 64GB of memory. The overall implementation process is as follows: Figure 1 As shown.
[0070] Step 1: Multimodal Data Acquisition For 110kV main transformers, according to Figure 2 The system employs a multi-source data acquisition logic, deploying a sensor array to collect real-time operational data. This includes current transformers, voltage transformers, platinum resistance oil temperature sensors, and piezoelectric vibration sensors, acquiring voltage, current, top oil temperature, and body vibration signals at sampling frequencies of 50Hz, 50Hz, 1Hz, and 100Hz, for a continuous 720-hour period. It also exports nearly five years of historical maintenance data for the transformer from the substation's operation and maintenance management system, including structured oil chromatography test records and defect handling forms, as well as unstructured fault analysis reports and maintenance work logs, totaling 236 documents. Environmental data, including ambient temperature, relative humidity, atmospheric pressure, and electromagnetic interference intensity, is collected via a miniature weather station and UHF electromagnetic interference monitoring probes within the substation, at a sampling frequency of 1Hz. All collected data is time-stamped and stored on edge computing nodes, providing a foundation for subsequent preprocessing and feature extraction.
[0071] Step 2: Differentiated Preprocessing of Multimodal Data This step is as follows Figure 2 The preprocessing workflow performs targeted and refined preprocessing on three types of heterogeneous data. Each preprocessing step is executed in parallel, and the processing results are as follows: Figure 3 As shown.
[0072] Step 2.1: Real-time Data Preprocessing Time-series data of voltage, current, oil temperature, and vibration are input into a particle filter denoising model to construct a system state transition probability model. Sensor noise and electromagnetic interference noise are estimated and eliminated to complete the denoising process. For missing values in the denoised data caused by sensor disconnection, cubic spline interpolation is used to construct a smooth spline function to fill in the missing points. Figure 3 As shown in (c), the continuity and integrity of the data are effectively guaranteed, and a standardized real-time runtime sequence dataset is finally obtained.
[0073] Step 2.2: Preprocessing of historical maintenance data Unstructured fault reports and maintenance logs were processed using natural language processing, including word segmentation and part-of-speech tagging. A CRF (Conditional Random Field) entity recognition algorithm was used to extract transformer fault-related knowledge triples, such as "transformer-occurrence-partial discharge" and "insulating bushing-aging-insulation breakdown". A total of 892 valid knowledge triples were extracted. At the same time, the structured tables were cleaned to remove outliers and duplicates, and the text data was transformed into a structured data set, resulting in a standardized knowledge triple dataset.
[0074] Step 2.3: Environmental Data Preprocessing The exponentially weighted moving average method is used to process environmental data. The weighting coefficient α = 0.7 is set, and the weighted moving average is calculated according to formula (1). This effectively suppresses random fluctuations in environmental data and accurately captures the changing trends of environmental temperature, humidity, and electromagnetic interference, resulting in a preprocessed smooth environmental dataset. (1); In the formula, for Raw environmental data collected in real time. for Weighted moving average of environmental data over time. The weighting coefficients and , for The weighted moving average of environmental data at any given time.
[0075] Step 3: Multi-branch feature extraction according to Figure 2 The multi-branch feature extraction architecture constructs time series, text, and environment branches respectively, extracts fault-related features for various types of data, and finally outputs a feature vector with unified dimensions.
[0076] Step 3.1: Temporal Feature Extraction The temporal branch is constructed using an improved multi-scale CNN, with the network structure as follows: Figure 4 As shown, the network consists of an input layer, three convolutional layers, two max-pooling layers, and a fully connected layer. The convolutional layers use 3×1, 5×1, and 7×1 kernels of various sizes, with ReLU as the activation function and a pooling stride of 2. Preprocessed real-time operating data is input into this network, and multi-scale convolution fully extracts temporal-dependent features and subtle fault features such as the continuous trend of oil temperature changes, the pulse interval of partial discharge signals, and minute current fluctuations. Finally, a 64-dimensional temporal feature vector is obtained. .
[0077] Step 3.2: Fault Theme Feature Extraction The text branch employs an improved TF-IDF algorithm to optimize the LDA topic model, setting weight enhancement factors for core fault terms (such as insulation aging, partial discharge, and core fault). Common vocabulary The word weights are calculated according to formula (2); the weighted TF-IDF matrix is input into the LDA model, the number of topics is set to K=8, the model parameters are learned through Gibbs sampling, the fault topic features in the maintenance data are deeply mined, and a text feature vector with a dimension of 64 is obtained. : (2); In the formula, For vocabulary In the document Improved value; For vocabulary In the document Word frequency in; For vocabulary Inverse document frequency; For vocabulary Fault-related weight enhancement factor.
[0078] Step 3.3: Extraction of environmental attenuation features The environmental branch employs a GBDT decision tree with 100 decision trees, a learning rate of γ=0.1, and a mean squared error loss function. Using preprocessed environmental data as input, and transformer insulation resistance attenuation rate and oil gas growth rate as equipment performance degradation indicators, a complex nonlinear relationship between environmental factors and equipment performance degradation is fitted. Attenuation features of environmental influences such as high temperature and humidity, and electromagnetic interference are extracted, resulting in a 64-dimensional environmental feature vector. .
[0079] Step 4: Dynamic weight fusion based on the improved sparrow search algorithm Principal Component Analysis (PCA) was used to verify the dimensionality of the three types of feature vectors mentioned above (in this embodiment, the dimensionality has been standardized to 64 dimensions, so no additional dimensionality reduction is required); an ISSA dynamic weight allocation module was constructed, with the population size N=30 and the maximum number of iterations set. With a safety threshold ST=0.8, the Pearson correlation coefficient between each modal feature and the fault label is used as the fitness function to iteratively calculate the fault contribution of each modal feature. The trend of the change of each modal weight with monitoring time is as follows: Figure 5 As shown, in this embodiment, the temporal feature weights are obtained after iterative convergence. Text feature weights Environmental feature weights ,satisfy , The fitness function is calculated using the following formula: (4); In the formula, Based on feature weights The fitness function value of the variable; It is a multimodal fusion feature matrix With fault labels covariance; Multimodal fusion feature matrix Standard deviation; Fault Label The standard deviation.
[0080] The three types of feature vectors are weighted and fused according to formula (3) to obtain a 64-dimensional multimodal fusion feature matrix. This enables accurate fusion of multimodal features, providing comprehensive feature input for fault early warning models. (3); In the formula, This is a multimodal fusion feature matrix; This is the time-series feature vector of real-time running data after dimensionality reduction by principal component analysis; This represents the fault theme feature vector of historical maintenance data after dimensionality reduction by principal component analysis. This is the environmental data decay feature vector after dimensionality reduction by principal component analysis.
[0081] Step 5: Fault warning and dynamic update based on optimization of the extreme learning machine Step 5.1: Early Warning Model Construction and Training according to Figure 1 The model training process is as follows: an optimized Extreme Learning Machine (ELM) fault warning model is constructed, using a single hidden layer feedforward neural network structure. The input layer is 64-dimensional (corresponding to the dimension of the fusion feature matrix), the number of hidden layer nodes is set to 128, the activation function is Sigmoid, and a regularization term is introduced to optimize the model. The regularization parameter is set to C=100. The multimodal fusion feature matrix is used as input, and the fault label matrix (one-hot encoded, including four categories: insulation aging, partial discharge, core fault, and normal operation) is used as output. The Moore-Penrose generalized inverse of the hidden layer output matrix is solved, and the output weight is calculated according to formula (5) to complete the model training and verification tuning. (5); In the formula, It is the output weight matrix from the hidden layer to the output layer of the Extreme Learning Machine model; This represents the output matrix of the hidden layer in the Extreme Learning Machine model. Output matrix for hidden layer The transpose of the matrix; This is the model regularization parameter, used to balance prediction error and weight norm; For the identity matrix, the dimensions are... Consistent; The fault label matrix uses one-hot encoding (e.g., "1, 0, 0" corresponds to insulation aging fault).
[0082] Step 5.2: Fault Early Warning and Dynamic Updates The multimodal fusion feature matrix, collected and processed in real time, is input into the trained early warning model. The model outputs the probability of a current transformer fault and the specific fault type. In this embodiment, when fusion feature data for a certain period is input, the model outputs a fault probability of 0.93 and a fault type of partial discharge, which is consistent with the on-site detection results.
[0083] according to Figure 1 The real-time operating condition monitoring logic acquires key operating condition parameters such as transformer load rate and voltage level in real time through the operating condition monitoring unit. When the monitored change in operating condition parameters exceeds a threshold (set at 15%, such as the load rate increasing from 30% to 48%), the ISSA dynamic weight allocation module is immediately triggered to recalculate the weights of each modal characteristic. The weight change trend is as follows: Figure 5 As shown, the fusion feature matrix and extreme learning machine model parameters are updated synchronously to enable the model to adapt to changes in working conditions in real time and continuously output accurate early warning results.
[0084] In this embodiment, the method was applied to fault early warning of a 110kV main transformer. Compared with the traditional single oil chromatography analysis early warning method, the fault identification accuracy was improved to 96.8%, the false alarm rate was reduced to 2.1%, and the missed alarm rate was reduced to 1.5%. For fault early warning under complex operating conditions, combined with… Figure 5 The dynamic weight adjustment mechanism significantly improves the adaptability and reliability of the method, enabling accurate early warning of transformer faults 2-4 hours in advance. This provides sufficient time for maintenance teams to handle faults, effectively preventing power accidents and reducing transformer maintenance costs, thus fully verifying the practicality and effectiveness of the method.
[0085] In the preferred embodiment, in step 1, the real-time operating data is collected by a sensor array, the environmental data is collected by the substation meteorological station and electromagnetic interference monitoring equipment, and the historical maintenance data is exported from the transformer operation and maintenance system. This multi-source data acquisition comprehensively obtains information related to the transformer's operating status, covering real-time operation, environmental impact, and historical maintenance, providing a rich and accurate data foundation for subsequent analysis. This avoids incomplete analysis due to limited data, thereby enabling a more precise understanding of the transformer's actual condition and providing a reliable basis for fault early warning.
[0086] In the preferred embodiment, step 2 involves the following differentiated preprocessing methods: For real-time operational data, particle filtering is used for denoising, combined with cubic spline interpolation to fill in missing values; for historical maintenance data, a conditional random field algorithm is used to extract knowledge triples, while simultaneously cleaning structured tables; for environmental data, an exponentially weighted moving average method is used to suppress data fluctuations. These settings, tailored to different data characteristics, effectively improve data quality. Denoising and filling in missing values make real-time operational data more accurate, knowledge extraction and cleaning make historical maintenance data more valuable, and suppressing fluctuations makes environmental data more stable, providing high-quality data for subsequent feature extraction and fault analysis, and improving the reliability of the analysis results.
[0087] In the preferred embodiment, step 3 of the multi-branch feature extraction architecture includes a time-series branch, a text branch, and an environment branch. The time-series branch uses an improved convolutional neural network to extract time-series features from real-time operating data. The text branch extracts fault theme features from historical maintenance data based on a topic model optimized by an improved word frequency-inverse document frequency algorithm. The environment branch uses a gradient boosting decision tree to extract the attenuation features of environmental data on transformer equipment performance. This multi-branch architecture deeply mines features from different types of data, comprehensively capturing various information during transformer operation. Time-series features reflect operational dynamics, fault theme features reveal historical problems, and attenuation features reflect environmental impacts. These multi-dimensional features provide strong support for accurate fault early warning, improving the accuracy of early warnings.
[0088] In a preferred embodiment, the improved term frequency-inverse document frequency algorithm introduces a fault-related word weight enhancement factor. The core vocabulary of faults Values range from 1.5 to 2.0 for common vocabulary. The value is set to 1.0. This setting, by enhancing the weight of core fault terms, highlights information closely related to faults in historical maintenance data. When extracting fault-related features, it makes fault-related content easier to identify and extract, avoiding interference from common terms. This allows for a more accurate grasp of historical fault situations, providing a more targeted reference for assessing current transformer fault risks and improving the accuracy of fault warnings.
[0089] In the preferred scheme, in step 4, principal component analysis is first used to reduce the dimensionality of the three types of feature vectors to the same dimension. Then, the temporal feature weights are obtained through iterative optimization of the improved sparrow search algorithm. Text feature weights Environmental feature weights ,and The above settings, including principal component analysis for dimensionality reduction, can reduce data dimensionality and computational complexity while retaining key information. The improved sparrow search algorithm iteratively optimizes weights, rationally allocating weights based on the degree of influence of different features on faults. This allows each feature to play its appropriate role in fault analysis, improving efficiency and accuracy, and enabling more precise assessment of transformer fault risks.
[0090] In a preferred embodiment, the improved sparrow search algorithm uses the Pearson correlation coefficient between each modal feature and the fault label as the fitness function. This setting, using the Pearson correlation coefficient as the fitness function, accurately measures the degree of correlation between each modal feature and the fault. During the iterative process of the improved sparrow search algorithm, optimizing feature weights based on this coefficient makes the weight allocation more scientific and reasonable, allowing features with a greater impact on the fault to receive higher weights. This more accurately reflects the contribution of each feature to the fault, improving the reliability and accuracy of fault early warning.
[0091] In the preferred embodiment, in step 5, the threshold for the change in operating parameters is 15%. The optimized Extreme Learning Machine (ELM) model is a single-hidden-layer feedforward neural network structure, with the Sigmoid function used as the activation function in the hidden layer. Furthermore, a regularization term is introduced to optimize the parameters. With these settings, the 15% threshold can reasonably define the range of operating condition changes and trigger timely model updates. The optimized ELM model has a simple structure and high computational efficiency. The Sigmoid function better fits the data relationships, and the regularization term prevents overfitting, improving the model's generalization ability. This allows the model to accurately adapt to different operating conditions, precisely predict transformer fault probabilities, and ensure the stable operation of the power system.
[0092] In the preferred embodiment, in step 5, the fault probability ranges from 0 to 1, with a value closer to 1 indicating a higher risk of transformer fault. The specific fault types include at least one of insulation aging, partial discharge, and core fault. This setting quantifies the fault probability within the 0-1 range, intuitively reflecting the degree of fault risk and facilitating rapid judgment and decision-making by maintenance personnel. Clearly defining the specific fault types allows for corresponding maintenance measures to be taken for different faults, improving maintenance efficiency, reducing fault losses, ensuring normal transformer operation, and enhancing the reliability and stability of power supply.
[0093] In summary, this invention proposes a transformer fault early warning method based on multimodal data fusion and dynamic weighting, effectively solving the key problems of poor generalization and insufficient weight adaptability in the field of transformer fault early warning technology. Traditional methods are often limited by a single data source, failing to comprehensively capture the complex operating states of transformers, and using fixed weights for data fusion makes it difficult to adapt to dynamic changes under different operating conditions, resulting in limited accuracy and timeliness of fault early warning. To address these technical shortcomings, this invention overcomes the limitations of poor multimodal data fusion and insufficient adaptability to operating conditions caused by fixed weights in existing technologies by comprehensively utilizing real-time transformer operating data, environmental data, and historical maintenance data.
[0094] To achieve the above objectives, this invention employs a series of technical solutions. First, multi-source data on transformers is collected through multiple channels, including sensor arrays, weather stations, and electromagnetic interference monitoring equipment. This data is then subjected to refined preprocessing, such as particle filtering for denoising, conditional random field entity recognition, and exponentially weighted moving average. For real-time operational data, particle filtering and cubic spline interpolation are used; for text data, conditional random field entity recognition is employed; and for environmental data, exponentially weighted moving average is used. This differentiated preprocessing effectively addresses issues such as noise and missing data in time-series data, unstructured text data, and fluctuations in environmental data, ensuring the integrity, structure, and effectiveness of the data. This lays a solid foundation for subsequent accurate feature extraction and fault early warning.
[0095] Subsequently, a multi-branch feature extraction architecture was constructed, utilizing an improved convolutional neural network (ICNN), a topic model based on word frequency-inverse document frequency, and gradient boosting decision tree (GBDT) to extract temporal features, textual features, and environmental features, respectively. Appropriate algorithms were selected for feature extraction based on the characteristics of different data, enabling the full mining of hidden fault-related features in different types of data, including temporal dependency features, fault topic features, and environmental decay features. This achieved thorough mining and accurate extraction of fault information, significantly improving the accuracy of fault early warning.
[0096] Furthermore, a dynamic weight allocation module based on an improved Sparrow Search Algorithm (ISSA) was designed. This module iteratively calculates the fault contribution of each modal characteristic based on transformer current load rate, voltage level, and other operating parameters, generating optimal weight coefficients to achieve adaptive adjustment of weights according to operating conditions. This dynamic weight allocation mechanism dynamically adjusts the data fusion strategy based on the actual operating conditions of the transformer, solving the problem of insufficient adaptability of traditional fixed weights. This enables the fault early warning model to better adapt to complex and changing operating environments, significantly improving the reliability and adaptability of fault early warning, and promoting the transformation of transformer operation and maintenance from passive handling to proactive early warning.
[0097] Finally, a fault early warning model based on extreme learning machine optimization is constructed. Using a fused feature matrix as input, it predicts the probability and type of faults, and dynamically updates the weights through real-time operating condition monitoring to ensure the timeliness and accuracy of the early warning results. This invention comprehensively reflects the transformer status from multiple dimensions by integrating multi-source heterogeneous data, providing a more comprehensive and richer data foundation for transformer fault early warning. It solves the problem of inaccurate fault early warning caused by the low utilization rate of single-modal information in traditional methods, and provides a completely new approach and method for transformer fault early warning.
Claims
1. A transformer fault early warning method based on multimodal data fusion and dynamic weighting, characterized in that, Includes the following steps: Step 1: Collect real-time operating data, environmental data, and historical maintenance data of the transformer; Step 2: Perform differentiated preprocessing based on the heterogeneity of the three types of data; Step 3: Construct a multi-branch feature extraction architecture to extract fault-related features from the three types of data and obtain corresponding feature vectors; Step 4: Construct a dynamic weight allocation module based on the improved sparrow search algorithm, calculate the fault contribution of each modality feature and generate dynamic weight coefficients, and weight and fuse the three types of feature vectors to obtain a multimodal fusion feature matrix; Step 5: Construct an optimized extreme learning machine fault warning model, using the fusion feature matrix as input and output transformer fault probability and specific fault type. When the change in operating parameters exceeds the threshold, recalculate the weights and update the model, output and visualize the warning results.
2. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 1, characterized in that: In step 1, the real-time operating data is collected by a sensor array, the environmental data is collected by a substation meteorological station and electromagnetic interference monitoring equipment, and the historical maintenance data is exported from the transformer operation and maintenance system.
3. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 1, characterized in that, In step 2, the specific method of differential preprocessing is as follows: for real-time running data, particle filtering is used to denoise and cubic spline interpolation is used to fill missing values; for historical maintenance data, the conditional random field algorithm is used to extract knowledge triples, while cleaning the structured table; for environmental data, the exponential weighted moving average method is used to suppress data fluctuations.
4. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 3, characterized in that, The calculation formula for the exponentially weighted moving average method is as follows: (1); In the formula, for Raw environmental data collected in real time. for Weighted moving average of environmental data over time. These are the weighting coefficients. for The weighted moving average of environmental data at any given time.
5. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 1, characterized in that: In step 3, the multi-branch feature extraction architecture includes a time-series branch, a text branch, and an environment branch. The time-series branch uses an improved convolutional neural network to extract time-series features of real-time operating data. The text branch uses a topic model optimized by an improved word frequency-inverse document frequency algorithm to extract fault topic features of historical maintenance data. The environment branch uses a gradient boosting decision tree to extract the attenuation features of environmental data on transformer equipment performance.
6. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 5, characterized in that, The improved term frequency-inverse document frequency algorithm introduces a fault-related word weight enhancement factor. Improved The formula for calculating the value is: (2); In the formula, For vocabulary In the document Improved value; For vocabulary In the document Word frequency in; For vocabulary Inverse document frequency; For vocabulary Fault-related weight enhancement factor.
7. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 1, characterized in that, In step 4, principal component analysis is first used to reduce the dimensionality of the three types of feature vectors to the same dimension. Then, the temporal feature weights are obtained through iterative optimization of the improved sparrow search algorithm. Text feature weights Environmental feature weights The formula for calculating the multimodal fusion feature matrix is as follows: (3); In the formula, This is a multimodal fusion feature matrix; This is the time-series feature vector of real-time running data after dimensionality reduction by principal component analysis; This represents the fault theme feature vector of historical maintenance data after dimensionality reduction by principal component analysis. This is the environmental data decay feature vector after dimensionality reduction by principal component analysis.
8. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 7, characterized in that, The improved sparrow search algorithm uses the Pearson correlation coefficient between each modal feature and the fault label as the fitness function. The fitness function is calculated as follows: (4); In the formula, Based on feature weights The fitness function value of the variable; It is a multimodal fusion feature matrix With fault labels covariance; Multimodal fusion feature matrix Standard deviation; Fault Label The standard deviation.
9. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 1, characterized in that, In step 5, the optimized extreme learning machine model is a single-hidden-layer feedforward neural network structure. The hidden layer uses the Sigmoid function as the activation function, and the model introduces a regularization term to optimize the parameters. The output weight calculation formula is as follows: (5); In the formula, It is the output weight matrix from the hidden layer to the output layer of the Extreme Learning Machine model; This represents the output matrix of the hidden layer in the Extreme Learning Machine model. Output matrix for hidden layer The transpose of the matrix; These are the model regularization parameters; It is the identity matrix; This is a fault label matrix.
10. The transformer fault early warning method based on multimodal data fusion and dynamic weighting according to claim 1, characterized in that: In step 5, the fault probability ranges from 0 to 1, and the closer the value is to 1, the higher the risk of transformer fault. The specific fault types include at least one of insulation aging, partial discharge, and core fault.
Citation Information
Patent Citations
Transformer fault early warning method and system
CN119312124A
Transformer fault prediction method based on multi-source data fusion and related device
CN120448913A