Intelligent monitoring and analyzing method for operation of power equipment based on artificial intelligence
By acquiring cross-modal, multi-dimensional monitoring data and optimizing artificial intelligence models, the limitations of traditional power equipment monitoring methods have been overcome, enabling intelligent operation and maintenance of power equipment and improving the stability of power systems. In particular, it has achieved high precision and high efficiency in fault identification, prediction, and source tracing.
Patent Information
- Application Number
- CN202511425662.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-09
AI Technical Summary
Traditional power equipment monitoring methods fail to fully cover the full range of equipment characteristics from external appearance to internal operation, resulting in insufficient fault identification, inability to dynamically adapt to the entire equipment lifecycle, lack of a system-wide perspective, low real-time performance and resource utilization efficiency, and inability to meet the need for rapid response to equipment faults.
By acquiring cross-modal, multi-dimensional monitoring data and integrating attention allocation rules with cross-type data transformation, an artificial intelligence monitoring and prediction model is constructed. This model is then dynamically optimized using a digital simulation model to achieve identification of abnormal equipment operation features, fault prediction, and source tracing. The model is then processed collaboratively on an edge computing cluster and a cloud platform.
It achieves deep integration of multiple types of data, improves the dimensions of fault identification, the accuracy of fault prediction and location, ensures the stability of the power system and the intelligent operation and maintenance of equipment, and avoids unplanned shutdowns and system-level power outages.
Smart Images

Figure CN121302008A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system operation and maintenance technology, and in particular to an intelligent monitoring and analysis method for power equipment operation based on artificial intelligence. Background Technology
[0002] As the core carrier of the power system, the operating status of power equipment directly determines the reliability of power supply. With the expansion of the power grid and the increasing aging of equipment, the requirements for intelligent and comprehensive monitoring methods are rising. Traditional power equipment monitoring methods rely on single parameter acquisition and fixed threshold judgment as their core logic. While these methods provided support for basic operation and maintenance in the early stages of power system development, their limitations can be gradually deduced from the technical logic, considering the current complex power grid conditions and equipment characteristics. Traditional methods often focus on numerical parameters such as voltage, current, and temperature, failing to include text records generated during equipment operation and maintenance, images of the equipment's appearance, and vibration and sound characteristics during operation within the monitoring scope. If the equipment has a hidden fault where the value does not exceed the threshold but the appearance has already shown signs of corrosion, relying solely on numerical parameters will inevitably fail to capture such risks. It can be inferred that a single data source leads to information blind spots in monitoring, making it difficult to cover the full state characteristics of the equipment from its external appearance to its internal operation, resulting in insufficient comprehensiveness in fault identification.
[0003] Directly classifying a transformer winding temperature exceeding 80℃ as a fault fails to consider the impact of equipment age and real-time load on the normal parameter range. In other words, the insulation aging degree of a newly commissioned transformer differs from that of a transformer that has been in operation for 10 years, and the normal temperature threshold should naturally differ. The fixed model cannot dynamically adapt to this change. At the same time, for newly connected equipment of the same type, traditional models need to collect a large amount of data for full training, and the deployment cycle is usually as long as 1-2 months. It can be inferred that the static attributes of the model make it poorly adaptable to the entire life cycle of the equipment, resulting in high debugging costs and low efficiency when connecting new equipment.
[0004] When the switchgear contacts overheat, a shutdown for maintenance is triggered directly without considering the switchgear's topological location in the power system. For example, if the switchgear is the main power supply for an industrial user and there is no backup line to transfer the load, a hasty shutdown will cause the user to stop production. It can be inferred that such decisions lack a system-wide perspective, fail to assess the impact of equipment failure on downstream loads and related lines, and are prone to causing a global power outage due to improper local handling, which contradicts the current operation and maintenance goal of prioritizing power grid safety and minimizing power outages.
[0005] Traditional monitoring often employs a centralized model of on-site data collection and cloud-based analysis, requiring all collected data to be uploaded to a cloud platform. For substations in remote areas, network bandwidth is typically limited, and the synchronous transmission of large amounts of data such as vibration and images is prone to latency, potentially missing the optimal window for initial fault handling. Furthermore, the bandwidth resources consumed by redundant data increase communication costs. Therefore, it can be inferred that the centralized architecture suffers from both real-time bottlenecks and resource waste, failing to meet the demand for rapid response to equipment faults. To address this, we propose an intelligent monitoring and analysis method for power equipment operation based on artificial intelligence. Summary of the Invention
[0006] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides an intelligent monitoring and analysis method for power equipment operation based on artificial intelligence. This method enables deep fusion of multiple types of data, dynamic optimization of monitoring models, accurate prediction and tracing of faults, and comprehensive collaborative processing, thereby improving the intelligence level of power equipment operation and maintenance and the stability of power system operation.
[0007] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: An intelligent monitoring and analysis method for power equipment operation based on artificial intelligence is proposed to acquire cross-modal multi-dimensional monitoring data of power equipment. The cross-modal multi-dimensional monitoring data includes equipment operating parameters recorded in numerical form, equipment maintenance and operation records recorded in text form, equipment appearance status data presented in image form, and vibration waveform data and sound feature data generated during equipment operation. Artificial intelligence processing methods that integrate attention allocation rules and cross-type data transformation are used to perform preliminary cleaning, supplementation and deep integration at the feature level on the cross-modal multi-dimensional monitoring data, generating a standardized fusion feature set; Based on the standardized fusion feature set, past equipment failure case data, and digital simulation model constructed according to the actual specifications and parameters of the equipment, an artificial intelligence monitoring and prediction model is built, and the model is dynamically adjusted and optimized through multi-channel information feedback rules. The cross-modal fusion data collected and processed in real time is input into the optimized artificial intelligence monitoring and prediction model. Combined with the simulation operation function of the digital simulation model, it completes the identification of abnormal equipment operation characteristics, the prediction of fault development trend, and the tracking of the root cause of fault, and outputs abnormal assessment values, fault risk level and fault cause tracking map. Based on the overall impact of the equipment on the entire power system, and combined with anomaly assessment values and fault risk levels, a collaborative processing and maintenance plan is formulated to ensure both the safe operation of individual equipment and the overall stability of the power system.
[0008] Preferably, the feature-level deep integration step of the cross-modal multi-dimensional monitoring data includes: Construct a cross-modal data association map, synchronize data from different sources through a unified time stamp, achieve spatial matching of data through the positional correspondence of various components of the equipment, and establish spatiotemporal correlations between numerical parameters, text records, image features, vibration and sound data. By employing data type conversion processing methods, text-based maintenance and operation records are converted into semantic feature vectors that can be used for analysis, image-based appearance data is converted into visual feature vectors, and vibration waveform and sound feature data are converted into feature vectors arranged in time order. A dynamic attention allocation rule is introduced, which dynamically adjusts the weight of different types of data in the integration process based on the real-time power load of the equipment and the correlation between the equipment's past failures and various types of data. Specifically, the weight of vibration and sound data of the core components of the equipment is increased by 20% to 50%, while the weight of environmental data related to non-core areas of the equipment is reduced by 10% to 30%. By concatenating different types of feature vectors and focusing attention, a standardized fusion feature set containing multifaceted information is generated. This solves the problem of incomplete information caused by relying on only a single type of data for analysis and reduces the error of the analysis results.
[0009] Preferably, the steps for building and dynamically optimizing the artificial intelligence monitoring and prediction model include: Based on the equipment's 3D design drawings, manufacturing material parameters, and past operating data, a digital simulation model with a 1:1 scale to the actual equipment size is constructed to achieve real-time visualization of the equipment's actual operating status. By inputting a standardized set of integrated features into a digital simulation model, the changing process of equipment operating status under different fault causes is simulated, generating virtual fault sample data to supplement the lack of actual fault case data. Establish closed-loop optimization rules for model prediction results, simulation model operation results, and actual equipment operation: when the deviation between the model prediction results and the simulation model operation results exceeds 5%, the parameters of the simulation model are automatically corrected by calling the actual equipment operation data; when the deviation between the simulation model operation results and the actual equipment operation on site exceeds 3%, the incremental training process of the model is started. By adopting a learning approach that applies the knowledge of the trained model to the monitoring scenario of new equipment, combined with a learning method that quickly adapts to new scenarios, for newly connected power equipment of the same type, only 10% to 20% of the local operating data of the equipment is needed to complete the model adaptation and adjustment. Compared with the traditional transfer learning method, the training efficiency is improved by 3 to 5 times.
[0010] Preferably, the steps for predicting the fault development trend and tracing the root cause of the fault include: Based on the fault development time series data generated by the digital simulation model, a multi-stage fault prediction model was built, which divides the fault development process into four stages: the initiation stage, the development stage, the critical stage, and the fault occurrence stage, and clearly marks the characteristic numerical limits corresponding to each stage. By combining the rate of change of real-time fusion features and the current power load of the equipment, and by using prediction methods that analyze time series data, the remaining time required for a fault to develop from the current stage to the critical stage can be accurately predicted with an error controlled within 1 hour. By employing an improved fault tree analysis method and combining the correlation between various components of the equipment in the digital simulation model, the direct and indirect causes corresponding to abnormal features are tracked, and a visualized cause tracking map containing fault inducing factors, involved components, and scope of impact is generated. The fault location accuracy can reach the level of a single equipment component.
[0011] Preferably, the steps for formulating the collaborative processing and maintenance plan include: A global impact assessment model for the power system is established, taking into account the load proportion of the power supply line where the equipment is located, the type of users served by the line, and the load transfer capacity within the power system, to calculate the impact weight of the equipment on the entire power system. Establish a three-dimensional decision-making standard based on anomaly assessment values, fault risk levels, and global impact weights: When the global impact weight of the equipment is ≥0.6 or the fault risk level is severe, the emergency load transfer and priority repair plan is activated, and the dispatching instructions are pushed to the power grid control center. When the global impact weight of the equipment is between 0.3 and 0.6 or the fault risk level is moderate, formulate a staggered maintenance and spare parts preparation plan, and determine the maintenance time in combination with the period when the power grid load is at its lowest. When the global impact weight of the equipment is less than 0.3 or the fault risk level is minor, a regular tracking and monitoring and automatic early warning plan can be formulated without stopping the equipment to intervene. The formulated processing plan is simultaneously sent to the equipment operation and maintenance management platform, the power grid dispatching system, and the spare parts management system to achieve coordinated response from multiple systems.
[0012] Preferably, it also includes a collaborative processing step between the edge computing cluster and the cloud platform: Edge computing clusters are deployed in power facilities such as substations and power distribution rooms. A main edge computing node coordinates multiple sub-edge computing nodes. The sub-edge computing nodes are responsible for real-time data collection and preliminary anomaly screening of one or more devices, with a response time controlled within 100 milliseconds. When a child edge computing node detects a complex anomaly, it uses a dedicated communication channel based on the fifth-generation mobile communication network to achieve low-latency data transmission between edge nodes and shares the abnormal data with the main edge computing node. The main edge computing node then calls upon the cluster's computing power for collaborative analysis. If the problem still cannot be resolved, the data is uploaded to the cloud platform. The cloud platform builds a model training resource pool, generates dedicated predictive and diagnostic models based on the device types and operating scenarios monitored by the edge computing cluster, and distributes some modules of the model to the edge computing nodes through incremental updates, realizing a closed-loop process of cloud-customized models, edge node execution analysis, and result feedback optimization.
[0013] Preferably, the steps for acquiring the cross-modal multi-dimensional monitoring data include: Operating parameters such as voltage, current, and temperature are collected and recorded in numerical form by numerical sensors deployed on and around the equipment. Capture images of the device's appearance using a high-definition camera to obtain data on the device's appearance in image form; The sound characteristics of the equipment during operation are recorded by the sound feature acquisition device, and the vibration waveform data of the equipment during operation is collected by the vibration sensor; The operation and maintenance terminal records the daily maintenance and troubleshooting of the equipment in written form. Add uniform time stamps and device component location identifiers to the collected data to complete the initial data processing.
[0014] Preferably, the incremental training step of the model includes: Regularly collect the latest operating data and newly added fault cases from the equipment, and process them into standardized fusion feature data; The processed new data is used as an incremental training dataset and merged with the original training data; Keep the network structure of the model unchanged, adjust the parameters inside the model, and adapt the model to the characteristics of the new data through repeated iterative calculations; After training is completed, compare the accuracy of the model in identifying fault features before and after training. If the accuracy improves by 5% or more, save the optimized model; otherwise, readjust the training parameters and train again.
[0015] Preferably, the management steps of the edge computing cluster include: Real-time monitoring of the load, network connectivity status, and data processing capabilities of each edge node; The master edge node is dynamically elected based on monitoring data, and nodes with sufficient computing power and stable network connection are given priority as master nodes. When the primary edge node fails, the backup node election mechanism is automatically triggered, and the primary node switch is completed within 10 seconds. The main edge node periodically reports the cluster's operating status and anomaly handling results to the cloud platform, and receives model update instructions issued by the cloud platform.
[0016] Preferably, the step of generating the visualized cause-tracing map includes: Based on the results of tracing the root cause of the fault, determine the cause of the fault, the equipment components involved, and the scope of power supply affected; The tree structure displays the relationship between the causes of failure and equipment components, while the hierarchical structure displays the scope of the impact. Use different colors to indicate the severity of the fault, and use arrows to indicate the path of the fault propagation; The generated maps are converted into an image format that can be displayed on the operation and maintenance terminal and the scheduling platform, and then pushed to the relevant management terminal simultaneously.
[0017] (III) Beneficial Effects 1. By implementing a cross-modal, multi-dimensional data acquisition solution, integrating numerical operating parameters, text-based maintenance records, image-based appearance data, and vibration / sound characteristic data, the limitations of traditional single-numerical parameter acquisition are overcome. Simultaneously, by constructing a spatiotemporal correlation map, time synchronization and spatial matching of multi-source data are achieved. Data weights are dynamically allocated based on equipment load and fault correlation, with the weight of vibration / sound data from core components increased by 20%-50%, while the weight of non-core environmental data is reduced by 10%-30%, effectively highlighting key characteristics. This solves the information fragmentation problem caused by the heterogeneity of multiple data types, enabling fault... The identification dimension has been upgraded from single-parameter judgment to multi-modal feature collaborative analysis. Secondly, by constructing a 1:1 digital simulation model, the evolution of the equipment's operating state under different fault causes is simulated, generating virtual fault samples to supplement the lack of real cases, thus solving the problem of poor model generalization ability caused by the lack of rare fault samples. At the same time, an adaptation scheme combining transfer learning and meta-learning is implemented to transfer the knowledge of the trained model to new equipment scenarios, requiring only 10%-20% of local data to complete the adaptation. In addition, incremental training is automatically started when the deviation between simulation and reality exceeds 3%, ensuring that the model continuously adapts to changes in equipment state.
[0018] 2. Implement a multi-stage fault prediction scheme, dividing fault development into four stages: initiation, development, criticality, and failure, and defining characteristic thresholds. Combine the change rate of real-time fusion features and equipment load, and predict the remaining time for the fault to develop to the critical stage through time series analysis. At the same time, based on the component correlation in the digital simulation model, use an improved fault tree analysis method to trace direct and indirect causes and generate a visual source map. The implementation of this technical solution shifts fault prediction from post-event identification to pre-event warning, controlling prediction errors to within one hour and achieving fault location accuracy at the individual component level. This means that maintenance personnel can formulate maintenance plans in advance based on the predicted time and accurately replace faulty components based on the source map, effectively avoiding unplanned downtime. Secondly, by building a global impact assessment model, the global impact weight is calculated by comprehensively considering the load ratio of the line where the equipment is located, the type of service users, and the system load transfer capability. A three-dimensional decision matrix is established, consisting of anomaly assessment value, fault risk level, and global impact weight: for serious faults in high-load equipment on the main line, emergency load transfer is initiated; for minor anomalies in equipment on branch lines, only tracking and warning are initiated. At the same time, the solution is synchronized to multiple systems such as maintenance, scheduling, and spare parts to achieve coordinated response. The implementation of this decision-making solution avoids system-level power outages caused by traditional one-size-fits-all shutdown maintenance, and significantly improves the power supply reliability, especially for industrial users and important backup loads.
[0019] 3. Deploy edge computing clusters in substations and power distribution rooms. Sub-edge nodes are responsible for local real-time data collection and preliminary anomaly screening. Only complex anomaly data is uploaded to the main edge node or the cloud. The cloud generates device-specific models and distributes them incrementally to the edge. At the same time, the continuity of edge processing is ensured through dynamic election of the main edge node and a backup node switching mechanism within 10 seconds. Attached Figure Description
[0020] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0021] Figure 1 This is an overall flowchart of an embodiment of the present invention. Detailed Implementation
[0022] This application provides an AI-based intelligent monitoring and analysis method for power equipment operation, achieving deep fusion of multiple data types, dynamic optimization of monitoring models, accurate fault prediction and tracing, and comprehensive collaborative processing. This improves the intelligence level of power equipment operation and maintenance and the stability of the power system. By implementing a cross-modal, multi-dimensional data acquisition scheme, it integrates numerical operating parameters, text-based maintenance records, image-based appearance data, and vibration / sound characteristic data, overcoming the limitations of traditional single-parameter data acquisition. Simultaneously, it constructs a spatiotemporal correlation map to achieve temporal synchronization and spatial matching of multi-source data, and dynamically allocates data weights based on equipment load and fault correlation, increasing the weight of vibration / sound data of core components by 20%-50%, and so on. The weight of core environmental data is reduced by 10%-30%, effectively highlighting key features; it solves the information fragmentation problem caused by the heterogeneity of multiple types of data, upgrading the fault identification dimension from single parameter judgment to multi-modal feature collaborative analysis; secondly, by constructing a 1:1 digital simulation model, the evolution of the equipment's operating state under different fault causes is simulated, generating virtual fault samples to supplement the lack of real cases, solving the problem of poor model generalization ability caused by the lack of rare fault samples; at the same time, an adaptation scheme combining transfer learning and meta-learning is implemented to transfer the knowledge of the trained model to new equipment scenarios, requiring only 10%-20% of local data to complete the adaptation; in addition, incremental training is automatically started when the deviation between simulation and reality exceeds 3%, ensuring that the model continuously adapts to changes in equipment state.
[0023] Example: The technical solution in this application aims to achieve deep fusion of multiple types of data, dynamic optimization of monitoring models, accurate prediction and tracing of faults, and comprehensive collaborative processing, thereby improving the intelligence level of power equipment operation and maintenance and the stability of power system operation. The overall approach is as follows: As the core carrier of the power system, the operating status of power equipment directly determines the reliability of power supply. With the expansion of the power grid and the increase in the proportion of new energy grid connection, the limitations of traditional monitoring methods have gradually become apparent: ① Data acquisition only focuses on single numerical parameters such as voltage and current, resulting in a failure rate of over 25% for hidden faults; ② The analysis model uses fixed parameters, requiring 45-60 days to adapt to new equipment, which cannot cope with parameter drift caused by equipment aging; ③ Fault handling only targets single equipment and does not consider the system impact, making it easy for main line equipment failures to cause regional power outages; ④ All data is uploaded to the cloud, and insufficient bandwidth in remote substations leads to a data packet loss rate of over 15%.
[0024] In recent years, artificial intelligence and digital twin technologies have made progress in the application of power operation and maintenance, but existing methods still have shortcomings: multi-source data fusion uses fixed weights, which cannot adapt to load changes; simulation models are disconnected from measured data and lack closed-loop optimization; decision-making lacks quantitative basis and relies on human experience.
[0025] To address the problems existing in the prior art, this invention provides an intelligent monitoring and analysis method for the operation of power equipment based on artificial intelligence. This analysis method includes: 1. Architecture: (1) Cross-modal multi-dimensional data fusion: Multi-dimensional data fusion is used to address the information fragmentation caused by the heterogeneity of various data types, including numerical, text, image, and vibration / sound data. It achieves deep fusion through four steps: data acquisition, modality transformation, dynamic weighting, and feature integration. The core of this process is the dynamic weight quantization formula. Details are as follows: First, data acquisition was deployed. Numerical parameter acquisition used voltage, current, and temperature sensors: a CYL-110 voltage sensor (0-120kV), an ACL-2000 current sensor (0-2000A), and a PT100 temperature sensor (-20-150℃). All three sensors sampled at 1 time per second, and the data was stored in JSON format. Text records were entered using an industrial tablet, encoded in UTF-8. Each record was associated with a timestamp accurate to the second and a component identifier, such as the transformer core A phase or switchgear contacts. Image data was acquired using three high-definition cameras (DS-2CD3T86FWDV2-I5S, 25fps), aimed at the transformer core, tank, and terminals. For switchgear, the camera is aimed at the contacts and circuit breakers; for transmission lines, it is aimed at the towers and insulators. The shooting frequency is 1 time per minute, the image format is JPEG, and the resolution is 1920×1080. In high-dust scenarios, such as substations near cement plants, the camera is equipped with a dust cover and the shooting frequency is increased to 1 time per 30 seconds to ensure image clarity. Vibration / sound data acquisition uses vibration sensors and acoustic fingerprint collectors. The vibration sensor model is YD-100, with a measurement range of 0-500Hz and a sensitivity of 100mV / g. It is deployed on the top of the transformer core, the side of the switchgear cabinet, and the bottom of the transmission line towers. The acoustic fingerprint collector model is ASR-200, with a sampling rate of 44.1kHz. It is deployed near the cooling system and the switchgear operating mechanism. Both have a sampling frequency of 10 times per second, and the data formats are CSV (vibration waveform) and WAV (sound).
[0026] The data preprocessing stage uses timestamps as a basis to achieve spatiotemporal alignment, associating four types of data from the same component at the same time to form structured data tuples. For example, at 08:30:00 on 2025-09-06, the associated data for the switch cabinet contacts are 10kV voltage, 800A current, 58℃ temperature, text record of slight contact heating, contact image, 180Hz vibration waveform, and 65dB sound signature.
[0027] The modality transformation process employs various tools to process different types of data: text data is converted into 128-dimensional semantic vectors using a pre-trained BERT model with weights based on bert-base-chinese. BERT is a pre-trained language model based on a Transformer encoder, capturing text semantics through bidirectional context learning. Its core structure is a 12-layer Transformer encoder, each layer containing 12 self-attention heads, learning general language representations through masked language modeling (MLM) and next-sentence prediction (NSP) pre-training tasks. The pre-trained weights based on bert-base-chinese... This refers to optimization for Chinese scenarios. The input text length is limited to 128 characters to adapt to the length of maintenance records. For example, if a minor abnormal noise was found in the iron core during an inspection on September 1, 2025, the output is a 128-dimensional semantic vector, corresponding to the hidden state of the [CLS] token in the last layer of the model. The batch size is 32, the learning rate is 2e-5, and the number of fine-tuning iterations is 50. Only the output layer is adjusted to adapt to the fault text classification. Traditional text processing methods, such as TF-IDF, only count word frequencies and cannot distinguish the semantic differences between minor and severe abnormal noises. BERT can capture these degree differences through a bidirectional attention mechanism, making it more suitable for the refined semantic representation of maintenance records.
[0028] Image data is processed using a ResNet50 model with pre-trained ImageNet weights to extract 256-dimensional visual vectors. For high-dust scenes, an additional image denoising algorithm, bilateral filtering, is added. ResNet50 is a deep neural network with 50 convolutional layers. Its core innovation is residual connections, which solve the gradient vanishing problem in deep networks by skipping connections across layers, effectively extracting local image details such as equipment corrosion and contact oxidation. The model uses pre-trained ImageNet weights with 1000 image classes. The input image size is 224×224, and 1920×1080 images are resized to this size. Convolutional layers have a stride of 1-2, with shallow layers using a stride of 1 to preserve details and deep layers using a stride of 2 to compress the image size. Max pooling is used, outputting a 2048-dimensional image feature vector, which is then reduced to 256 dimensions by a fully connected layer to adapt and fuse the feature dimensions. For high-dust scenes, a bilateral filtering algorithm is additionally applied, with spatial domain sigma... s =50, grayscale sigma r =100, using spatial proximity and gray-level similarity weighted filtering; traditional models such as VGG16 have no residual connections, and their accuracy drops by 12% after 50 layers; ResNet50 maintains high accuracy at 50 layers through residual connections, making it suitable for feature extraction from power equipment images.
[0029] Vibration / sound data is converted into a 384-dimensional time-series vector using a Short-Time Fourier Transform (STFT). The STFT segments the time-domain signal, applies a window, and then performs a Fourier Transform. By balancing time and frequency resolution through a sliding window, it can convert one-dimensional time-domain vibration / sound data into a two-dimensional frequency-domain feature map, such as converting a vibration waveform into a waveform representing frequency, time, and amplitude. The window function used is the Hanning window, with a window length of 256, a window overlap of 50%, a step size of 128, and 512 Fourier transform points. The frequency resolution is equal to the number of samples. Rate / Number of Points = 44100Hz / 512≈86Hz, which can identify fault frequencies of 50-500Hz; vibration data sampling rate is 10Hz, corresponding to low-frequency vibration of the equipment, and sound data sampling rate is 44.1kHz, corresponding to high-frequency abnormal noise. Finally, the frequency domain feature map is flattened into a 384-dimensional time-series vector; Fourier transform (FT) has no time resolution and cannot capture the change of frequency over time, such as the increase in vibration frequency caused by the development of faults; STFT can dynamically track frequency changes through a sliding window and is the core processing method for vibration / sound data.
[0030] Dynamic weighting and fusion stage based on formula W m =L×C m ×W 0m Calculate the weights, where W m is the final weight for the m-th data category, where m = 1-4 corresponds to the four data categories: numerical, text, image, and vibration / sound, with a value range of [0,1]; L is the real-time load factor of the equipment, calculated as the ratio of actual power to rated power, with a value range of [0.3,1.0]. L = 0.3 under light load and L = 1.0 under full load; C m The correlation coefficient between the m-th type of data and the fault is calculated as the ratio of the number of fault identifications containing the m-th type of data to the total number of fault identifications, with a value range of [0.1, 0.9]. The calculation requires a sample size of ≥100 sets of historical fault data to ensure statistical significance. The statistical period is the past 3 years. Extreme fault data, such as sudden short circuits, must be excluded. For example, the C4 of vibration data and core faults is 0.85, based on statistics from 120 sets of core fault data, of which 102 sets were identified through vibration data. 0m The baseline weight for the m-th type of data is the vibration / sound data W corresponding to the core component. 04 =0.8, Environmental data W for transformer core, switchgear contacts, and non-core areas. 01 =0.2, equipment casing, value range [0.2, 0.8]; taking the switchgear contact vibration data under full load as an example, L=1.0, C4=0.82, based on the statistics of 115 sets of switchgear fault data, W 04 =0.7, the baseline weight of the switch cabinet contact vibration data, substituting into the formula, we get W4=1.0×0.82×0.7=0.574; The weight design must satisfy the following conditions: the higher the load, the more important the core data; the stronger the fault correlation, the greater the data contribution. Therefore, a load coefficient L × correlation coefficient C is constructed. m ×Benchmark Weight W 0m The product model; the load factor L is based on power conservation, the equipment load rate is positively correlated with the failure risk, and the load rate is the ratio of actual power to rated power, therefore L=P 实际 / P 额定 The value ranges from [0.3, 1.0]. When the load rate is <30%, the failure risk is low (L=0.3); when the load rate is >100%, L=1.0; the correlation coefficient C... m Based on statistical probability, C m = Number of faults identified including the m-th class of data / Total number of faults identified. The sample size must be ≥100 groups, the statistical period 3 years, and extreme faults, such as sudden short circuits, need to be excluded to avoid outlier interference; the baseline weight W. 0m It is based on the equipment structure and core components, such as transformer cores and switchgear contacts. 0m =0.7-0.8, non-core data, such as ambient temperature, W 0m =0.1-0.2, adjusted after verification using 100 sets of fault data; taking switchgear contact vibration data as an example: given P 实际 =640A, P 额定 =800A, L=0.8; 84 out of 115 switchgear faults were identified through vibration data, C4=84 / 115≈0.73; contact vibration data W 04 =0.7, then W4=0.8×0.73×0.7≈0.409.
[0031] Feature fusion through formula To achieve this, a 768-dimensional fusion feature vector is constructed by concatenating a 128-dimensional semantic vector W2=0.12, a 256-dimensional visual vector W3=0.18, a 384-dimensional temporal vector W4=0.574, and a 64-dimensional numerical vector W1=0.126. This fusion is then stored in the local database of the edge node. The fusion process must reflect the weight differences of each modality. A weighted summation model is used, linearly superimposing feature vectors of different dimensions—128-dimensional text, 256-dimensional image, 384-dimensional vibration / sound, and 64-dimensional numerical—according to their weights. This ensures that high-weight features, such as vibration, have a higher proportion in the fusion vector. To avoid dimensional inconsistencies, all feature vectors are first normalized to the [0,1] interval, from Min to Max: F m '=(F m -F min ) / F max -F minThen multiply by the weights and sum; after normalization, the text vector F2' is 128-dimensional, W2=0.12, the image vector F3' is 256-dimensional, W3=0.18, the vibration vector F4' is 384-dimensional, W4=0.409, and the numerical vector F1' is 64-dimensional, W1=0.291. Then F... 融合 =0.12F2'+0.18F3'+0.409F4'+0.291F1', finally generating a 768-dimensional fused feature vector.
[0032] The core of this solution lies in the dual-dimensional weighting of load adaptation and fault correlation. Traditional methods with fixed weights cannot adapt to load changes, while L and C... m The weights are dynamically adjusted to ensure that core features are not diluted. The BERT model is chosen for text processing because it can capture semantic differences; ResNet50 is selected for image processing because it can extract local features; and Short-Time Fourier Transform is used for vibration processing because it can convert time-domain waveforms into frequency-domain features, facilitating the identification of fault characteristic frequencies. Applicable scenarios cover power equipment such as transformers, switchgear, GIS equipment, and transmission lines. When expanding to transmission lines, image acquisition is changed to drone inspection images (DJI M300RTK, equipped with a 100-megapixel camera, sampling frequency 1 time / 5 minutes). When expanding to distribution transformers, the sensor parameter range is adjusted, with the voltage sensor range being 0-35kV. Regarding parameter adjustment, the baseline weight W... 0m Adaptation is required based on equipment differences; switchgear contact vibration data W 04 Adjusted to 0.7, image data W of substations with high dust levels 03 Decreased to 0.1; correlation coefficient C m Automatic updates are made quarterly based on newly added fault data, using the SQL statement SELECT COUNT(*) FROM fault_data WHERE data_type='vibration' AND fault_type='contact fault'. The sampling frequency is increased to 2 times / second for numerical data and 20 times / second for vibration data during peak load periods, and reduced to 1 time / 5 minutes for image data during off-peak load periods. In high dust scenarios, the image sampling frequency is increased to 1 time / 30 seconds.
[0033] (2) Model building and optimization driven by digital simulation: To address the issues of insufficient fault samples and model drift (specifically, fewer than 50 rare fault samples) and decreased accuracy after one year of operation, this study utilizes a digital twin model to generate virtual samples, LSTM and Attention models to predict faults, and a closed-loop optimization of the error correction formula. The core approach involves a time-series model capturing fault trends, a simulation model supplementing samples, and an error formula ensuring accuracy. This, combined with the error correction formula, achieves iterative optimization from model to simulation and then to actual measurement. The digital twin model is constructed using 3D equipment design drawings (AutoCAD format, 1:1 scale), material parameters, and references to GB / T 6451-2015 "Technical Parameters and Requirements for Oil-Immersed Power Transformers" and GB / T 3906-2020 3.6kV~40.5kV AC Metal-Enclosed Switchgear and Controlgear," as well as five years of historical operating data (voltage, current, and temperature time-series data from 2020-2024). A 1:1 virtual model is built using the Unity3D 2022.3 engine. The core model for fault prediction is the LSTM+Attention model. The LSTM layer, a Long Short-Term Memory network, addresses the long-term dependency problem of traditional RNNs through input gates, forget gates, output gates, and cell states: the input gate controls the entry of new information into the cell state, the forget gate determines the discarding of old information from the cell state, and the output gate controls the contribution of the cell state to the output. It is suitable for time-series processes such as fault development, where the temperature rises from 50℃ to 75℃ over several hours. The Attention layer assigns weights to the features of each time step in the LSTM output, highlighting key features related to the fault, such as the 250Hz frequency band in vibration data or the area near the 65℃ threshold in temperature data. The weights are calculated as follows: , where a t Let s represent the importance of the LSTM's output features at time step t for fault prediction. t For the similarity between the LSTM output and the query vector, exp(s) tThe similarity score is calculated using the exponential operation, where T is the total number of time steps. The LSTM layer has 256 hidden units, 2 layers, and a dropout of 0.2 to prevent overfitting. The time step T = 10. It takes the fused feature vector from the first 10 time steps as input and predicts the fault state at the 11th time step. The activation function is tanh. The Attention layer has a query vector dimension of 256, consistent with the LSTM hidden layer dimension. Similarity is calculated using the dot product, and the output is a weighted 256-dimensional temporal feature. The output layer uses a soft... The model employs the max activation function for fault type classification (12 categories) and the sigmoid activation function for fault severity regression (1-5 levels). The loss function is a combination of cross-entropy and MSE, the optimizer is Adam, the learning rate is 0.001, and β1=0.9 and β2=0.999. CNN models only process spatial features and cannot capture the temporal evolution of faults. RNN models exhibit gradient vanishing after long sequences, resulting in fault prediction errors exceeding 15%. The combination of LSTM and Attention, through gating and attention weighting, is suitable for predicting the development trend of power equipment faults.
[0034] The core model of a digital twin is a 1:1 virtual mapping of the physical entity of the equipment. It is constructed through a combination of geometric modeling, physical modeling, and behavioral modeling. Geometric modeling is based on AutoCAD 3D drawings and uses Unity3D mesh modeling. The modeling accuracy for key components such as transformer cores and switchgear contacts is 0.1mm, while the accuracy for non-critical components is 1mm. Physical modeling is based on heat transfer using Fourier's law and mechanics using Hooke's law to establish physical equations, such as the transformer core temperature equation T(t) = T0 + (P...). 损耗 ×t) / (C×m), where T(t) is the core temperature at time t, T0 is the initial temperature, and P 损耗 Where t represents power loss, C represents specific heat capacity, and m represents mass; behavioral modeling is based on training a state transition model using historical fault data, such as the mapping relationship between 0.5mm core wear and a vibration frequency of 250Hz; material parameters refer to national standards, with a transformer silicon steel sheet loss coefficient of 1.2W / kg, insulating paper temperature resistance of 105℃, and switchgear contact copper resistivity of 1.7×10⁻⁶. -8 Ω・m; Simulation step size: 1 second / step in normal scenarios, 0.1 seconds / step in fault simulation to capture rapid faults; Accuracy verification: Comparison between actual measurement and simulation, temperature error <3℃, vibration frequency error <5Hz, otherwise adjust physical parameters, such as correcting the silicon steel sheet loss coefficient to 1.22W / kg.
[0035] The core formula for closed-loop optimization, namely the error correction formula and the error calculation formula, is: model error Δ1 = | -Y 仿真 , For example, Y represents the predicted value of a temperature of 65℃, obtained by combining LSTM and Attention. 仿真 For digital twin simulation values, such as a temperature of 68℃; simulation-to-measurement error: Δ2=|Y 仿真 -Y 实际 |,Y 实际 The parameters are: Sensor measured values, such as a temperature of 70℃; Formula design logic: Δ1 threshold 5%, model prediction error exceeding 5% will lead to misjudgment of faults, such as predicting 68℃ instead of 65℃, potentially missing the critical stage, requiring incremental training, iteration count 20, learning rate 0.0005; Δ2 threshold 3%, simulation and measured error exceeding 3% will lead to virtual sample distortion, such as simulated temperature 68℃ vs. measured 70℃, requiring correction of physical parameters, such as correcting the insulation paper's temperature resistance from 105℃ to 103℃; Optimization cycle: executed daily at 3 AM during off-peak load to avoid affecting equipment operation. Specifically includes: Transformer model: Silicon steel sheet loss coefficient 1.2W / kg, insulation paper temperature resistance 105℃, capable of simulating 6 types of faults including core wear, insulation aging, and tap changer failure; Switchgear model: Contact resistance normally <100μΩ, fault condition >500μΩ, capable of simulating 4 types of faults including contact oxidation, circuit breaker jamming, and arc discharge; Transmission line model: Tower material Q235 steel, insulator creepage distance 25mm / kV, capable of simulating 2 types of faults including tower tilt, insulator contamination, and conductor strand breakage; The model needs to achieve two core functions: one is real-time mapping of measured data, such as the model core temperature at a measured temperature of 62℃. The simulation displays 62℃ simultaneously; secondly, it simulates fault conditions. The fault parameters are set with reference to the "Power Equipment Fault Diagnosis Manual". For example, 0.5mm of iron core wear corresponds to the average aging amount of equipment that has been running for 5 years, and 0.2mm of contact oxidation thickness corresponds to the common condition of a switchgear that has not been maintained for 3 years. The simulated output vibration frequency is 250Hz and the temperature rise rate is 3℃ / h. The accuracy of the model needs to be verified by actual measurement, and the error needs to be <3%. For example, when the actual vibration frequency is 248Hz, the simulated value should be around 250Hz, and the error should be controlled within 0.8%. If the error exceeds the standard, the material parameters should be adjusted, such as correcting the silicon steel sheet loss coefficient to 1.22W / kg.
[0036] A hybrid dataset was constructed during the AI model training and virtual sample generation phases. It included 500 real samples (fault data from 2020-2024, covering 12 fault types) and 1000 virtual samples generated by a simulation model, with 80-90 samples per fault type. Parameter settings conformed to actual aging patterns. The 12 fault types are: ① transformer core wear, ② transformer insulation aging, ③ transformer tap changer fault, ④ transformer bushing flashover, ⑤ switchgear contact oxidation, ⑥ switchgear circuit breaker jamming, ⑦ switchgear arc discharge, ⑧ switchgear grounding fault, ⑨ transmission line tower tilting, ⑩ transmission line insulator contamination, ⑪ transmission line conductor strand breakage, and ⑫ general cooling system fault. The feature thresholds for each fault type are shown in Table 1.
[0037] The AI model uses an LSTM (Long Short-Term Memory) network combined with an Attention architecture. It takes a 768-dimensional fused feature vector as input and outputs 12 fault types and fault severity levels 1-5. Training parameters are set to batch size=32, learning rate=0.001, and iterations=100. The loss function is cross-entropy. PyTorch 2.0 is used for training, and the training environment is configured with an NVIDIA A100 GPU. Closed-loop optimization is automatically executed at 3 AM daily, using the formula: Δ1=| -Y 仿真 | and Δ2 = |Y 仿真 -Y 实际 |Calculate the error, where, Y is the predicted value of the AI model. 仿真 Y represents the simulation value of the digital twin model. 实际 The data used are actual measured values from the equipment. If Δ1 > 5%, such as a model predicting a temperature of 65℃ and a simulated temperature of 68℃, then the measured data is used for incremental training, with 20 iterations. If Δ2 > 3%, such as a simulated temperature of 68℃ and a measured temperature of 70℃, then the model material parameters are corrected, such as adjusting the temperature resistance of the insulation paper to 103℃. When adapting to new equipment, the model parameters are initialized based on transfer learning, and the trained LSTM weights are loaded. Only 50 sets of local data are needed for one day of collection, which is convenient for subsequent meta-learning adjustments. The learning rate is 0.0005, the number of iterations is 10, and the adaptation can be completed within 24 hours. For older equipment that has been in operation for more than 5 years, an additional 20 sets of aging data are added, such as the temperature and time curves after insulation aging, for fine-tuning to ensure that the model accuracy is > 90%.
[0038] (3) Failure time prediction and global collaborative decision-making: The management system, which moves from passive alarm to proactive prevention and from localized handling to systemic decision-making through collaborative decision-making, is based on the fault time prediction formula and the global impact weight quantification formula. Fault stage classification is based on historical evolution data of 12 types of faults, clearly dividing fault development into four stages and setting characteristic thresholds: In the nascent stage, transformer temperature is 40-50℃ / switchgear 45-55℃, vibration frequency is 50-100Hz, severity level 1, and the handling requirement is regular monitoring; in the development stage, transformer temperature is 50-65℃ / switchgear 55-70℃, vibration frequency is 100-200Hz, severity level 2, and the handling requirement is early warning; in the critical stage, transformer temperature is 65-75℃ / switchgear 70-85℃, vibration frequency is 200-300Hz, severity level 3, and the handling requirement is preparation for maintenance; in the fault stage, temperature > transformer 75℃ > switchgear 85℃, vibration frequency > 300Hz, severity level 4-5, and the handling requirement is emergency shutdown.
[0039] Failure time prediction is achieved using formula T 剩余 = (T 临界 -T 当前 ) / v T ×k L Calculate, where T 剩余 The remaining time for a fault to develop to the critical stage, in hours; T 临界 The fault critical temperature threshold is preset according to the equipment model and national standards, such as T for transformer core faults. 临界 =75℃, switchgear contact fault T 临界 =85℃; T 当前 The current temperature of the equipment, such as the real-time collected contact temperature of the switchgear, which is 68℃; v T This refers to the rate of temperature change, expressed in °C / hour. It is calculated as the difference between the current temperature and the temperature one hour ago. For example, if the current temperature is 68°C and the temperature one hour ago was 65°C, then v = (68°C / hour). T =3℃ / hour; k L This is the load correction factor, which is positively correlated with the real-time load of the equipment, and is calculated using the formula k. L =0.5+0.5L, where L is the load factor; under light load, L=0.3 corresponds to k L =0.65, L=1.0 at full load corresponds to k L =1.0; Taking a switchgear contact fault as an example, substitute the known condition T 临界 =85℃, T 当前 =68℃、v T =3℃ / hour, L=0.8, k L =0.5 + 0.5 × 0.8 = 0.9, so T is calculated. 剩余 = (85-68) / (3×0.9)≈6.3 hours, and the final output failure will enter the critical stage after 6.2-6.4 hours; Among them, the fault development time is directly proportional to the temperature difference; the larger the temperature difference, the longer the remaining time. It is inversely proportional to the temperature change rate multiplied by the load correction factor; the faster the rate and the higher the load, the shorter the remaining time. 临界 The critical fault temperature is referenced from GB / T standards, such as 75℃ for transformer cores and 85℃ for switchgear contacts. For older equipment (those that have been in operation for more than 5 years), the temperature should be reduced by 5℃, as insulation aging leads to a decrease in temperature resistance. T v is the rate of temperature change. T = (T 当前 -T 前1小时 ) / 1, if v T If the temperature decreases by ≤0, then T 剩余 =∞, fault mitigation; k L k is the load correction factor. L=0.5 + 0.5L, where L is the load rate. The higher the load rate, the worse the heat dissipation. k L The larger T is 剩余 The smaller.
[0040] The global impact weight calculation is based on the substation topology diagram and user data to determine three types of coefficients. The core of the global impact assessment adopts the Analytic Hierarchy Process (AHP). AHP decomposes complex decision problems into a hierarchical structure of objective layer, criterion layer, and alternative layer. A judgment matrix is constructed through expert scoring, criterion weights are calculated, and consistency checks are performed to avoid subjective judgment errors. The objective layer calculates the global impact weight G of the equipment. 全局 The criteria layer consists of the line load proportion coefficient A, the user type coefficient B, and the load transfer capacity coefficient D; the scheme layer consists of core equipment, general equipment, and secondary equipment. Implementation steps: ① Construct a judgment matrix. Invite 5 operation and maintenance experts with more than 10 years of experience to score the importance of the criterion layer using the 1-9 scale, where 1 = equally important and 9 = extremely important. Take the average value to obtain the judgment matrix: Among them, A is 4 times more important than B, A is 7 times more important than D, and B is 3 times more important than D; ② Calculate the weights and solve for the largest eigenvalue λ of M using the eigenvalue method. max ≈3.018, corresponding to the weights [0.7, 0.2, 0.1] obtained after normalizing the feature vector; ③ Consistency test, calculate the consistency index CI = (λ) max -n) / (n-1)=(3.018-3) / (3-1)=0.009, random consistency index RI=0.58, n=3, consistency ratio CR=CI / RI≈0.016<0.1, the test is passed.
[0041] The weights of the three types of coefficients are 0.7, 0.2, and 0.1, respectively, determined using the Analytic Hierarchy Process (AHP). Five power system experts with over 10 years of operation and maintenance experience were invited to score the importance of the coefficients, with a consistency check CR < 0.1. The line load share coefficient A is set based on the load share of the line where the equipment is located: A = 1.0 for main lines with a load share ≥ 60%, A = 0.5 for branch lines with a load share 30%-60%, and A = 0.1 for terminal lines with a load share < 30%. The user type coefficient B is set based on the type of user served: B = 1.0 for critical users (hospitals, data centers), B = 0.8 for industrial users, and B = 0.5 for residential users. The load transfer capacity coefficient D is set based on the availability of backup lines: D = 1.0 for 100% load transfer if backup lines are available, D = 0.5 for partial transfer (50%-100%), and D = 0.1 for no transfer capability. The global impact weight is determined using formula G. 全局=A×0.7+B×0.2+D×0.1; Taking a 10kV switchgear as an example, this switchgear supplies power to a branch line, with a load share of 45%, A=0.5, serves one residential community, B=0.5, has partial transfer capability, D=0.5, substituting into the formula, we get G. 全局 =0.5×0.7+0.5×0.2+0.5×0.1=0.5, which belongs to general equipment, i.e., 0.3≤G 全局 <0.6.
[0042] Collaborative decision-making based on G 全局 Matching handling strategies to fault severity: Core equipment G 全局 For faults ≥0.6, Level 1 faults should be tracked regularly, Level 2 faults should be prepared for off-peak maintenance, Level 3 faults should be prioritized for emergency repair, and Level 4-5 faults should be addressed with emergency shutdown combined with load transfer; for general equipment, 0.3≤G 全局 <0.6, Level 1 faults require no intervention, Level 2 faults require regular tracking, Level 3 faults require off-peak maintenance, and Level 4-5 faults require planned shutdown; secondary equipment G 全局 <0.3, level 1-2 faults require no intervention, level 3 faults require regular monitoring, and level 4-5 faults require off-peak maintenance; in the above switchgear case, G 全局 =0.5, fault severity level 2, matched with a periodic tracking strategy, the system is automatically set to collect data every 30 minutes, and triggers an alarm when the temperature rise rate is >4℃ / hour.
[0043] (4) Edge and cloud collaborative processing: To address the issues of poor real-time performance and bandwidth waste caused by full data uploads, a three-tier architecture combining local filtering, collaborative analysis, and model deployment is adopted, along with a data transmission volume optimization formula to achieve efficient resource utilization. Edge node hardware deployment follows a hierarchical configuration principle: one main edge node and five sub-edge nodes are deployed at a 110kV substation; one main edge node and two sub-edge nodes are deployed at a 10kV distribution room; and independent sub-edge nodes are deployed near transmission line towers, powered by solar energy. Industrial servers are used for the main edge node; embedded terminals are used for the sub-edge nodes; and the transmission lines... Each sub-node is additionally equipped with a 4G / 5G module and a lithium battery; responsibilities are clearly defined: sub-edge nodes are responsible for data collection and simple anomaly screening for a single device, such as temperature > 50℃ but < 65℃, vibration frequency > 100Hz but < 200Hz; the main edge node coordinates sub-node data, performs complex anomaly analysis, such as temperature > 65℃, vibration > 200Hz and sound > 70dB, and communicates with the cloud; transmission line sub-nodes periodically upload data to the main node once every 5 minutes, cache data locally when the network is interrupted, with a maximum cache size of 10GB, and resume transmission after network recovery.
[0044] The dynamic election of the primary edge node adopts a combined algorithm of weighted round-robin and priority sorting. The election is based on three quantitative indicators: ① CPU load rate C with a weight of 0.4 and a threshold of <60%; ② Memory utilization rate M with a weight of 0.3 and a threshold of <50%; ③ Network latency L with a weight of 0.3 and a threshold of <50ms. The election process is as follows: ① Filter candidate nodes that meet the thresholds; ② Calculate the weighted score of the candidate nodes, score = (1-C)×0.4 + (1-M)×0.3 + (1-L / 50)×0.3; ③ Select the node with the highest score as the primary node, and recalculate the score every 10 minutes to ensure that the primary node is always in the optimal state. The fault detection and switching mechanism adopts a combination of heartbeat packets and timeout reconnection: the primary node sends a heartbeat packet to the child node every 10 seconds, and the child node reports its status to the primary node every 10 seconds; if the child node does not receive a heartbeat packet for 3 consecutive times, i.e., timeout 200ms / timeout, the primary node is judged to be faulty, and the primary node is immediately switched to the backup primary node. The node is the candidate node with the second highest score; the backup master node maintains incremental synchronization with the master node, only synchronizing newly added data. During the switchover process, a local cache queue of 1000 records is used to prevent data loss. The data hierarchical processing and transmission process is divided into three scenarios: In simple anomaly handling, after the child node collects data and identifies it as a simple anomaly, such as a temperature of 55℃ or vibration of 150Hz, a warning is generated locally and stored in the local database, without uploading data to the master node; In complex anomaly handling, after the child node identifies multiple features simultaneously as anomalies, it uses 5G slicing with a bandwidth of 10Mbps to upload the data, including a 768-dimensional feature vector, a 10-second vibration waveform, and an image, to the master node. The master node calls the computing power of the three child nodes for collaborative analysis. Based on the Spark distributed computing framework, after confirming the anomaly, it uploads the anomaly data to the cloud and combines it with the analysis results; In unknown anomaly handling, anomaly data that the master node cannot identify is directly uploaded to the cloud for in-depth analysis; The data transmission volume is expressed by the formula Q. 传输 =Q 总 ×(1-α)×β optimization, where Q 传输 Q represents the amount of data transmitted from edge nodes to the cloud, measured in MB / hour. 总 The total amount of data collected for edge nodes, such as vibration data Q. 总 =500MB / hour; α is the local processing rate of edge nodes. For simple anomalies, α=0.8, 80% of the data is processed locally; for complex anomalies, α=0.3, 30% of the data is processed locally. The value range is [0.3, 0.8]. When the proportion of simple anomalies in the switchgear is 60%, α is adjusted to 0.6; β is the data compression coefficient. The LZ4 compression algorithm is used. For text data, β=0.2, compressed to 20% of the original size; for image / vibration data, β=0.3, compressed to 30% of the original size; for 4K drone images, β=0.2, compression ratio 5:1. The value range is [0.2, 0.3]. When the bandwidth of a remote substation is only 2Mbps, β is adjusted to 0.2; Taking the vibration data of the switchgear under complex anomalies as an example, Q 总=400MB / hour, α=0.6, β=0.3, substituting into the formula, we get Q 传输 =400×(1-0.6)×0.3=48MB / hour, which is 88% less than the full upload; Cloud model customization and edge update form a closed loop: The cloud platform uses Alibaba Cloud ECS server, configured with 32-core CPU and NVIDIA V100 GPU, to build a model training resource pool, and train special models for transformers, switch cabinets and transmission lines respectively.
[0045] The cloud generates an incremental model module every week, containing only updated weight parameters, with a size of <100MB, and distributes it to the main edge node via the MQTT protocol. After receiving the module, the main edge node verifies the model accuracy. Once the verification is successful, the module is synchronized to the child nodes, with an update time of <30 minutes. In extreme scenarios, such as high temperature or high dust environments, the cloud generates an additional scenario adaptation module, such as a high temperature environment temperature compensation model, and distributes it to the edge nodes urgently to ensure that the model accuracy drops by no more than 5%.
[0046] The core algorithm for edge node management, the dynamic election algorithm, employs weighted round-robin and priority sorting to select the optimal master node through quantitative indicators, ensuring sufficient computing power for the master node and network stability. The steps are as follows: ①Indicator definition: CPU load rate C, weight 0.4; memory utilization rate M, weight 0.3; network latency L, weight 0.3; ② Candidate screening: Nodes with C < 60%, M < 50%, and L < 50ms are selected as candidate nodes; ③ Score calculation: S=(1-C)×0.4+(1-M)×0.3+(1-L / 50)×0.3, the higher the score, the better the node status; ④ Master node election: Select the candidate node with the highest score as the master node, and recalculate the score every 10 minutes to adapt to changes in node status; ⑤ Failover: The master node sends a heartbeat packet every 10 seconds. If a child node times out three times consecutively (200ms / time), it is considered a failure and the node is switched to the second-highest scoring backup node. The backup node synchronizes with the master node using incremental timestamps, only synchronizing data with timestamps greater than the last synchronization time. The switchover time is less than 10 seconds. Algorithm Example: Candidate node 1, C=50%, M=40%, L=40ms, score S=(0.5×0.4)+(0.6×0.3)+(0.2×0.3)=0.2+0.18+0.06=0.44; Candidate node 2, C=40%, M=30%, L=30ms, score S=(0.6×0.4)+(0.7×0.3)+(0.4×0.3)=0.24+0.21+0.12=0.57; elect node 2 as the master node and node 1 as the backup node.
[0047] Finally, it should be noted that the above embodiments are merely examples for clearly illustrating the present invention and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for intelligent monitoring and analysis of power equipment operation based on artificial intelligence, characterized in that, The analytical method includes: Acquire cross-modal multi-dimensional monitoring data of power equipment. The cross-modal multi-dimensional monitoring data includes equipment operating parameters recorded in numerical form, equipment maintenance and operation records recorded in text form, equipment appearance status data presented in image form, and vibration waveform data and sound characteristic data generated during equipment operation. Artificial intelligence processing methods that integrate attention allocation rules and cross-type data transformation are used to perform preliminary cleaning, supplementation and deep integration at the feature level on cross-modal multi-dimensional monitoring data, generating a standardized fusion feature set; Based on a standardized fusion feature set, past equipment failure case data, and a digital simulation model built according to the actual specifications and parameters of the equipment, an artificial intelligence monitoring and prediction model is built, and the model is dynamically adjusted and optimized through multi-channel information feedback rules. The cross-modal fusion data collected and processed in real time is input into the optimized artificial intelligence monitoring and prediction model. Combined with the simulation operation function of the digital simulation model, it completes the identification of abnormal equipment operation characteristics, the prediction of fault development trend, and the tracking of the root cause of fault, and outputs abnormal assessment values, fault risk level and fault cause tracking map. Based on the overall impact of the equipment on the entire power system, and in conjunction with anomaly assessment values and fault risk levels, a collaborative handling and maintenance plan is developed.
2. The method according to claim 1, characterized in that, The steps for deep integration of cross-modal, multi-dimensional monitoring data at the feature level include: Construct a cross-modal data association map, synchronize data from different sources through a unified time stamp, achieve spatial matching of data through the positional correspondence of various components of the equipment, and establish spatiotemporal correlations between numerical parameters, text records, image features, vibration and sound data. By employing data type conversion processing methods, text-based maintenance and operation records are converted into semantic feature vectors that can be used for analysis, image-based appearance data is converted into visual feature vectors, and vibration waveform and sound feature data are converted into feature vectors arranged in time order. A dynamic attention allocation rule is introduced to dynamically adjust the weight of different types of data in the integration process based on the real-time power load of the equipment and the correlation between the equipment's past failures and various types of data. Specifically, the weight of vibration and sound data of the core components of the equipment is increased, while the weight of environmental data in non-core areas of the equipment is decreased. By concatenating different types of feature vectors and applying attention-focused processing, a standardized fusion feature set containing multifaceted information is generated.
3. The method according to claim 1, characterized in that, The steps for building and dynamically optimizing an artificial intelligence monitoring and prediction model include: Based on the equipment's 3D design drawings, manufacturing material parameters, and past operating data, a digital simulation model with a 1:1 scale to the actual equipment size is constructed. The standardized fusion feature set is input into the digital simulation model to simulate the change process of equipment operating status under different fault causes, and generate virtual fault sample data. Establish optimization rules for model prediction results, simulation model running results, and actual equipment operation: When the deviation between the model prediction results and the simulation model running results exceeds the first threshold, the parameters of the simulation model are automatically corrected by calling the actual equipment operation data; when the deviation between the simulation model running results and the actual equipment operation on site exceeds the second threshold, the incremental training process of the model is started. The learning approach involves applying the knowledge of the trained model to the monitoring scenario of new equipment. For newly connected power equipment of the same type, the model is adapted and adjusted using the local operating data of the corresponding equipment.
4. The method according to claim 1, characterized in that, The steps for predicting the development trend of a fault and tracing its root cause include: Based on the fault development time series data generated by the digital simulation model, a multi-stage fault prediction model was built, which divides the fault development process into four stages: the initiation stage, the development stage, the critical stage, and the fault occurrence stage, and clearly marks the characteristic numerical limits corresponding to each stage. By combining the rate of change of real-time fusion features and the current power load of the equipment, the remaining time required for a fault to develop from the current stage to the critical stage is predicted by analyzing time series data using prediction methods. By employing an improved fault tree analysis method and combining the correlation between various components of the equipment in the digital simulation model, the direct and indirect causes corresponding to abnormal characteristics are traced, generating a visualized cause tracing map that includes fault inducing factors, involved components, and scope of impact.
5. The method according to claim 1, characterized in that, The steps for developing a collaborative processing and maintenance plan include: A global impact assessment model for the power system is established, taking into account the load proportion of the power supply line where the equipment is located, the type of users served by the line, and the load transfer capacity within the power system, to calculate the impact weight of the equipment on the entire power system. Establish a three-dimensional decision-making standard that includes anomaly assessment values, fault risk levels, and global impact weights: When the global impact weight of the equipment is ≥0.6 or the fault risk level is severe, the emergency load transfer and priority repair coordination scheme is activated, and the dispatching instructions are pushed to the power grid control center. When the global impact weight of the equipment is between 0.3 and 0.6 or the fault risk level is moderate, formulate a coordinated plan for off-peak maintenance and spare parts preparation, and determine the maintenance time in combination with the period when the power grid load is at its lowest. When the global impact weight of the equipment is less than 0.3 or the fault risk level is minor, a coordinated plan for regular tracking and monitoring and automatic early warning can be formulated without stopping the equipment to intervene. The formulated handling plan will be simultaneously sent to the equipment operation and maintenance management platform, the power grid dispatching system, and the spare parts management system.
6. The method according to claim 1, characterized in that, It also includes the collaborative processing steps between edge computing clusters and cloud platforms: Deploy edge computing clusters at power facility sites, with one master edge computing node coordinating multiple sub-edge computing nodes. The sub-edge computing nodes are responsible for real-time data collection from one or more devices and preliminary anomaly screening. When a child edge computing node detects multiple anomalies, it transmits low-latency data between edge nodes through a dedicated communication channel based on the fifth-generation mobile communication network. The abnormal data is then shared with the main edge computing node, which uses the cluster's computing power for collaborative analysis. If the problem still cannot be resolved, the data is uploaded to the cloud platform. The cloud platform builds a model training resource pool, generates dedicated prediction and diagnostic models based on the device types and operating scenarios monitored by the edge computing cluster, and distributes some modules of the model to the edge computing nodes through incremental updates.
7. The method according to claim 1, characterized in that, The steps for acquiring cross-modal, multi-dimensional monitoring data include: Operating parameters, including voltage, current, and temperature, are collected by numerical sensors deployed on and around the equipment. High-definition cameras capture images of the device's exterior, obtaining data on the device's appearance in image form; Sound feature acquisition equipment is used to record the sound features during equipment operation, and vibration waveform data during equipment operation is collected by vibration sensors; Use the maintenance terminal to record maintenance records, including routine equipment maintenance and troubleshooting. Add uniform time stamps and device component location identifiers to the collected data to complete the initial data processing.
8. The method according to claim 3, characterized in that, The incremental training steps for the model include: Regularly collect the latest operating data and newly added fault cases from the equipment, and process them into standardized fusion feature data; The processed new data is used as an incremental training dataset and merged with the original training data; Keep the network structure of the model unchanged, adjust the parameters inside the model, and adapt the model to the characteristics of the new data through repeated iterative calculations; After training is completed, the accuracy of the model in identifying fault features before and after training is compared. If the accuracy reaches or exceeds the preset accuracy threshold, the optimized model is saved; otherwise, the training parameters are readjusted and training is performed again.
9. The method according to claim 6, characterized in that, The management steps for an edge computing cluster include: Real-time monitoring of the load, network connectivity status, and data processing capabilities of each edge node; The master edge node is dynamically elected based on monitoring data, and nodes with sufficient computing power and stable network connection are given priority as master nodes. When the primary edge node fails, the standby node election mechanism is automatically triggered to switch the primary node. The main edge node periodically reports the cluster's operating status and anomaly handling results to the cloud platform, and receives model update instructions issued by the cloud platform.
10. The method according to claim 4, characterized in that, The steps for generating a visual cause-of-death mapping diagram include: Based on the results of tracing the root cause of the fault, determine the cause of the fault, the equipment components involved, and the scope of power supply affected; The tree structure displays the relationship between the causes of failure and equipment components, while the hierarchical structure displays the scope of the impact. Use different colors to indicate the severity of the fault, and use arrows to indicate the path of the fault propagation; The generated maps are converted into an image format that can be displayed on the operation and maintenance terminal and the scheduling platform, and then pushed to the relevant management terminal simultaneously.
Citation Information
Cited By
Substation equipment fault early warning method and system based on digital twinning
CN121502492A
AI-enabled data-driven simulation correction and fault prediction method
CN122113635A