Secondary equipment health diagnosis system and method based on multi-source data

By integrating multi-source data for real-time hierarchical acquisition and intelligent feature extraction, and combining LSTM networks with expert rules for collaborative prediction, the reliability and real-time issues of data acquisition in secondary equipment health diagnosis are solved, achieving efficient health status assessment and low-cost equipment monitoring.

CN120993082APending Publication Date: 2025-11-21GUODIAN NANJING AUTOMATION

Patent Information

Application Number
CN202511152591.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies for secondary equipment health diagnosis suffer from problems such as insufficient data acquisition reliability, excessive reliance on manual annotation, poor real-time performance, and high deployment costs. They also struggle to support multi-protocol heterogeneous data access and achieve full-scenario coverage, and lack effective multi-rule data processing methods, resulting in inaccurate diagnostic results and high maintenance costs.

Method used

It adopts a real-time hierarchical acquisition system based on multi-source data, intelligent feature extraction and health status quantitative assessment, combined with LSTM network and expert rule collaborative prediction and correction. Through the real-time hierarchical acquisition system of multi-source data fusion, it supports standardized data access of multiple protocols such as IEC61850, Modbus, and SNMP. It uses wavelet transform to extract transient features of electrical quantities, establishes a historical database module, and adopts a system architecture that does not require the deployment of intelligent detection chips at the device end.

Benefits of technology

It achieves full coverage acquisition of key data such as electrical quantities, communication data, self-test signals, and environmental parameters, improving the accuracy and interpretability of diagnostic results, reducing system deployment complexity and cost, and supporting improvements in real-time performance and data validity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120993082A_ABST
    Figure CN120993082A_ABST
Patent Text Reader

Abstract

The invention discloses a secondary equipment health diagnosis system and method based on multi-source data, and relates to the technical field of secondary equipment monitoring, and the system comprises a data collection module which is used for obtaining multi-source data based on a standard communication protocol, and carrying out the hierarchical collection according to a priority order; the data processing module is used for acquiring the processed standardized data and extracting electrical quantity transient characteristics through wavelet transform; the historical database module is used for establishing a historical data sample storage system and providing multi-source data set samples for model training; the equipment health diagnosis module is used for carrying out periodic prediction by utilizing multi-dimensional equipment feature differentiation fitting and combining a long-short-term memory network model, and correcting to obtain a real equipment health index; and the early warning and decision module is used for analyzing and positioning potential fault elements and generating a maintenance strategy. The method has the advantages of multi-source data real-time grading collection, intelligent feature extraction and health state quantitative evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of secondary equipment monitoring technology, and more specifically, to a secondary equipment health diagnosis system and method based on multi-source data. Background Technology

[0002] With the rapid development of high-speed railway networks and the widespread application of smart substations, the intelligence level of power secondary equipment has been significantly improved, but this has also brought about a substantial increase in system complexity. Traditional cable connection methods are gradually being replaced by digital signals such as GOOSE communication and SV sampling. The lack of visibility of digital signals greatly increases the difficulty of analyzing and handling communication faults and maintaining secondary equipment. During long-term operation, secondary equipment may experience various problems such as component aging, communication anomalies, and logic errors. Timely detection and accurate diagnosis of these problems are crucial for ensuring the safe and stable operation of the power system. Therefore, there is an urgent need for a secondary equipment health diagnosis technology that can achieve intelligent monitoring, accurately assess equipment health status, and provide predictive maintenance decisions.

[0003] To address the issue of secondary equipment health monitoring, existing technologies have proposed various solutions. For example, patent application number 202311378487.2 provides a method for predictive diagnosis of equipment health through multi-source data fusion. This method utilizes a data acquisition terminal to collect multi-source data from multiple power devices, organizes and summarizes the collected multi-source data through a data processing system, and uses a learning system to receive and classify the integrated data information. After obtaining standard data, it is compared and analyzed with a standard database to determine the health status of the power equipment. Another example is patent application number 202311629196.6, which provides a three-layer discrimination method for the health diagnosis of power distribution equipment. This method uses end devices for initial defect identification of power distribution equipment, sends the initial judgment results to edge devices for fine-tuning, and then sends the fine-tuning results to the cloud for final judgment. This three-layer hierarchical discrimination mechanism (end-edge-cloud) enables the identification of power distribution equipment defects and health assessment.

[0004] However, existing technical solutions still have significant shortcomings. First, they lack reliability assurance mechanisms during data acquisition and transmission, making them prone to data loss or transmission errors, thus affecting the accuracy of diagnostic results. Second, existing solutions rely excessively on expert experience and manually labeled standard databases, requiring frequent manual calibration and maintenance, increasing operational costs and making them difficult to adapt to complex and changing real-world operating environments. Furthermore, the edge-cloud three-layer hierarchical discrimination mechanism used in some solutions may lead to cumulative latency issues, especially under conditions of limited network bandwidth or high device load, resulting in poor real-time performance. Additionally, the deployment of intelligent detection chips on each device is costly and difficult for large-scale applications. More importantly, existing technical solutions lack real-time hierarchical acquisition strategies based on multi-source data fusion, making it difficult to support heterogeneous data access from multiple protocols such as IEC61850, Modbus, and SNMP. This hinders the ability to ensure comprehensive data coverage and timely data delivery, and the lack of effective multi-rule data processing methods makes it difficult to fill in missing values, filter outliers, and extract time-frequency domain features. This results in insufficient data sample reliability and excessive reliance on manual labeling, limiting the widespread application of existing technologies in practical engineering.

[0005] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention

[0006] To address the problems in related technologies, this invention proposes a secondary equipment health diagnosis system and method based on multi-source data. It has the advantages of real-time hierarchical acquisition of multi-source data, intelligent feature extraction and quantitative assessment of health status, and collaborative prediction and correction by LSTM network and expert rules. This solves the problems of insufficient data acquisition reliability, excessive reliance on manual annotation, poor real-time performance and high deployment cost in existing technologies.

[0007] Therefore, the specific technical solution adopted by the present invention is as follows:

[0008] According to one aspect of the present invention, a secondary equipment health diagnosis system based on multi-source data is provided, the system comprising:

[0009] The data acquisition module is used to acquire multi-source data based on standard communication protocols and to perform hierarchical acquisition according to priority order;

[0010] The data processing module is used to obtain standardized data after processing from the collected raw data, and to extract transient features of electrical quantities through wavelet transform.

[0011] The historical database module is used to establish a historical data sample storage system based on the processed standardized data, and to provide multi-source dataset samples for model training through the data query function;

[0012] The equipment health diagnosis module is used to perform periodic predictions by using multidimensional equipment feature differential fitting and combining it with a long short-term memory network model to obtain an initial equipment health index within a preset range, and then correct it to obtain the actual equipment health index.

[0013] The early warning and decision-making module is used to analyze and locate potential faulty components based on real equipment health indices, and generate maintenance strategies to be sent to the operation and maintenance personnel's terminals.

[0014] Furthermore, the data acquisition module includes: a multi-source acquisition module, used to acquire multi-source data based on standard communication protocols; specifically, it acquires voltage, current, electrical quantities, and protection setting parameters according to the communication protocol; acquires CPU load and memory status self-test signal data through the self-diagnostic interface built into the Ethernet connection device; and acquires packet loss rate and latency communication data using network packet capture tools or network management protocols; a priority setting module, used to establish a real-time hierarchical priority system based on the impact of different data types on the device's health status, and acquires sensor temperature and humidity environmental parameters through the protocol interface, utilizing the operation and maintenance system. The interface acquires historical equipment operation and maintenance data; the hierarchical data acquisition module is used to acquire high-frequency, medium-frequency, and low-frequency data according to a real-time hierarchical priority system, and dynamically adjusts the data acquisition strategy according to network status; the timing control module is used to sample high-frequency data at the microsecond level, poll medium-frequency data at the second level, and update low-frequency data at the minute or hour level according to data accuracy and timeliness requirements, and execute the acquisition tasks according to priority order; among them, high-frequency data includes electrical quantities; medium-frequency data includes communication data, self-test signal data, and environmental parameters; low-frequency data includes protection parameters and operation and maintenance records.

[0015] Furthermore, the data processing module includes: a data cleaning module, used to clean the collected raw data based on a multi-rule data cleaning strategy, including using linear interpolation to fill in the missing parts of time-series data and deduplicating duplicate data according to timestamps; an anomaly filtering module, used to dynamically filter out outliers within a preset time range based on the three-standard-deviation principle, and remove abnormal data with deviations exceeding the threshold; and a feature extraction module, used to perform time-frequency analysis processing on electrical quantities at the time of the fault based on continuous wavelet transform, extract transient features of electrical quantities, and store the processed standardized data into the historical database module.

[0016] Furthermore, the feature extraction module, when performing time-frequency analysis on electrical quantities at the fault time based on continuous wavelet transform and extracting transient features of electrical quantities, includes: converting an infinitely long trigonometric function basis into a finite-length decaying wavelet basis based on a basis-changing operation to obtain a mother wavelet function; adjusting the mother wavelet function using scaling and translation parameters to establish a wavelet basis function set adapted to different time-frequency scales; performing convolution operations between the wavelet basis function set and the electrical quantity signal at the fault time to obtain multi-scale time-frequency transform coefficients; and performing progressive refinement based on the multi-scale time-frequency transform coefficients to extract transient features of electrical quantities at the fault time.

[0017] Furthermore, based on the multi-scale time-frequency transformation coefficients, progressive refinement is performed to extract the transient electrical quantity features at the fault moment. This includes: progressively refining the electrical quantity signal through scaling and translation operations based on the multi-scale time-frequency transformation coefficients to obtain the time-frequency distribution at different scales; determining the start and end times corresponding to specific frequencies based on the time-frequency distribution, and establishing the correspondence between frequency and time windows; using the correspondence between frequency and time windows, subdividing the signal into high-frequency and low-frequency components, and identifying key frequency bands sensitive to faults; and extracting transient electrical quantity features containing fault moment characteristic information based on the amplitude changes and frequency distribution of the key frequency bands.

[0018] Furthermore, the equipment health diagnosis module includes: a feature optimization module, which uses normalization processing to obtain standardized feature data based on multi-source data samples from different types of secondary equipment, and removes low-contribution features through multi-dimensional equipment feature differentiation fitting to obtain optimized feature data samples; a health assessment module, which uses an LSTM network model to perform periodic predictions based on the optimized feature data samples, and obtains an initial equipment health index through the synergistic effect of multiple gating mechanisms; and a rule correction module, which determines the set of activated influencing factors based on an expert knowledge rule base, and corrects the initial equipment health index according to rules to obtain the true equipment health index.

[0019] Furthermore, when the feature optimization module obtains optimized feature data samples by differential fitting of multi-dimensional equipment features and eliminating low-contribution features, it includes: calculating the contribution weight of each data sample to the health status of secondary equipment using a feature evaluation model based on multi-source heterogeneous data samples acquired by the data acquisition module, and establishing a differential feature weight matrix according to the equipment model; calculating the fitting distance between the current sample data and the feature reference node by constructing historical normal operation data as feature reference nodes; calculating the sum of weighted fitting distances between each sample data and the feature reference node based on the fitting distance and combined with the differential feature weight matrix; retaining the sample data when the sum of weighted fitting distances is less than or equal to a preset threshold; and determining it as a low-contribution feature and removing it from the feature data sample when the sum of weighted fitting distances is greater than the preset threshold; and generating multi-dimensional feature data samples by extracting fault mode feature vectors from self-test data, deterioration trend feature vectors from environmental monitoring data, performance degradation feature vectors from protection action data, and network anomaly feature vectors from communication status data according to the functional type and operating environment of the secondary equipment, and generating multi-dimensional feature data samples through feature fusion algorithms.

[0020] Furthermore, the multi-source heterogeneous data samples include self-test data, environmental monitoring data, protection action data, and communication status data; the fault mode feature vector includes the number of self-test failures, the frequency of specific fault codes, the duration of communication interruption, and the deviation value of the self-test cycle; the degradation trend feature vector includes the mean, maximum and minimum values, and rate of change of temperature and humidity, as well as the temperature distribution characteristics of internal hot spots; the performance degradation feature vector includes the number of protection actions, the dispersion of action time, the frequency of failure to act and false action, the protection setting value, the deviation of actual operating parameters, and the accuracy rate of protection action timing; the network anomaly feature vector includes communication delay, packet loss rate, error frame rate, the mean and variance of communication traffic, and the frequency of burst communication.

[0021] Furthermore, when the rule correction module determines the set of activated influencing factors based on the expert knowledge rule base and corrects the initial equipment health index to obtain the true equipment health index, it includes: establishing a quantitative mapping relationship between fault symptoms and health index decay based on equipment failure modes and fault evolution patterns in the historical fault database; constructing an expert knowledge rule base and setting corresponding influencing factors according to the severity of different fault types; calculating the probability value of triggering various fault rules in the current equipment state by analyzing the similarity between real-time monitoring data and historical fault precursor data based on the expert knowledge rule base, and determining the set of activated influencing factors based on the comparison result of the probability value and the preset probability threshold; fusing the influencing factors in the activated influencing factor set using Bayesian inference method, and calculating a comprehensive correction coefficient by combining the equipment operating years, maintenance records, and environmental factors; and obtaining the true equipment health index by weighting the initial equipment health index with the comprehensive correction coefficient.

[0022] According to another aspect of the present invention, a method for secondary equipment health diagnosis based on multi-source data is also provided, the method comprising:

[0023] Data from multiple sources is acquired based on standard communication protocols and collected in a hierarchical manner according to priority.

[0024] Based on the collected raw data, the processed standardized data is obtained, and the transient features of electrical quantities are extracted through wavelet transform.

[0025] Based on the processed standardized data, a historical data sample storage system is established, and multi-source dataset samples are provided for model training through the data query function.

[0026] By using multidimensional device feature differential fitting and combining it with a long short-term memory network model for periodic prediction, an initial device health index within a preset range is obtained, and then corrected to obtain the true device health index.

[0027] Based on real equipment health indices, potential faulty components are analyzed and located, and maintenance strategies are generated and sent to the terminals of maintenance personnel.

[0028] The beneficial effects of this invention are as follows:

[0029] (1) By constructing a real-time hierarchical acquisition system that integrates multiple sources of data, it supports standardized data access for multiple protocols such as IEC61850, Modbus, and SNMP, and achieves full coverage acquisition of six major categories of key equipment data, including electrical quantities, communication data, self-test signals, environmental parameters, protection parameters, and operation and maintenance records. It also establishes a differentiated acquisition strategy of microsecond-level sampling for high-frequency data, second-level polling for medium-frequency data, and minute-level or hour-level updating for low-frequency data, which maximizes the reduction of system resource consumption while ensuring the real-time performance of key data.

[0030] (2) By using the time-frequency domain feature extraction technology based on continuous wavelet transform and the multi-dimensional equipment feature differential fitting technology, the infinitely long trigonometric function basis is converted into a finite-length wavelet basis that decays, and the transient electrical quantity features at the fault moment are accurately extracted. The contribution weight of each data sample is calculated using the random forest algorithm, and a differential feature weight matrix for different equipment models is established. Low contribution features are eliminated by calculating the weighted fitting degree distance, which significantly improves the effectiveness of feature data and the efficiency of model training.

[0031] (3) By using the collaborative mechanism of LSTM network model and expert knowledge rule base, the initial health index is obtained by periodic prediction based on 72 hours of historical data. Then, the influence factors of the expert knowledge rule base are dynamically corrected, realizing the deep integration of data-driven and knowledge-driven approaches. A historical database module with a hierarchical storage architecture is established, realizing the second-level query response of TB-level data, which improves the accuracy and interpretability of health status assessment.

[0032] (4) Through a graded early warning mechanism and differentiated maintenance strategy, when the equipment health index is lower than the preset threshold, different levels of early warning are automatically triggered. Combined with the fault tree analysis method, potential faulty components are accurately located and differentiated maintenance suggestions are generated. At the same time, a system architecture design that does not require the deployment of intelligent detection chips on the equipment side is adopted. The health diagnosis function can be realized by simply obtaining equipment data through the standard communication protocol, which greatly reduces the system deployment complexity and cost. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a schematic diagram of a secondary equipment health diagnosis system based on multi-source data according to an embodiment of the present invention;

[0035] Figure 2 This is a schematic block diagram of a data acquisition module in a secondary equipment health diagnosis system based on multi-source data according to an embodiment of the present invention;

[0036] Figure 3 This is a flowchart of the data processing module in a secondary equipment health diagnosis system based on multi-source data according to an embodiment of the present invention;

[0037] Figure 4 This is a schematic diagram of the LSTM network structure used in the equipment health diagnosis module of a secondary equipment health diagnosis system based on multi-source data according to an embodiment of the present invention.

[0038] Figure 5 This is a detailed implementation diagram of an equipment health diagnosis module in a secondary equipment health diagnosis system based on multi-source data according to an embodiment of the present invention;

[0039] Figure 6 This is a schematic flowchart of a secondary equipment health diagnosis method based on multi-source data according to an embodiment of the present invention.

[0040] In the picture:

[0041] 1. Data acquisition module; 2. Data processing module; 3. Historical database module; 4. Equipment health diagnosis module; 5. Early warning and decision-making module. Detailed Implementation

[0042] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.

[0043] According to embodiments of the present invention, a secondary equipment health diagnosis system and method based on multi-source data are provided.

[0044] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, according to an embodiment of the present invention, a secondary equipment health diagnosis system based on multi-source data is provided, the secondary equipment health diagnosis system based on multi-source data includes:

[0045] Data acquisition module 1 is used to acquire multi-source data based on standard communication protocols and to perform hierarchical acquisition according to priority order;

[0046] Data processing module 2 is used to obtain processed standardized data based on the collected raw data, and extract transient features of electrical quantities through wavelet transform;

[0047] Historical database module 3 is used to establish a historical data sample storage system based on the processed standardized data, and to provide multi-source dataset samples for model training through data query function;

[0048] Equipment health diagnosis module 4 is used to use multi-dimensional equipment feature differential fitting, combined with long short-term memory network model to make periodic predictions, to obtain an initial equipment health index within a preset range, and then correct it to obtain the real equipment health index.

[0049] The early warning and decision-making module 5 is used to analyze and locate potential faulty components based on the real equipment health index, and generate maintenance strategies to be sent to the operation and maintenance personnel's terminal.

[0050] Specifically, with the widespread application of smart substations, the intelligence level of secondary equipment has been rapidly improved, and the complexity of secondary equipment has also increased significantly. To better achieve real-time monitoring, early fault warning, and quantitative assessment of health status, thereby improving power grid reliability, this invention provides a secondary equipment health diagnosis system and method that integrates multi-source data and intelligent algorithms. This system enables real-time monitoring, early fault warning, and quantitative assessment of health status, thus improving power grid reliability. Figure 1 As shown, the secondary equipment health diagnosis system consists of five modules: data acquisition, data processing, historical database, equipment health diagnosis, and early warning decision-making.

[0051] Specifically, the functions of the five modules are as follows: Data Acquisition Module 1, to ensure the reliability of health diagnosis, supports standard communication protocols such as 61850 and Modbus, and is responsible for collecting multi-source data, including real-time and historical data, such as electrical quantities like voltage and current during real-time operation of secondary equipment, self-test signals like CPU load and memory status, communication data like packet loss rate and latency, environmental parameters like sensor temperature and humidity, and historical equipment maintenance data obtained from the maintenance terminal; Data Processing Module 2, responsible for data cleaning, removing outliers from the collected data, standardizing data from various types of equipment, and extracting electrical parameters through wavelet transform. The system includes several modules: a gas volume transient characteristic module (Module 3), a communication stability index (Module 4), a historical database module (Module 5), and an equipment health diagnosis module (Module 6). Module 4 uses a deep learning-based evaluation model and an LSTM network to fuse time-series data (e.g., continuous 72-hour operation data). The output range is 0-100, and alarm thresholds are dynamically adjusted based on equipment type and operating years. Module 5 combines an expert knowledge base to analyze and locate potential faulty components and generate maintenance suggestions, such as "immediate repair" and "observe operation," which are then sent to maintenance personnel terminals.

[0052] In one embodiment, the data acquisition module 1 includes: a multi-source acquisition module, used to acquire multi-source data based on a standard communication protocol; specifically, it includes acquiring voltage, current, electrical quantities, and protection setting parameters according to the communication protocol; acquiring CPU load and memory status self-test signal data through the self-diagnostic interface built into the Ethernet connection device; and acquiring packet loss rate and latency communication data using network packet capture tools or network management protocols; and a priority setting module, used to establish a real-time hierarchical priority system based on the degree of impact of different data types on the health status of the equipment, and acquiring sensor temperature and humidity environmental parameters through the protocol interface, and utilizing maintenance... The system interface acquires historical equipment operation and maintenance data; the hierarchical acquisition module is used to acquire high-frequency, medium-frequency, and low-frequency data according to a real-time hierarchical priority system, and dynamically adjusts the data acquisition strategy according to network status; the timing control module is used to sample high-frequency data at the microsecond level, poll medium-frequency data at the second level, and update low-frequency data at the minute or hour level according to data accuracy and timeliness requirements, and execute acquisition tasks according to priority order; among them, high-frequency data includes electrical quantities; medium-frequency data includes communication data, self-test signal data, and environmental parameters; low-frequency data includes protection parameters and operation and maintenance records.

[0053] Specifically, data acquisition module 1 is used to achieve real-time, hierarchical, multi-source data acquisition. The health of secondary equipment in smart substations is affected by numerous factors, involving many objective factors such as environment, electrical, mechanical, and communication. In addition, during equipment operation and maintenance, failure to replace components that have reached the end of their lifespan in a timely manner, maintenance only after each failure leading to the long-term accumulation of latent faults, and accelerated overall deterioration; failure to upgrade firmware in a timely manner leading to logical vulnerabilities, such as incorrect protection settings causing malfunctions, frequent malfunctions increasing hardware wear and tear, frequent plugging and unplugging of interfaces during debugging, and unauthorized modification of parameter settings, which may directly damage the hardware. These operation and maintenance management scenarios also affect the health of secondary equipment. Therefore, a complete set of equipment characteristics cannot be extracted from data from a single factor; it is necessary to collect multi-source data, including environmental, electrical, mechanical, communication, and operation and maintenance history data. Figure 2 As shown, the data acquisition module in this solution acquires multi-source data in the following ways: 1) acquiring electrical quantities such as voltage and current, and protection parameters such as protection settings through the 61850 communication protocol; 2) acquiring self-test signal data such as CPU load and memory status through the self-diagnostic interface built into the Ethernet connection device; 3) acquiring communication data such as packet loss rate and delay through network packet capture tools or SNMP protocol; 4) acquiring environmental parameters such as sensor temperature and humidity through the Modbus protocol; 5) the intelligent substation operation and maintenance system usually records the operation and maintenance records of each secondary device, and the data acquisition module can obtain the historical operation and maintenance data of the equipment through the operation and maintenance system.

[0054] Specifically, the aforementioned multi-source data is diverse, and different data have varying impacts on the equipment, thus their priorities also differ. Firstly, network status is crucial for data acquisition; a good network environment ensures the accuracy and completeness of the collected data. Therefore, communication data such as packet loss rate and latency have higher priority. Maintenance records, on the other hand, are generally updated only when maintenance personnel perform equipment maintenance, resulting in a low update frequency and thus lower priority. Based on the maintenance experience of secondary equipment in operational smart substations, the priority of the aforementioned multi-source data acquisition in this solution is arranged from highest to lowest as follows: 1) Communication data; 2) Electrical quantities; 3) Self-test signal data; 4) Environmental parameters; 5) Protection parameters; 6) Maintenance records.

[0055] Specifically, the precision and timeliness of different data vary. For example, fault analysis requires fault time accuracy down to the microsecond level, necessitating the recording of instantaneous changes. Electrical quantities such as voltage and current have high precision and strong timeliness, but environmental parameters such as temperature and humidity do not require such high precision. If the same strategy is used to collect all these data, a large amount of invalid data will be generated, and outliers may also exist, negatively impacting subsequent data processing. Therefore, this solution adopts a real-time hierarchical acquisition method. High-frequency data such as electrical quantities are sampled at the microsecond level, medium-frequency data such as communication, self-test signals, and environmental parameters are polled at the second level, and low-frequency data such as protection parameters and maintenance records are updated at the minute or hour level according to the actual operating environment. If multiple acquisition tasks need to be executed at the same time, such as acquiring communication data and electrical quantities, the task of acquiring communication data is executed first, followed by the task of acquiring electrical quantities, since communication data has higher priority.

[0056] In one embodiment, the data processing module 2 includes: a data cleaning module, used to clean the collected raw data based on a multi-rule data cleaning strategy, including filling in the missing parts of time-series data using linear interpolation and deduplicating duplicate data according to timestamps; an anomaly filtering module, used to dynamically filter out anomalies within a preset time range according to the three-standard-deviation principle and remove anomalies with deviations exceeding a threshold; and a feature extraction module, used to perform time-frequency analysis processing on electrical quantities at the fault time based on continuous wavelet transform, extract transient features of electrical quantities, and store the processed standardized data into a historical database module.

[0057] In one embodiment, the feature extraction module, when performing time-frequency analysis on electrical quantities at the fault time based on continuous wavelet transform to extract transient features of electrical quantities, includes: converting an infinitely long trigonometric function basis into a finite-length decaying wavelet basis based on a basis-changing operation to obtain a mother wavelet function; adjusting the mother wavelet function using scaling and translation parameters to establish a set of wavelet basis functions adapted to different time-frequency scales; performing convolution operations between the wavelet basis function set and the electrical quantity signal at the fault time to obtain multi-scale time-frequency transform coefficients; and performing progressive refinement processing based on the multi-scale time-frequency transform coefficients to extract transient features of electrical quantities at the fault time.

[0058] In one embodiment, the stepwise refinement process based on multi-scale time-frequency transformation coefficients and the extraction of transient electrical quantity features at the fault moment includes: based on multi-scale time-frequency transformation coefficients, the electrical quantity signal is gradually refined through scaling and translation operations to obtain time-frequency distributions at different scales; the start and end times corresponding to specific frequencies are determined according to the time-frequency distributions, and a correspondence between frequency and time windows is established; using the correspondence between frequency and time windows, the signal is subdivided into high-frequency components and low-frequency components, and key frequency bands sensitive to faults are identified; based on the amplitude changes and frequency distributions of the key frequency bands, transient electrical quantity features containing fault moment characteristic information are extracted.

[0059] Specifically, data processing module 2 is used to perform multi-rule data cleaning and processing. The data collected by data acquisition module 1 is raw data. During the acquisition process, various reasons (data loss due to network fluctuations, abnormal electrical quantity jumps, duplicate data uploads, etc.) may cause abnormal data to be mixed in. Therefore, data processing module 2 needs to clean the collected data. First, network fluctuations or sensor anomalies may cause the loss of time-series data, such as the loss of temperature data at a certain moment. For this data loss problem, this solution uses a linear interpolation method to fill in the missing data, that is, to take the average value within a certain time interval before and after the loss, such as... Figure 3 As shown; then, duplicate data caused by message retransmission is deduplicated based on their timestamps.

[0060] Specifically, data collected under normal operating conditions may contain outliers that differ significantly from normal values. According to the 3σ principle, if the data follows a normal distribution, an outlier is defined as a value whose deviation from the mean exceeds three standard deviations. In other words, under the assumption of a normal distribution, the probability of a value exceeding three standard deviations from the mean is less than 0.003, and therefore it can be considered an outlier. This solution dynamically filters outliers based on the 3σ principle. Within a certain time range (which can be adjusted according to specific equipment and operating environment), if a data point x satisfies the formula |x-μ|>3σ, it is identified as an outlier and removed from the dataset. μ is the rated electrical quantity of the equipment, and σ is the standard deviation of the electrical quantity.

[0061] Specifically, after data cleaning, preprocessing is performed to more comprehensively extract data features and remove noise and redundancy. In data preprocessing, most signals can be transformed using Fourier transform, but directly processing data using the Fourier transform function will lose time data, making it impossible to determine the start and end times corresponding to specific frequencies. Therefore, this invention uses continuous wavelet transform (CWT) to preprocess electrical quantities at the fault time, extracting transient features such as the amplitude-frequency joint features at the fault time, and screening key frequency bands sensitive to faults. Continuous wavelet transform provides a "time-frequency" window that changes with frequency. Through basis-changing operations, it replaces the infinitely long trigonometric function basis with a finite-length decaying wavelet basis, thus obtaining not only frequency but also time, enabling time-frequency analysis. Furthermore, through scaling and translation operations, the signal is progressively refined across multiple scales, ultimately achieving a subdivision of high and low frequencies, adaptively analyzing time-frequency signals. The calculation formula for a single electrical quantity is shown below:

[0062]

[0063] In the formula, f(t) is the input signal. This is the mother wavelet, where 'a' is a scaling parameter (not equal to 0) used to control the scaling of the wavelet, and 'b' is a translation parameter controlling the position of the wavelet on the time axis. Among common mother wavelets, the Morlet wavelet is suitable for frequency analysis; therefore, the mother wavelet here is... The formula is shown below:

[0064]

[0065] In the formula, ω0 is the center frequency. Finally, the data samples processed by the above rules are stored in the historical database to provide the necessary data samples for equipment model training and health status prediction.

[0066] It should be noted that: ① The wavelet basis function set is formed by scaling the mother wavelet function at different scales and translating it along the time axis. The scaling parameter 'a' controls the frequency resolution of the wavelet, and the translation parameter 'b' controls the position of the wavelet on the time axis, thus constructing a function set suitable for analysis at different time and frequency scales. ② Convolution operation refers to the mathematical convolution calculation between the wavelet basis function set and the electrical signal at the fault time. Similarity is calculated point-by-point using a sliding window method to extract the signal's feature information at different time and frequency scales. ③ The multi-scale time-frequency transform coefficients are a coefficient matrix obtained through continuous wavelet transform. These coefficients reflect the energy distribution of the electrical signal at different time points and frequency scales, providing basic data for subsequent refinement processing. ④ The correspondence between frequency and time window refers to the ability of continuous wavelet transform to accurately determine the specific time period in which a specific frequency component appears in the signal, solving the problem that traditional Fourier transform cannot simultaneously obtain time and frequency information.

[0067] In one embodiment, the equipment health diagnosis module 4 includes: a feature optimization module, used to obtain standardized feature data based on multi-source data samples of different types of secondary equipment using normalization processing, and to remove low-contribution features through multi-dimensional equipment feature differentiation fitting to obtain optimized feature data samples; a health assessment module, used to perform periodic prediction using an LSTM network model based on the optimized feature data samples, and to obtain an initial equipment health index through the synergistic effect of multiple gating mechanisms; and a rule correction module, used to determine the set of activated influencing factors based on an expert knowledge rule base, and to perform rule correction on the initial equipment health index to obtain the true equipment health index.

[0068] In one embodiment, when the feature optimization module obtains optimized feature data samples by differential fitting of multi-dimensional device features and eliminating low-contribution features, the following steps are included: based on the multi-source heterogeneous data samples acquired by the data acquisition module, calculating the contribution weight of each data sample to the health status of the secondary equipment using a feature evaluation model, and establishing a differential feature weight matrix according to the equipment model; constructing historical normal operation data as feature reference nodes, and calculating the fitting distance between the current sample data and the feature reference nodes; wherein, a larger fitting distance value indicates that the sample data is closer to the normal state, and a smaller fitting distance value indicates that the sample data deviates from the normal state; based on the fitting distance... The algorithm combines the differential feature weight matrix to calculate the sum of weighted fitting distances between each sample data and the feature reference node. When the sum of weighted fitting distances is less than or equal to a preset threshold, the sample data is retained; when the sum of weighted fitting distances is greater than the preset threshold, it is determined to be a low-contribution feature and removed from the feature data sample. Based on the functional type and operating environment of the secondary equipment, fault mode feature vectors are extracted from self-test data, degradation trend feature vectors are extracted from environmental monitoring data, performance degradation feature vectors are extracted from protection action data, and network anomaly feature vectors are extracted from communication status data. Multi-dimensional feature data samples are generated through feature fusion algorithms.

[0069] In one embodiment, the multi-source heterogeneous data sample includes self-test data, environmental monitoring data, protection action data, and communication status data; the fault mode feature vector includes the number of self-test failures, the frequency of specific fault codes, the duration of communication interruption, and the self-test cycle deviation; the degradation trend feature vector includes the mean, maximum and minimum values, and rate of change of temperature and humidity, as well as the temperature distribution characteristics of internal hot spots; the performance degradation feature vector includes the number of protection actions, the dispersion of action time, the frequency of failure to act and false action, the protection setting, the deviation of actual operating parameters, and the accuracy of protection action timing; the network anomaly feature vector includes communication delay, packet loss rate, error frame rate, the mean and variance of communication traffic, and the frequency of burst communication.

[0070] In one embodiment, when the rule correction module determines the set of activated influencing factors based on the expert knowledge rule base and corrects the initial equipment health index to obtain the true equipment health index, the following steps are included: establishing a quantitative mapping relationship between fault symptoms and health index decay based on equipment failure modes and fault evolution patterns in a historical fault database; constructing an expert knowledge rule base and setting corresponding influencing factors according to the severity of different fault types; calculating the probability value of triggering various fault rules in the current equipment state by analyzing the similarity between real-time monitoring data and historical fault precursor data based on the expert knowledge rule base, and determining the set of activated influencing factors based on the comparison result of the probability value and the preset probability threshold; fusing the influencing factors in the activated influencing factor set using Bayesian inference methods, and calculating a comprehensive correction coefficient by combining the equipment's operating years, maintenance records, and environmental factors; and obtaining the true equipment health index by weighting the initial equipment health index with the comprehensive correction coefficient.

[0071] Specifically, the equipment health diagnosis module 4 is used to implement health status prediction driven by a hybrid model. For different types or models of secondary equipment, the factors affecting their health may differ. For example, protection device A is designed with heat dissipation in mind, resulting in stronger heat dissipation performance than general equipment, thus temperature has a smaller impact on device A. Better heat dissipation performance may mean increased impact of humidity on equipment health. To better extract features from the secondary equipment model, this invention uses a random forest to evaluate feature importance and remove low-contribution features, as well as features with importance higher than a set threshold, such as... Figure 5 As shown. Due to the inconsistent types of various data, this invention first normalizes the data samples when constructing the random forest model:

[0072] s c (x, p)=1-|v(x, c)-v(p, c)|;

[0073] In the formula, s c This represents the output of the random forest model after normalizing the data samples; x represents a certain type of data sample; p represents a random reference sample, which is derived from the data sample; c represents the feature parameter of the data sample x; v(x, c) represents the normalized value of the data sample x on feature c; v(p, c) represents the feature parameter of the reference sample.

[0074] Specifically, when extracting feature values ​​from different types of data, this invention can obtain counting features related to secondary equipment self-test data from self-test data such as the number of self-test failures and the frequency of specific fault codes; extract state features related to self-test data from self-test data such as communication interruption duration and self-test cycle deviation; extract features related to environmental monitoring data from data such as the mean, maximum and minimum values, rate of change, and internal hotspot temperature distribution; extract features related to protection action data from data such as the number of protection actions, action time dispersion, frequency of failure to operate / false operation, deviation between protection settings and actual operating parameters, and accuracy of protection action timing; and extract features related to communication status data from data such as communication delay, packet loss rate, error frame rate, mean and variance of communication traffic, and burst communication frequency.

[0075] Specifically, secondly, for any sample data, its distance to the node can be expressed as:

[0076]

[0077] In the formula, d[g0, g(x)] i ] represents any sample data x i The distance between d[g0, g(x) and node g0]. i The larger the value of ], the stronger the sample data x. i The higher the goodness of fit of the feature parameters of the reference sample corresponding to g0, the better; conversely, the better the goodness of fit of d[g0, g(x)]. i The smaller the value of ], the stronger the sample data x. i The lower the goodness of fit of the feature parameters of the reference sample corresponding to g0.

[0078] Specifically, the final distance calculation method between any sample data and different nodes in the random forest model can be expressed as:

[0079]

[0080] In the formula, D[g j g(x) i [] represents the sum of distances between any sample data and different nodes in the random forest model. i and j are sample indices, and their values ​​range from 1 to the total number of samples N. It is necessary to calculate the distance values ​​of all N*N possible (i,j) pairs (including i=j), with the upper and lower limits depending on the total number of samples. When D[g] j g(x) i When D[g] is less than or equal to the maximum allowable threshold of the sample data, j g(x) i The reference sample with the smallest value is used as the sample data x. i The integrated object. When D[g j g(x)i If the value of x exceeds the maximum allowable threshold for sample data integration, discard the current sample data. i For example, after evaluation using a random forest, the historical data of device A yields an importance of 0.35 for CPU load rate, 0.51 for ambient temperature, and 0.83 for humidity. Given a maximum allowable threshold of 0.8 for sample data, humidity data is a low-importance feature for device A. Therefore, when analyzing the health of device A, humidity data should not be included in the feature data sample to reduce noise interference and improve model accuracy.

[0081] Specifically, after extracting the optimized device feature data samples, this invention triggers LSTM prediction with a 72-hour cycle. The LSTM network model (Long Short-Term Memory network model) consists of an input layer, hidden layers, and an output layer, as illustrated in the diagram below. Figure 4 As shown, the input layer consists of data from the feature data samples, and the hidden layer is composed of recurrently connected hidden units. Each neural unit contains three gates (input gate σ). i Output gate σ o And the forgetting gate σ f Each of these is responsible for write, read, and select operations, respectively. The forward propagation formula is as follows:

[0082]

[0083] In the formula, W xi W hi W ci W represents the weight matrix of the input gate. xf W hf The weight matrix representing the forget gate; W xc W hc The weight matrix representing the memory unit; W xo W ho W co The sum of the weight matrices representing the output gates; b i b if b represents the bias vector of the input gate; c The bias vector representing the memory unit; b o The bias vector representing the output gate; x t and h t These are the hidden layer input and output of the LSTM, respectively; c t For memory units; σ represents the sigmoid activation function, the specific formula of which is σ(x)=1 / (1-e -x It can map a real number to the interval (0, 1). Input gate σ i The input quantity x is determined t How much can be stored in memory; the gate of forgetting σ fThe decision was made regarding the previous memory c t-1 How much can be retained in this memory? t In LSTM, historical information can be selectively retained or forgotten, which is a key component and can avoid the gradient vanishing and exploding problems caused during gradient backpropagation; the output gate σ o The memory unit c is determined t The output h of the current neuron t The impact of this is to determine how much historical memory needs to be called for the current output, and finally the output layer outputs the device health index.

[0084] Specifically, the equipment health diagnosis and prediction module 4 reads feature data samples from the historical database within 72 hours for model inference and outputs an equipment health index ranging from 0 to 100. Since model inference can introduce errors and misjudgments, this invention utilizes the aforementioned maintenance data combined with an expert knowledge base for rule verification and correction. An LSTM and expert rule collaboration mechanism is established for several common maintenance scenarios. An influencing factor is given to the impact on device health, and expert rules are triggered based on maintenance records within the equipment health calculation period. The health index is corrected based on the health inference derived by the LSTM prediction, combined with the influencing factor μ. An expert knowledge base is constructed based on historical maintenance experience, and a rule base is designed, with corresponding influencing factors μ given according to specific rules. The influencing factor μ is formulated based on historical operating data analysis of the same model of secondary equipment and records of manual corrections in historical cases. When formulating the influencing factor, it is first necessary to clarify the specific maintenance scenario or equipment status defined by the rule, and define trigger thresholds according to different scenarios and statuses. These thresholds are typically derived from equipment technical specifications, industry standards, manufacturer recommendations, and historical fault statistical analysis. Secondly, from the historical operations and maintenance database and health status records, all cases that meet the triggering conditions of this rule are selected. The correlation between the health index predicted by the model (or the health status assessed at the time) and the subsequent actual equipment status (whether a failure will occur soon, whether performance will significantly decline, or whether emergency maintenance is required) is analyzed. Historical cases are reviewed to see how experienced operations and maintenance engineers subjectively adjusted their judgments of equipment health when such situations occurred, and the proportion or magnitude of these subjective corrections is collected. Finally, the proportion of actual equipment failures or significant performance declines (requiring intervention) within a certain period after the rule is triggered (e.g., within the next maintenance cycle, within one week, within one month) is statistically analyzed and compared with the baseline proportion when the rule is not triggered to determine a conservative and reasonable influencing factor. If multiple rules are triggered simultaneously within the same time period, the multiple influencing factors need to be multiplied to obtain a comprehensive factor. The calculation formula is as follows:

[0085]

[0086] In the formula, y t This is the final device health index obtained at the current moment.

[0087] Specifically, the expert knowledge base in this invention is pre-set with the following three rules:

[0088] 1) Correction rule for contact resistance exceeding limits: If the contact resistance is greater than 0.1Ω, a correction factor of 0.7 is given;

[0089] 2) Correction rule for frequent device malfunctions: If the number of recent malfunctions exceeds a certain number, a correction factor of 0.5 is given;

[0090] 3) Correction rules for high humidity environments: Considering the insulation aging problem, if the ambient humidity is greater than 75% and the service life is greater than 5 years, a correction factor of 0.8 is given.

[0091] Specifically, if during a certain period of device operation, the relay contact resistance is 0.12Ω, the humidity is 85%, and there are two false trips, triggering the above three rules, then the influence factor μ = 0.7 × 0.5 × 0.8 = 0.28. The final equipment health index is obtained by multiplying the health index predicted by the model by 0.28. If none of the above rules are triggered during this period, the influence factor defaults to 1. The hybrid model drives the quantitative assessment of equipment health status and provides corresponding operation and maintenance strategies. Based on this assessment, maintenance personnel can understand the current health level of the equipment and implement appropriate management measures according to the operation and maintenance strategies.

[0092] It should be noted that: ① The feature evaluation model uses the random forest algorithm to assess the importance of each data sample to the secondary equipment health status, quantifying its contribution weight by calculating the information gain of the feature during the decision tree splitting process. ② The contribution weight refers to the degree of influence of each type of data sample on the accuracy of equipment health status prediction, represented by the feature importance score calculated by the random forest model. The higher the importance score, the greater the contribution of the feature to the health status judgment. ③ The differentiated feature weight matrix is ​​a matrix constructed by assigning different weight coefficients to various features based on the technical characteristics and operating environment differences of different equipment models. For example, equipment with good heat dissipation performance has a lower weight for temperature features and a relatively higher weight for humidity features. ④ The feature fusion algorithm merges multiple feature vectors into a unified multi-dimensional feature data sample through weighted averaging. During the fusion process, the contribution weight of each feature vector is considered to ensure that important features occupy a larger proportion in the fusion result. ⑤ The quantitative mapping relationship between fault symptoms and health index decay is established by analyzing the various symptoms before equipment failure in the historical fault database, establishing a mathematical relationship between the intensity of symptoms and the decline in the health index, providing a quantitative basis for setting the influence factors of expert rules. ⑥ Similarity calculation uses Euclidean distance or cosine similarity algorithms. By comparing the similarity between real-time monitoring data and historical fault precursor data in the feature vector space, it determines whether the current equipment state is close to the historical fault mode. ⑦ The comparison between probability values ​​and preset probability thresholds refers to comparing the calculated fault rule trigger probability with the system's preset threshold. When the probability value exceeds the threshold, the fault rule is determined to be activated, and the corresponding influencing factor is included in the influencing factor set. ⑧ The Bayesian inference method, through the calculation of prior probability and likelihood probability, combined with the equipment's historical operating data and current state information, infers the comprehensive influence of each influencing factor on the equipment health index, achieving intelligent fusion of multiple influencing factors. ⑨ Weighted operation refers to multiplying the initial equipment health index with the comprehensive correction coefficient to obtain the true equipment health index corrected by expert knowledge. The calculation formula is: True health index = Initial health index × Comprehensive correction coefficient.

[0093] To facilitate understanding of the above-mentioned technical solution of the present invention, the following is a detailed description using a 220kV relay protection device of a substation as an example:

[0094] First, data acquisition module 1 acquires multi-source data from the relay protection device via the IEC 61850 standard communication protocol. Based on a real-time priority hierarchy, the system sets electrical quantities such as voltage and current as high-frequency data, acquiring them at a microsecond-level sampling frequency; CPU load, memory status, communication latency, and packet loss rate are set as medium-frequency data, acquired using a second-level polling method; and protection settings and maintenance records are set as low-frequency data, acquired using a minute- or hourly update frequency. When the network is in good condition, the system acquires data at the standard frequency; when the network is congested, it automatically reduces the acquisition frequency of medium and low-frequency data, prioritizing the real-time performance of high-frequency data.

[0095] Next, data processing module 2 cleans the collected raw data. The system employs a multi-rule data cleaning strategy, using linear interpolation to fill in missing time-series data and deduplicating duplicate data based on timestamps. The anomaly filtering module dynamically identifies outliers within a preset one-hour time window based on the three-standard-deviation principle. For example, if the current value at a certain moment exceeds the mean ± 3σ range, the system automatically removes the outlier data. The feature extraction module performs time-frequency analysis on the electrical quantities at the fault time based on continuous wavelet transform. Through basis transformation, it converts the infinitely long trigonometric function basis into a finite-length decaying wavelet basis. Using scaling and translation parameters, it establishes a set of wavelet basis functions adapted to different time-frequency scales, ultimately extracting the transient features of electrical quantities containing fault time characteristics.

[0096] Subsequently, historical database module 3 establishes a historical data sample storage system based on the processed standardized data. This module adopts a hierarchical storage architecture, storing the most recent month's data in high-speed storage media to support fast queries, storing data from 3 to 12 months in medium-speed storage media, and storing historical data older than one year in low-cost long-term storage media. The data query function supports multi-dimensional retrieval, including data filtering by time range, device type, fault type, environmental conditions, etc. The system has established an indexing mechanism, capable of completing TB-level data query operations within seconds, providing efficient multi-source dataset sample acquisition capabilities for model training. Simultaneously, this module also has data backup and recovery functions, ensuring data reliability and availability through a master-slave database architecture.

[0097] Then, the equipment health diagnosis module 4 utilizes multi-dimensional equipment feature differential fitting technology, combined with an LSTM network model, to perform periodic predictions. The feature optimization module, based on the technical characteristics of relay protection devices, uses the random forest algorithm to calculate the contribution weight of each data sample and establish a differential feature weight matrix. For example, for device A with good heat dissipation performance, the temperature feature weight is set to 0.35, the humidity feature weight to 0.83, and the CPU load rate weight to 0.51. The system constructs historical normal operation data as feature reference nodes, calculates the fitting distance between the current sample data and the reference nodes, and when the sum of the weighted fitting distances exceeds a preset threshold of 0.8, it is determined to be a low-contribution feature and is removed. The health assessment module uses a 72-hour prediction period and utilizes the input gate, forget gate, and output gate collaborative mechanism of the LSTM network model to output an initial equipment health index within the range of 0-100. The rule correction module corrects rules based on an expert knowledge rule base. When the contact resistance exceeds 0.1Ω, a correction factor of 0.7 is applied; when the device frequently malfunctions, a correction factor of 0.5 is applied; and when the device is in a high humidity environment and has been in operation for more than 5 years, a correction factor of 0.8 is applied. Finally, the true equipment health index is calculated by multiplying multiple influencing factors.

[0098] Finally, the early warning and decision-making module 5 locates potential faulty components based on the analysis of the actual equipment health index and generates maintenance strategies. This module adopts a tiered early warning mechanism: a yellow warning is triggered when the equipment health index is below 95, an orange warning below 85, and a red warning below 70. The system uses fault tree analysis, combined with the changing trend of the equipment health index and abnormal conditions of key characteristic parameters, to pinpoint the specific location of potential faulty components. For example, when the health index of a relay protection device drops from 85 to 72, and the contact resistance abnormally increases and the frequency of malfunctions increases, the system automatically determines that the contact components have a potential fault risk. The decision-making module generates differentiated maintenance strategies based on factors such as the equipment's importance level, the degree of fault risk, and maintenance costs, including different levels of recommendations such as immediate maintenance, planned maintenance, and status monitoring. The generated maintenance strategies are sent to the maintenance personnel's terminals in real time via SMS, email, and APP push notifications, ensuring that fault risks can be addressed promptly, thereby improving the operational reliability and maintenance efficiency of secondary equipment.

[0099] like Figure 6 As shown, according to another embodiment of the present invention, a secondary equipment health diagnosis method based on multi-source data is also provided, the method comprising:

[0100] S 1. Acquire multi-source data based on standard communication protocols and perform hierarchical collection according to priority order;

[0101] S2. Based on the collected raw data, obtain the processed standardized data, and extract the transient features of electrical quantities through wavelet transform;

[0102] S3. Based on the processed standardized data, establish a historical data sample storage system, and provide multi-source dataset samples for model training through data query function;

[0103] S4. Utilize multidimensional device feature differential fitting, combine with long short-term memory network model for periodic prediction, obtain the initial device health index within the preset range, and correct it to obtain the real device health index.

[0104] S5. Based on the actual equipment health index, analyze and locate potential faulty components, and generate maintenance strategies to send to the operation and maintenance personnel's terminals.

[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A secondary equipment health diagnosis system based on multi-source data, characterized in that, The application relates to a device health diagnosis method and device. The application comprises: a data acquisition module for acquiring multi-source data based on a standard communication protocol and performing hierarchical acquisition according to a priority order; a data processing module for acquiring processed standardized data from the acquired original data and extracting electrical quantity transient characteristics through wavelet transformation; a historical database module for establishing a historical data sample storage system based on the processed standardized data and providing multi-source data set samples for model training through a data query function; a device health diagnosis module for obtaining an initial device health index within a preset range through periodic prediction by combining a multi-dimensional device feature differentiation fitting and a long short-term memory network model and correcting the initial device health index to obtain a real device health index; 2. The secondary equipment health diagnostic system based on multi-source data according to claim 1, wherein, an early warning and decision module for analyzing and positioning potential fault elements based on the real device health index and generating a maintenance strategy to be sent to an operation and maintenance personnel terminal. The data acquisition module comprises: a multi-source acquisition module for acquiring multi-source data based on a standard communication protocol; specifically, voltage, current electrical quantities and protection parameter protection values are acquired according to a communication protocol, CPU load and memory state self-check signal data are acquired through a self-diagnosis interface built in an Ethernet connection device, and message packet loss rate and delay communication data are acquired by using a network packet capturing tool or a network management protocol; a priority setting module for establishing a real-time hierarchical priority system according to the influence degree of different data types on the device health state, acquiring sensor temperature and humidity environmental parameters through a protocol interface and acquiring device operation and maintenance historical data by using an operation and maintenance system interface; a hierarchical acquisition module for acquiring high-frequency data, medium-frequency data and low-frequency data according to the priority order based on the real-time hierarchical priority system and dynamically adjusting a data acquisition strategy according to a network state; a timing control module for performing microsecond-level sampling on high-frequency data, second-level polling on medium-frequency data and minute or hour-level updating on low-frequency data according to data precision and timeliness requirements and executing acquisition tasks according to the priority order; 3. The secondary equipment health diagnostic system based on multi-source data according to claim 1, wherein, wherein the high-frequency data comprise electrical quantities, the medium-frequency data comprise communication data, self-check signal data and environmental parameters and the low-frequency data comprise protection parameters and operation and maintenance records. The data processing module comprises: a data cleaning module for cleaning the acquired original data based on a multi-rule data cleaning strategy, including filling in lost time series data by using a linear interpolation method and removing repeated data according to a time stamp; an abnormality filtering module for dynamically filtering abnormal values within a preset time range according to a three-standard-deviation principle and removing abnormal data with a deviation exceeding a threshold value; 4. The secondary equipment health diagnostic system based on multi-source data according to claim 3, wherein, a feature extraction module for performing time-frequency analysis and processing on electrical quantities at a fault moment based on continuous wavelet transformation, extracting electrical quantity transient characteristics and storing the processed standardized data into the historical database module. The method for performing time-frequency analysis and processing on electrical quantities at a fault moment based on continuous wavelet transformation and extracting electrical quantity transient characteristics comprises: converting an infinite long trigonometric function base into a finite long wavelet base which attenuates by using a variable base operation to obtain a mother wavelet function; The mother wavelet function is adjusted by scaling and translation parameters to establish a wavelet base function group suitable for different time-frequency scales; The electrical quantity signal at the fault time is convoluted with the wavelet base function group to obtain multi-scale time-frequency transform coefficients; The electrical quantity transient state feature at the fault time is extracted through step-by-step refinement based on the multi-scale time-frequency transform coefficients.

5. The secondary equipment health diagnostic system based on multi-source data according to claim 4, wherein, The step-by-step refinement based on the multi-scale time-frequency transform coefficients and the extraction of the electrical quantity transient state feature at the fault time include: The electrical quantity signal is refined step by step through scaling and translation operation based on the multi-scale time-frequency transform coefficients to obtain time-frequency distribution under different scales; The start time and end time corresponding to a specific frequency are determined according to the time-frequency distribution to establish the corresponding relationship between the frequency and the time window; The signal is subdivided into high-frequency components and low-frequency components using the corresponding relationship between the frequency and the time window, and the key frequency band sensitive to the fault is identified; The electrical quantity transient state feature containing the feature information of the fault time is extracted based on the amplitude change and frequency distribution of the key frequency band.

6. The secondary device health diagnostic system based on multi-source data of claim 1, wherein, The device health diagnosis module includes: The feature optimization module is configured to obtain standardized feature data by normalization based on multi-source data samples of different types of secondary devices, and to obtain optimized feature data samples by multi-dimensional device feature differentiation fitting and low-contribution feature elimination. The health assessment module is configured to perform periodic prediction by using an LSTM network model based on the optimized feature data samples, and to obtain an initial device health index by synergistic action of a multi-gate mechanism. The rule correction module is configured to determine an activated influence factor set according to an expert knowledge rule base, and to correct the initial device health index to obtain a real device health index.

7. The secondary equipment health diagnostic system based on multi-source data according to claim 6, wherein, The step of obtaining optimized feature data samples by multi-dimensional device feature differentiation fitting and low-contribution feature elimination includes: Based on the multi-source heterogeneous data samples obtained by the data acquisition module, the feature evaluation model is used to calculate the contribution weight of each data sample to the health state of the secondary device, and a differentiated feature weight matrix is established according to the device model. The fitting distance between the current sample data and the feature reference node is calculated by constructing the historical normal operation data as a feature reference node. Based on the fitting distance, the sum of the weighted fitting distances between each sample data and the feature reference node is calculated in combination with the differentiated feature weight matrix. When the sum of the weighted fitting distances is less than or equal to a preset threshold, the sample data is retained. When the sum of the weighted fitting distances is greater than the preset threshold, the sample data is determined as a low-contribution feature and is eliminated from the feature data samples. According to the function type and operating environment of the secondary device, the fault mode feature vector is extracted from the self-checking data, the degradation trend feature vector is extracted from the environmental monitoring data, the performance degradation feature vector is extracted from the protection action data, and the network anomaly feature vector is extracted from the communication state data. A multi-dimensional feature data sample is generated by a feature fusion algorithm.

8. The secondary equipment health diagnostic system based on multi-source data according to claim 7, wherein, The multi-source heterogeneous data samples include self-checking data, environmental monitoring data, protection action data, and communication state data. The failure mode feature vector includes the number of self-test failures, the frequency of specific fault codes, the duration of communication interruption, and the self-test cycle deviation value; the degradation trend feature vector includes the mean, maximum, rate of change of temperature and humidity, and internal hot spot temperature distribution characteristics; the performance degradation feature vector includes the number of protection actions, action time dispersion, frequency of refusal and misoperation, protection setting value, actual operating parameter deviation, and protection action timing accuracy; and the network anomaly feature vector includes communication delay, packet loss rate, error frame rate, mean and variance of communication traffic, and burst communication frequency.

9. The secondary device health diagnostic system based on multi-source data according to claim 6, wherein, The method comprises: Based on the device failure mode and failure evolution law in the historical failure database, a quantitative mapping relationship between failure symptoms and health index decay is established, an expert knowledge rule base is constructed, and corresponding influence factors are set according to the hazard degree of different fault types; According to the expert knowledge rule base, the probability value of triggering each type of fault rule is calculated by analyzing the similarity of real-time monitoring data and historical failure precursor data, and the activated influence factor set is determined according to the comparison result of the probability value and the preset probability threshold value; The influence factors in the activated influence factor set are fused by using the Bayesian inference method, and the comprehensive correction coefficient is calculated combined with the device running time, maintenance record and environmental factors, and the real device health index is obtained by weighted operation of the initial device health index and the comprehensive correction coefficient.

10. A method for secondary equipment health diagnosis based on multi-source data, using the secondary equipment health diagnosis system based on multi-source data according to any one of claims 1-9, characterized in that, The method comprises: Based on the standard communication protocol, multi-source data are acquired, and hierarchical collection is performed according to the priority order; Based on the collected raw data, processed standardized data are acquired, and electrical quantity transient characteristics are extracted by wavelet transform; Based on the processed standardized data, a historical data sample storage system is established, and a multi-source data set sample is provided for model training through data query function; The initial device health index within the preset range is obtained by using multi-dimensional device feature differentiation fitting combined with long short-term memory network model periodic prediction, and the real device health index is obtained by correction; Based on the real device health index, potential fault elements are analyzed and located, and a maintenance strategy is generated and sent to the operation and maintenance personnel terminal.

Citation Information

Patent Citations

  • Equipment health prediction diagnosis method and diagnosis device based on multivariate data fusion

    CN117113282A

  • Power distribution equipment health diagnosis three-layer discrimination method, system, equipment and medium

    CN117725475A

Cited By

  • LED advertisement screen remote control system based on multi-terminal linkage

    CN121331036A

  • Intelligent chip fault detection method and system based on operation data

    CN121432152A

  • Network-based electromechanical equipment operating system and method

    CN121523298A

  • Metering type circuit breaker health monitoring method and device and metering type circuit breaker

    CN122220872A