Fault detection method and device, computer device and storage medium
Patent Information
- Application Number
- CN202610796912.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-25
AI Technical Summary
[0015]另一方面,提供了一种计算机程序产品,包括计算机程序,该计算机程序存储在计算机可读存储介质中,计算机设备的处理器从计算机可读存储介质读取该计算机程序,处理器执行该计算机程序,使得该计算机设备执行上述各个方面或者各个方面的各种可选实现方式中提供的故障检测方法。
Smart Images

Figure CN122817892A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a fault detection method, apparatus, computer equipment, and storage medium. Background Technology
[0002] With the continuous improvement of industrial automation, condition monitoring and fault diagnosis of various key equipment (such as high-voltage electrical equipment and rotating machinery) have become important links in ensuring production safety and continuous operation. During the operation of such equipment, the sensor signals (such as vibration, ultrasound, current, etc.) often contain a wealth of equipment health information. How to accurately identify fault characteristics from complex sensor signals is the core issue of fault diagnosis technology. Summary of the Invention
[0003] This application provides a fault detection method, apparatus, computer equipment, and storage medium, employing a dual verification mechanism of "anomaly detection + physical verification," which can improve the accuracy and efficiency of fault detection. The technical solution is as follows: On the one hand, a fault detection method is provided, the method comprising: Based on the signal reconstruction model, the first sensing signal collected by the device in operation is reconstructed to obtain the second sensing signal. The signal reconstruction model is trained with the sensing signal of the device when it is fault-free as the training sample and with the goal of minimizing the reconstruction error. Based on the first sensing signal and the second sensing signal, the reconstruction error of the signal reconstruction model is determined, and the reconstruction error is used to indicate the difference between the first sensing signal and the second sensing signal. If the reconstruction error is greater than the error threshold, feature extraction is performed on the first sensing signal to obtain key features, which are the fault-related physical characteristics in the sensing signal of the device. Based on the key features and physical fault characteristics, the fault detection result of the device is determined, wherein the physical fault characteristics are used to indicate the physical characteristics of the sensing signal when the device fails.
[0004] On the other hand, a fault detection device is provided, the device comprising: The reconstruction module is used to reconstruct the first sensing signal collected by the device in the operating state based on the signal reconstruction model to obtain the second sensing signal. The signal reconstruction model is trained with the sensing signal of the device when it is fault-free as the training sample and with the goal of minimizing the reconstruction error. The first determining module is used to determine the reconstruction error of the signal reconstruction model based on the first sensing signal and the second sensing signal, wherein the reconstruction error is used to indicate the difference between the first sensing signal and the second sensing signal; The feature extraction module is used to extract features from the first sensing signal when the reconstruction error is greater than the error threshold, and obtain key features, wherein the key features are the physical characteristics related to the fault in the sensing signal of the device. The second determining module is used to determine the fault detection result of the device based on the key features and the physical characteristics of the fault, wherein the physical characteristics of the fault are used to indicate the physical characteristics of the sensing signal when the device fails.
[0005] In some embodiments, the second determining module is configured to generate fault prompt words based on the key features; input the fault prompt words into a large language model, and compare the key features with the fault physical features through the large language model to obtain a fault detection report of the device. The fault detection report includes at least one of the following: the fault type of the device, the fault severity information, and the fault repair plan.
[0006] In some embodiments, the physical characteristics of the fault include the physical characteristics of the sensing signals under multiple fault types; The second determining module is used to determine that the device has a fault and output the fault type when the large language model detects that the key feature matches any physical feature corresponding to any fault type; and to determine that the device does not have a fault when the large language model detects that the key feature does not match any physical feature corresponding to any fault type.
[0007] In some embodiments, the first sensing signal includes sensing signals of multiple physical types; The second determining module is used to concatenate the key features corresponding to the multiple physical types of sensor signals to obtain the fault prompt words; input the fault prompt words into a large language model, and compare the key features of each sensor signal with the fault physical features of the corresponding physical type through the large language model to obtain the fault detection report of the device.
[0008] In some embodiments, the second determining module is configured to determine that the device is fault-free if the reconstruction error is not greater than the error threshold.
[0009] In some embodiments, the apparatus further includes: The third determining module is used to determine the error threshold based on the numerical distribution of the historical reconstruction error of the device in a fault-free state.
[0010] In some embodiments, the signal reconstruction model is a one-dimensional convolutional denoising autoencoder, which includes an encoder and a decoder. The encoder is used to extract features of the input signal, and the decoder is used to reconstruct the input signal based on the features output by the encoder. The reconstruction module is used to input the first sensing signal into the encoder, extract the signal feature vector through convolution and pooling operations of the one-dimensional convolutional layer in the encoder, and input the signal feature vector into the decoder, reconstruct the second sensing signal through upsampling and convolution operations of the one-dimensional transposed convolutional layer in the decoder.
[0011] In some embodiments, the sensing signal of the device when it is fault-free includes at least one of the following: The sensor signals collected by the device under normal operating conditions that are powered on and confirmed to be fault-free; The sensor signal segments are automatically filtered and extracted from the collected sensor signals. The sensor signal segments have a preset duration and do not contain amplitude pulses exceeding a preset value.
[0012] In some embodiments, the apparatus further includes: The training module is used to pre-train the initial reconstruction model with the third sensor signal collected by the device under power failure as a training sample, aiming to minimize the reconstruction error, to obtain an intermediate reconstruction model. The third sensor signal is a pure environmental noise sample. The fourth sensor signal collected by the device under normal operation with power and no faults is used as a training sample. With the goal of minimizing the reconstruction error, the parameters of the intermediate reconstruction model are adjusted to obtain the signal reconstruction model.
[0013] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the fault detection method in the embodiments of this application.
[0014] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to implement the fault detection method as described in the embodiments of this application.
[0015] On the other hand, a computer program product is provided, including a computer program stored in a computer-readable storage medium, a processor of a computer device reading the computer program from the computer-readable storage medium, and the processor executing the computer program, causing the computer device to perform the fault detection method provided in the above-described aspects or various alternative implementations of the above-described aspects.
[0016] This application provides a fault detection method that uses a signal reconstruction model to perform unsupervised reconstruction of sensor signals in operation and calculates the reconstruction error. This method can sensitively capture any abnormal signals that deviate from the normal pattern, avoiding missed or false alarms caused by changes in the noise environment. Furthermore, feature extraction and judgment based on fault physical characteristics are triggered only when the reconstruction error exceeds a threshold. This retains the high sensitivity of deep learning models to complex anomalies while eliminating non-fault interference false alarms common in purely data-driven methods through physical mechanism constraints. This dual mechanism of "anomaly perception + physical verification" significantly improves the accuracy and reliability of fault detection. At the same time, since subsequent calculations are only performed when an anomaly is suspected, the overall detection efficiency is greatly improved compared to methods that run complex rules or models throughout the process. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the implementation environment of a fault detection method provided in an embodiment of this application; Figure 2 This is a flowchart of a fault detection method provided according to an embodiment of this application; Figure 3 This is a flowchart of another fault detection method provided according to an embodiment of this application; Figure 4 This is a block diagram of a fault detection device provided according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a computer device provided according to an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0020] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.
[0021] In this application, the term "at least one" means one or more, and "multiple" means two or more.
[0022] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sensing signals and fault physical characteristics involved in this application were obtained with full authorization.
[0023] The fault detection method provided in this application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server. The following section first uses a server as an example to describe the implementation environment of the fault detection method provided in this application. Figure 1 This is a schematic diagram illustrating the implementation environment of a fault detection method according to an embodiment of this application. See also... Figure 1 The implementation environment includes terminal 101 and server 102. Terminal 101 and server 102 can be connected directly or indirectly via wired or wireless communication, which is not limited herein.
[0024] In some embodiments, terminal 101 is a data acquisition device, including but not limited to: an industrial computer, a programmable logic controller, a data acquisition card, an edge computing gateway, a smart sensor node, or an embedded monitoring device. Terminal 101 has an application or firmware installed and running for acquiring sensor signals and performing preliminary processing. This application can be any of a real-time data acquisition program, device status monitoring software, edge computing middleware, or an industrial IoT platform client. Indicatively, terminal 101 is deployed near equipment (such as high-voltage electrical equipment, rotating machinery, etc.) and connected to one or more sensors (such as vibration sensors, ultrasonic sensors, ultra-high frequency sensors, high-frequency current transformers, etc.) via wired or wireless means. Terminal 101 is used to continuously or trigger-based acquire sensor signals during device operation and perform preprocessing operations such as analog-to-digital conversion, filtering, amplification, and slicing on the acquired signals. In some embodiments, terminal 101 is also used to load and run a lightweight signal reconstruction model, calculate the reconstruction error, and trigger subsequent processing when the reconstruction error exceeds a threshold; if the reconstruction error does not exceed the threshold, the device is determined to be fault-free, and there is no need to upload data to the server, thereby saving communication bandwidth and cloud computing resources.
[0025] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more. This application does not limit the number of terminals or the type of device.
[0026] In some embodiments, server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), big data, and artificial intelligence platforms. Server 102 is used to provide background services for applications that support fault detection. In some embodiments, server 102 undertakes the main computational work, including: receiving the first sensor signal segment or extracted key features uploaded by terminal 101 when the reconstruction error exceeds a threshold; performing tasks such as feature extraction (if the terminal has not completed), fault physical feature matching, large language model invocation, and diagnostic report generation; and returning the diagnostic results to terminal 101 or pushing them to the user terminal. Terminal 101 undertakes secondary computational work, such as only performing signal acquisition, preprocessing, and reconstruction error calculation. Alternatively, server 102 undertakes secondary computational work, and terminal 101 undertakes the main computational work, such as the terminal completing feature extraction and feature matching, and only uploading the detection results to the server for recording. Alternatively, server 102 and terminal 101 can use a distributed computing architecture for collaborative computing, such as distributing the signal reconstruction model across the terminal and server for joint inference.
[0027] Figure 2 This is a flowchart of a fault detection method provided according to an embodiment of this application. See also... Figure 2 In this embodiment, the method is described using an example of execution by a server. The fault detection method includes the following steps: 201. The server reconstructs the first sensor signal collected by the device in operation based on the signal reconstruction model to obtain the second sensor signal. The signal reconstruction model is trained with the sensor signal of the device when there is no fault as the training sample and the goal of minimizing the reconstruction error.
[0028] In this embodiment, the signal reconstruction model is a pre-trained neural network model whose function is to reconstruct the input sensor signal. This model uses only the sensor signals from the device under fault-free conditions as training samples, and performs unsupervised training with the goal of minimizing reconstruction error, thereby learning and memorizing the signal patterns during normal device operation. The first sensor signal refers to the original time-series signal collected in real time by sensors (such as vibration sensors, ultrasonic sensors, UHF sensors, or high-frequency current transformers) during the actual operation of the device. It can be a vibration signal, ultrasonic signal, UHF signal, or high-frequency current signal, etc., and this embodiment does not limit this. After the server obtains the first sensor signal, it inputs it into the signal reconstruction model. The model performs encoding (feature extraction) and decoding (signal reconstruction) processing, and outputs a reconstructed signal of the same length as the input, i.e., the second sensor signal.
[0029] For example, for a high-voltage electrical device, the first sensing signal is an ultrasonic signal containing a partial discharge pulse. Since the signal reconstruction model has not learned the waveform of the discharge pulse, the output second sensing signal will not be able to accurately reconstruct the pulse part, thus producing a large reconstruction error in subsequent steps.
[0030] 202. The server determines the reconstruction error of the signal reconstruction model based on the first sensing signal and the second sensing signal. The reconstruction error is used to indicate the difference between the first sensing signal and the second sensing signal.
[0031] In this embodiment, reconstruction error is a quantitative indicator measuring the difference between the first and second sensing signals. Its magnitude reflects the degree to which the input signal deviates from the "normal pattern" learned by the model. The reconstruction error can be calculated using methods such as Mean Square Error (MSE), Mean Absolute Error (MAE), or Peak Error, and this embodiment does not impose any limitations on this. The server compares the first and second sensing signals point by point, calculates the difference at each sampling point or time window, and summarizes the results into a single value.
[0032] For example, for a segment with only background noise, the model reconstruction is effective, with a mean square error (MSE) of only 0.01. However, for anomalous segments containing partial discharge pulses or mechanical impacts, the MSE may rise to over 0.5 because the model cannot reconstruct the anomalous waveform. The value of the reconstruction error will serve as a key basis for determining whether to trigger in-depth analysis in subsequent steps.
[0033] 203. When the reconstruction error is greater than the error threshold, the server performs feature extraction on the first sensing signal to obtain key features, which are the physical characteristics of the device's sensing signal that are related to the fault.
[0034] In this embodiment, the error threshold is a preset numerical limit used to determine whether the reconstruction error is significant. This error threshold can be determined based on the historical reconstruction error distribution of the device under fault-free conditions. For example, the mean of the historical reconstruction errors can be added to three standard deviations to determine the error threshold, or a preset quantile can be used to determine the error threshold from the historical reconstruction errors. The server compares the reconstruction error calculated in step 202 with this error threshold. If the reconstruction error is greater than the error threshold, it is determined that the current signal segment has a suspected anomaly, thereby triggering the feature extraction process.
[0035] Feature extraction refers to the calculation of quantitative indicators, i.e. key features, from the original first sensing signal that can reflect the physical characteristics of equipment failure. For electrical equipment (such as transformers and cables), key characteristics may include, but are not limited to, kurtosis (characterizing the peaks in the signal amplitude distribution), spectral centroid (characterizing the centroid frequency of the spectrum), amplitude variation coefficient (characterizing the dispersion of the amplitude), RMS value (characterizing the overall energy level), mean (characterizing the DC component or static offset of the signal), peak value (characterizing the maximum instantaneous impulse intensity of the signal), peak-to-peak value (characterizing the dynamic range of the signal), variance (characterizing the fluctuation of the signal), standard deviation (characterizing the fluctuation of the signal), skewness (characterizing the asymmetry of the signal amplitude distribution), waveform factor (the ratio of the RMS value to the absolute mean, reflecting the waveform shape of the signal), peak factor (the ratio of the peak value to the RMS value, reflecting the impulse characteristics of the signal), impulse factor (the ratio of the peak value to the absolute mean, sensitive to sparse large-amplitude impulses), margin factor (the ratio of the peak value to the root square amplitude, often used to detect mechanical looseness or impact faults), mean square frequency (the second moment of the spectrum, reflecting the concentration of the signal frequency distribution), and frequency variance (the dispersion of the spectrum relative to the spectral centroid; a larger frequency variance indicates a higher signal frequency). Key features include: (1) more dispersed components; (2) fast spectral kurtosis map (a two-dimensional spectrum obtained by the fast spectral kurtosis algorithm, used to identify the most impactful frequency bands in non-stationary signals); (3) spectral kurtosis index (statistics extracted based on the fast spectral kurtosis map, quantifying the transient impact intensity of the signal); (4) time-frequency matrix energy density (energy values of each element in the time-frequency matrix obtained after performing short-time Fourier transform or wavelet transform on the signal, reflecting the energy distribution of the signal on the time-frequency plane); (5) time-frequency matrix coefficient of variation (the ratio of the standard deviation to the mean of the energy values in the time-frequency matrix, characterizing the uniformity of energy distribution); (6) envelope spectral entropy (calculating the entropy value of the envelope spectrum, quantifying the complexity of the envelope spectrum, the entropy value of the envelope spectrum decreases under fault conditions); (7) approximate entropy (an index that measures the complexity and irregularity of a signal sequence, fault signals usually have lower approximate entropy); (8) sample entropy (an improved version of approximate entropy, used to evaluate the self-similarity and regularity of the signal); for rotating machinery (such as motors and synchronous condensers), key features may also include, but are not limited to, peak factor, envelope spectrum fault characteristic frequency, rotational frequency amplitude, and sideband energy, etc., which are not limited in this application embodiment.
[0036] For example, when the reconstruction error of a signal exceeds the error threshold, the server calculates the kurtosis value of the signal. If the kurtosis is much higher than the theoretical value of the background noise (such as 3), it further confirms the existence of abnormal pulses and triggers the subsequent feature extraction process.
[0037] 204. The server determines the fault detection results of the device based on key features and physical fault features. The physical fault features are used to indicate the physical characteristics of the sensing signals when the device fails.
[0038] In this embodiment, the fault physical features are a pre-established set of knowledge rules used to describe the physical characteristics that the sensing signals should exhibit when different types of faults occur. Each fault physical feature may include constraints on one or more key features (such as threshold ranges, logical combinations, or trends), which are not limited in this embodiment. The server compares the values of the key features extracted in step 203 with the fault physical features in the knowledge base one by one. If at least one fault physical feature is matched, the device is determined to be faulty, and the fault type corresponding to that physical feature is output. If none of the fault physical features are matched, the device is still determined to be fault-free, even if the reconstruction error is large (possibly caused by unknown interference or non-faulty anomalies). This step, by introducing physical mechanism constraints, effectively eliminates common false alarms in purely data-driven methods, improving the reliability of the diagnostic results.
[0039] For example, the server extracts the first sensor signal with a kurtosis of 12.5 and a spectral centroid of 22kHz, which matches the fault physical characteristics "kurtosis > 5 and spectral centroid > 18kHz", and thus outputs the fault detection result as "point discharge".
[0040] This application provides a fault detection method that uses a signal reconstruction model to perform unsupervised reconstruction of sensor signals in operation and calculates the reconstruction error. This method can sensitively capture any abnormal signals that deviate from the normal pattern, avoiding missed or false alarms caused by changes in the noise environment. Furthermore, feature extraction and judgment based on fault physical characteristics are triggered only when the reconstruction error exceeds a threshold. This retains the high sensitivity of deep learning models to complex anomalies while eliminating non-fault interference false alarms common in purely data-driven methods through physical mechanism constraints. This dual mechanism of "anomaly perception + physical verification" significantly improves the accuracy and reliability of fault detection. At the same time, since subsequent calculations are only performed when an anomaly is suspected, the overall detection efficiency is greatly improved compared to methods that run complex rules or models throughout the process.
[0041] The above Figure 2 The diagram shown is merely the basic flow of this application. The following section will further elaborate on the solution provided in this application based on a specific implementation method. Figure 3This is a flowchart of another fault detection method provided according to an embodiment of this application, see [link / reference]. Figure 3 In this embodiment, the method is described using an example of execution by a server. The fault detection method includes the following steps: 301. The server reconstructs the first sensor signal collected by the device in operation based on the signal reconstruction model to obtain the second sensor signal. The signal reconstruction model is trained with the sensor signal of the device when there is no fault as the training sample and the goal of minimizing the reconstruction error.
[0042] In this embodiment, the server first acquires a first sensing signal collected by the device during operation. This signal is time-series waveform data, reflecting the real-time operating status of the device. The server processes this signal using a pre-trained signal reconstruction model. The core capability of the model is to "remember" the waveform pattern of the sensing signal during normal device operation and attempt to reconstruct the input signal according to the remembered normal pattern, outputting a second sensing signal. If the input signal contains abnormal waveforms that the model has not learned (such as partial discharge pulses or mechanical impacts), the model will be unable to accurately reconstruct these abnormal parts, resulting in a large reconstruction error in subsequent steps.
[0043] Since the first sensing signal may contain DC bias, environmental noise, power frequency interference, and amplitude differences with different dimensions, the server performs a series of preprocessing operations on the first sensing signal to improve the processing effect and numerical stability of the signal reconstruction model. In some embodiments, preprocessing may include operations such as mean removal, bandpass filtering, amplitude normalization, and full slicing.
[0044] Specifically, firstly, the server performs mean-reduction processing on the first sensor signal, calculating the arithmetic mean of the entire signal sequence and subtracting this mean from each sampling point to eliminate the DC bias of the sensor itself or the slowly varying components of the environment, thus making the signal meanless. Secondly, bandpass filtering is performed, setting the filter passband range according to the typical frequency band of the equipment fault signal. For example, for ultrasonic detection of partial discharge in electrical equipment, the bandpass filter is usually set to 20kHz~80kHz; for vibration signals of rotating machinery, the filter range (e.g., 10Hz~2kHz) can be set according to the equipment speed and fault characteristic frequency. Bandpass filtering can effectively eliminate power frequency interference, low-frequency vibration, or high-frequency electronic noise unrelated to the fault, highlighting the target fault characteristics. Then, amplitude normalization is performed, scaling the signal amplitude to a uniform range. For example, Min-Max normalization is used to linearly map the signal to the [-1, 1] interval, or Z-score normalization is used to make the signal mean 0 and variance 1. Normalization can eliminate amplitude differences between different sensors, different gains, or different acquisition times, allowing the model to focus on waveform morphology rather than absolute amplitude. Finally, full slicing is performed: Since signal reconstruction models typically require inputs of a fixed length (e.g., 1024 or 2048 sampling points), the server cuts the preprocessed long sequence signal into multiple short segments of equal length according to a preset window length and step size (the window length is fixed, and the step size can be equal to the window length to achieve non-overlapping slicing, or smaller than the window length to achieve overlapping slicing). Each segment serves as an independent sample for subsequent reconstruction error calculation and feature extraction. Full slicing ensures the standardization of the model input while increasing the temporal resolution of anomaly detection.
[0045] After preprocessing, the server inputs each slice (i.e., the preprocessed first sensor signal segment) into the signal reconstruction model. The signal reconstruction model is trained using the sensor signals of the device when it is fault-free as training samples, with the goal of minimizing reconstruction error. The model encodes (feature extraction) and decodes (signal reconstruction) the input signal, outputting a reconstructed signal of the same length as the input, i.e., the second sensor signal. If the input signal segment contains anomalous waveforms that the model has not learned (such as partial discharge pulses or mechanical shocks), the model will be unable to accurately reconstruct these anomalous parts, resulting in a large reconstruction error.
[0046] This application does not limit the architecture of the signal reconstruction model. The signal reconstruction model can be an autoencoder, a variational autoencoder (VAE), a denoising autoencoder (DAE), the encoder-decoder part of a generative adversarial network (GAN), or a reconstruction network based on a self-attention mechanism (such as a Transformer). As long as the model can compress, encode, decode, and reconstruct the input signal, and is trained in an unsupervised manner, it can be applied to this scheme.
[0047] In some embodiments, the signal reconstruction model is a one-dimensional convolutional denoising autoencoder (1D-CDAE). The 1D-CDAE includes an encoder and a decoder. The encoder extracts features from the input signal, and the decoder reconstructs the input signal based on the features output by the encoder. Accordingly, the server's signal reconstruction process includes: the server inputs a first sensing signal into the encoder, which performs convolution and pooling operations on a one-dimensional convolutional layer to extract a signal feature vector. Then, the server inputs the signal feature vector into the decoder, which performs upsampling and convolution operations on a one-dimensional transposed convolutional layer to reconstruct a second sensing signal. That is, after the server acquires the first sensing signal, it first inputs the signal into the encoder. The encoder contains multiple one-dimensional convolutional and pooling layers. The convolutional layers extract local temporal patterns of the signal (such as waveform rising edges, oscillation decay, etc.) through sliding convolution kernels, while the pooling layers reduce the feature dimension and expand the receptive field through downsampling. After layer-by-layer processing by the encoder, the original signal is compressed into a low-dimensional signal feature vector, which condenses the most essential morphological information of the signal. Subsequently, the server inputs this feature vector into the decoder. The decoder consists of multiple one-dimensional transposed convolutional layers (also known as deconvolutional layers) and upsampling layers. The transposed convolutional layers progressively restore the temporal resolution of the signal, while the upsampling layers enlarge the feature map to the length of the original signal. Finally, the decoder outputs the reconstructed signal, which is the second sensing signal.
[0048] For example, for an ultrasonic signal with a sampling rate of 1MHz, the encoder compresses 1024 sampling points into a 32-dimensional feature vector after three layers of convolutional pooling; the decoder then reconstructs the waveform back into 1024 points from this vector. Since the model only learns normal background signals without discharge during training, when the input signal contains discharge pulses, the reconstructed waveform will show obvious "collapse" or "smoothing" at the pulse position, thus achieving the detectability of anomalies.
[0049] The solution provided in this application specifically defines the signal reconstruction model as a one-dimensional convolutional denoising autoencoder. The encoder extracts local temporal features of the signal through one-dimensional convolution and pooling operations, while the decoder reconstructs the signal through one-dimensional transposed convolution and upsampling. This structure can effectively learn the spatiotemporal dependencies of the sensing signal, and the parameter sharing mechanism of the convolutional layers significantly reduces the number of model parameters, improving training and inference efficiency. Compared to fully connected networks, the one-dimensional convolutional structure is more suitable for processing one-dimensional time series, achieving better reconstruction accuracy with lower computational cost, thereby improving real-time performance while maintaining high detection accuracy.
[0050] The acquisition method of the sensing signals (samples) used in the signal reconstruction model training process in this application embodiment is not limited. In some embodiments, the sensing signals of the device when it is fault-free include at least one of the following. First, the sensing signal collected by the device in a normal operating state that is powered on and confirmed to be fault-free. That is, a continuous signal is directly collected as a training sample, which can ensure that the training sample is highly consistent with the subsequent detection conditions.
[0051] The second step involves automatically selecting and truncating sensor signal segments from the collected signals. These segments must reach a preset duration and contain no amplitude pulses exceeding a preset value. Specifically, an automatic selection algorithm is used to automatically extract signal segments from long-term monitoring data that have reached a preset duration and contain no pulses exceeding a preset amplitude threshold. These segments are then used as fault-free samples. The preset duration and amplitude pulse values can be set based on practical experience and requirements; no specific restrictions are imposed here. For example, for vibration signals from rotating machinery, the amplitude threshold can be set to three times the normal root mean square value to automatically select stable vibration segments without impact.
[0052] The solution provided in this application clarifies two reliable sources of fault-free sensing signals: signals collected when the device is powered on and confirmed to be fault-free, and pulse segments without large amplitude values automatically filtered from long-term signals. The former ensures that the training samples are consistent with the detection conditions, while the latter provides an automated sample acquisition method that does not require manual annotation. These two sources enable the low-cost, large-scale construction of high-quality training sets, avoiding the problem of the model misclassifying normal fluctuations as abnormal due to impure training samples, thereby improving the efficiency of model training and the accuracy of fault detection.
[0053] This application does not limit the training method of the signal reconstruction model. In some embodiments, the server directly uses the sensor signals collected by the device under normal operating conditions with power and no faults as training samples, aiming to minimize the reconstruction error, to train the initialized signal reconstruction model until the model converges. This training method is the most direct, as the training samples are completely consistent with the detection conditions, and the model can accurately learn the signal patterns when the device is operating normally. For example, for a newly commissioned wind turbine, vibration signals are continuously collected as training samples during the first month when its operating condition is stable and no bearing faults are manually confirmed, to train an autoencoder model. This model can then be used to detect abnormal vibrations of the wind turbine.
[0054] In other embodiments, to further improve the model's adaptability to real-world operating conditions, this application also provides a two-stage training method. In the first stage, the server uses the third sensor signal collected by the device under power-off conditions as training samples to pre-train the initial reconstruction model, aiming to minimize the reconstruction error, thus obtaining an intermediate reconstruction model. The third sensor signal is a pure environmental noise sample and does not contain any excitation response from the device itself. The signal under power-off conditions only contains basic noise such as thermal noise and environmental electromagnetic interference; the model learns the most basic noise patterns through pre-training. In the second stage, the server uses the fourth sensor signal collected by the device under normal, fault-free operation with power as training samples to adjust the parameters of the intermediate reconstruction model, aiming to minimize the reconstruction error, thus obtaining a signal reconstruction model. During parameter adjustment, the model learns the normal vibration and electromagnetic field changes of the device under power, thereby adapting to the real operating environment.
[0055] The solution provided in this application employs a two-stage training strategy: first, the model is pre-trained using pure environmental noise samples under power outage conditions to learn the basic noise pattern; then, it is fine-tuned using normal, fault-free samples under power conditions to adapt the model to the actual working environment. This strategy solves the problems that using only power outage samples may prevent the model from recognizing powered signals, and using only powered samples may introduce potential weak discharge hazards. Pre-training to obtain good initial parameters before fine-tuning significantly accelerates convergence, reduces the need for a large number of powered samples, thereby improving model training efficiency and ensuring the final model's detection accuracy under actual working conditions.
[0056] In other embodiments, for training the signal reconstruction model, the server can also employ online incremental learning: when the equipment operates for a long time and is consistently confirmed to be fault-free, newly acquired normal signals can be continuously used for fine-tuning and updating the model, enabling the model to adapt to slow changes in the equipment's operating conditions (such as normal bearing wear and seasonal changes in ambient temperature), avoiding false alarms due to outdated training samples. Another example is multi-task learning: in addition to the reconstruction task, the model is simultaneously trained to predict a specific operating parameter of the equipment (such as rotational speed or load) to assist in feature extraction. Specifically, during the model training phase, in addition to using minimizing the reconstruction error as the main task loss function, the server adds an auxiliary prediction task. This auxiliary task connects one or more fully connected layers above the feature vector output by the encoder, forming a prediction branch. It uses the actual operating parameters of the equipment (e.g., the current rotational speed or load rate read from the control system) as the prediction target and employs mean squared error or cross-entropy as the auxiliary loss function. The model's total loss function is a weighted sum of the main task loss and the auxiliary task loss, and the parameters of both the encoder and the prediction branch are updated simultaneously through backpropagation. The auxiliary task forces the encoder to extract features related to operating parameters from the input signal, enabling the model to distinguish between normal operating condition changes and genuine faults / anomalies. When equipment operating conditions change, the encoder can encode these changes as operating condition features rather than abnormal features, thus avoiding misjudging normal operating condition fluctuations as faults. During the inference phase, the server can discard the prediction branch and use only the encoder and decoder for signal reconstruction, or retain the prediction branch as an additional output for operating condition monitoring. Furthermore, model compression techniques (knowledge distillation, pruning, quantization) can be used to lightweight large models, enabling their deployment on edge computing terminals, reducing reliance on cloud servers and improving real-time performance.
[0057] 302. The server determines the reconstruction error of the signal reconstruction model based on the first sensing signal and the second sensing signal. The reconstruction error is used to indicate the difference between the first sensing signal and the second sensing signal.
[0058] In this embodiment, after the server completes signal reconstruction, it needs to quantify the difference between the first sensing signal (original signal) and the second sensing signal (reconstructed signal). This difference is called the reconstruction error, and its magnitude directly reflects the degree to which the input signal deviates from the "normal mode." The larger the reconstruction error, the more significant the abnormal components contained in the signal. The server compares the reconstruction error with a pre-set error threshold to determine whether to trigger the subsequent deep fault analysis process. If the reconstruction error is not greater than the error threshold, the server determines that the device is fault-free. If the reconstruction error is greater than the error threshold, the server extracts the key features of the first sensing signal and determines whether the device is faulty based on these key features, i.e., steps 303-304 are executed. Details can be found in subsequent content and will not be repeated here.
[0059] The solution provided in this application directly determines that the device is fault-free when the reconstruction error is no greater than a threshold, without performing subsequent computationally intensive steps such as feature extraction and rule matching. This early termination mechanism minimizes signal processing overhead during most normal periods, significantly improving the overall detection efficiency of the system. Especially in long-term continuous monitoring scenarios, where the device is in a healthy state most of the time, this method can save a significant amount of computing resources and energy consumption, while maintaining a sensitive response to abnormal signals, achieving a good balance between efficiency and accuracy.
[0060] The embodiments of this application do not impose any restrictions on the calculation method of the reconstruction error. For example, the reconstruction error can be calculated using mean square error, that is, the server calculates the square of the difference between the corresponding sampling points of the first sensor signal and the second sensor signal point by point, and then averages the square values of all sampling points. Alternatively, the reconstruction error can also be calculated using other forms such as mean absolute error (MAE) or peak error (i.e., maximum absolute difference).
[0061] For example, for a signal segment with 2048 sampling points, the server calculates the MSE as 0.0052, which is the reconstruction error of that segment. When the reconstruction error is no greater than the error threshold, the server directly determines that the device is fault-free and does not perform subsequent feature extraction, rule matching, and other steps. This early termination mechanism can significantly save computing resources because the device is in a healthy state most of the time. For example, in continuous 24-hour monitoring, the reconstruction error may exceed the threshold for less than 1% of the time period. Therefore, for 99% of the time period, only two lightweight steps, signal reconstruction and error comparison, need to be performed, resulting in extremely high detection efficiency.
[0062] The aforementioned error threshold can be a preset fixed value or dynamically determined based on the numerical distribution of historical reconstruction errors of the equipment under fault-free conditions. Specifically, the server can collect the reconstruction errors of all normal signal segments generated by the equipment during fault-free operation (e.g., the first month after commissioning), calculate the mean and standard deviation of these errors, and then set the error threshold as "mean + k times the standard deviation," where k is typically 3 (corresponding to a 99.7% confidence interval). The server can also use a quantile method, for example, taking the 99th quantile of the historical reconstruction error distribution as the threshold. This adaptive threshold mechanism can automatically adjust according to the equipment's own characteristics and the ambient noise environment, avoiding the problem of poor adaptability of fixed thresholds under different operating conditions. Compared with manual and repeated parameter tuning, this method significantly improves the efficiency and generalization ability of threshold setting, thereby maintaining high accuracy in anomaly detection across different equipment and different operating stages, reducing the risk of false alarms and missed alarms.
[0063] For example, a motor located in a noisy factory environment has a normal reconfiguration error with a mean of 0.01 and a standard deviation of 0.002. Taking k=3, the threshold is 0.016. On the other hand, another device located in a quiet laboratory has a mean of only 0.002, and the threshold is set to 0.008. Both can sensitively detect anomalies without generating a large number of false alarms due to differences in environmental noise.
[0064] 303. When the reconstruction error is greater than the error threshold, the server performs feature extraction on the first sensing signal to obtain key features, which are the physical characteristics of the device's sensing signal that are related to the fault.
[0065] In this embodiment, when the server determines that the reconstruction error is greater than the error threshold, it means that the current signal segment contains abnormal components that deviate from the normal pattern. At this time, the server does not immediately determine it as a fault, but instead proceeds to a deeper analysis: feature extraction is performed on the original first sensor signal, and several quantitative indicators that can reflect the physical nature of the device fault are calculated. These indicators are called key features. The design of key features is based on the unique characteristics of different fault types in the signal time domain, frequency domain, or time-frequency domain, such as the sharpness of the pulse, the centroid position of the spectrum, and the dispersion of energy. By extracting these features, the abstract anomalies in the original waveform can be transformed into numerical values with clear physical meaning, providing a basis for subsequent rule verification.
[0066] The key features extracted by the server from the first sensing signal may differ depending on the type of device. The following will take electrical equipment and rotating machinery as examples for detailed explanation, but it is by no means limited to these.
[0067] In some embodiments, the device is an electrical device, and accordingly, key features may include at least one of, but are not limited to, the kurtosis, spectral centroid, amplitude variation coefficient, and RMS value of the first sensing signal. Kurtosis is the ratio of the fourth central moment to the square of the variance, used to indicate the peaking of the amplitude distribution of the first sensing signal. The kurtosis of normally distributed noise is approximately 3, while the kurtosis of partial discharge pulses, due to their large amplitude and sparse occurrence, is typically much greater than 3 (reaching tens or even hundreds). The spectral centroid is the first moment of the signal spectrum, used to indicate the centroid frequency of the spectrum of the first sensing signal. The spectral centroid of normal background noise is typically concentrated in the low-frequency range (e.g., several kiloHz), while the spectral centroid of ultrasonic or ultra-high frequency signals generated by partial discharge will significantly shift upwards to above 16kHz or even higher. The amplitude variation coefficient is the ratio of the standard deviation to the mean, used to indicate the dispersion of the amplitude of the first sensing signal. The sparsity of partial discharge pulses leads to a significant increase in the variation coefficient. The RMS value is the root mean square value of the signal, used to indicate the overall energy level of the first sensing signal.
[0068] The solution provided in this application, targeting electrical equipment, uses kurtosis, spectral centroid, amplitude variation coefficient, and RMS value as key features. Kurtosis can sensitively reflect the peak characteristics of partial discharge pulses, the upward shift of the spectral centroid can effectively distinguish between high-frequency discharges and low-frequency interference, the amplitude variation coefficient characterizes the sparsity and dispersion of the pulse, and the RMS value provides an overall energy benchmark. By combining these features, the waveform essence of electrical faults can be comprehensively characterized from different physical dimensions. Compared with single threshold judgment, this significantly improves the accuracy of identifying typical electrical defects such as partial discharges, while avoiding false triggering caused by environmental noise fluctuations and improving detection efficiency.
[0069] For example, the server detects an ultrasonic signal and calculates a kurtosis of 15.2, a spectral centroid of 22.5 kHz, an amplitude variation coefficient of 1.8, and an effective value of 0.35 mV (the normal baseline is 0.12 mV). These characteristic values all point to the possible presence of tip discharge or free particle discharge.
[0070] In some embodiments, the first sensing signal of the electrical equipment is acquired. For the first sensing signal Pre-processing steps such as mean removal, bandpass filtering, normalization, and full slicing are performed to obtain a signal sequence containing multiple signal segments. The server can obtain the key features of electrical equipment using the following formulas (1)-(4).
[0071] (1) (2) (3) (4) in, The first sensing signal The first in A signal segment, The number of signal segments obtained after slicing the first sensing signal; The effective value of the first sensing signal; The kurtosis of the first sensing signal; The average amplitude of the first sensing signal; The standard deviation of the amplitude of the first sensing signal. ; The amplitude variation coefficient of the first sensing signal; The spectral centroid of the first sensing signal; After the signal undergoes a Fourier transform, at the th... Frequency values at each frequency point; For the first The spectral amplitude of the frequency value at each frequency point.
[0072] In other embodiments, the device is rotating machinery. Accordingly, key features may include, but are not limited to, at least one of the following: kurtosis, peak factor, envelope spectrum fault characteristic frequency, rotational frequency amplitude, and sideband energy of the first sensing signal. Kurtosis has the same meaning as in electrical equipment and is used to capture periodic impacts caused by early bearing failures. The peak factor is the ratio of the signal peak value to the effective value, used to indicate the impact characteristics of the first sensing signal. The peak factor of a normal vibration signal is approximately 3-5; when wear or loosening is present, the impact pulse can raise the peak factor to over 10. The envelope spectrum fault characteristic frequency refers to the characteristic frequency appearing in the envelope spectrum after Hilbert envelope demodulation of the signal, corresponding to faults in the bearing's inner ring, outer ring, rolling elements, or cage. It is used to indicate the energy concentration of specific frequency components in the spectrum of the first sensing signal after envelope demodulation. This energy concentration directly indicates the damage state of the corresponding component. The rotational frequency amplitude is the magnitude of the rotational frequency (i.e., the rotor rotational frequency) and its harmonics in the signal spectrum, used to assess the severity of rotor imbalance or misalignment. When unbalanced, the amplitude at 1x RPM increases significantly; when misaligned, the amplitude at 2x RPM increases relatively. Sideband energy refers to the sum of the energy in the sidebands on both sides of the carrier frequency (such as the gear meshing frequency) in the spectrum, used to indicate the degree of energy concentration in the sidebands on both sides of the carrier frequency in the spectrum of the first sensing signal. Sideband energy can be used to diagnose gear faults or modulation phenomena; that is, when a local fault occurs in the gear, the sideband energy will increase significantly.
[0073] In some embodiments, the first sensing signal of the rotating machinery is acquired. For the first sensing signal Pre-processing steps such as mean removal, bandpass filtering, normalization, and full slicing are performed to obtain a signal sequence containing multiple signal segments. The server can obtain the key features of the rotating machinery using the following formulas (5)-(12).
[0074] (5) (6) (7) (8) (9) (10) (11) (12) in, The first sensing signal The first in A signal segment, The number of signal segments obtained after slicing the first sensing signal; The first sensing signal The mean of the amplitude; The first sensing signal The steepness. The first sensing signal The maximum absolute value (i.e., peak value) in the time domain; The first sensing signal The effective value of is calculated using a formula similar to the aforementioned formula (1), and will not be repeated here. The number of rolling elements in rotating machinery; The rotational frequency of rotating machinery (i.e., the rotor rotational frequency, measured in Hz). The diameter of the rolling element in rotating machinery; The bearing pitch diameter for rotating machinery; This refers to the bearing contact angle; The characteristic frequency of bearing inner ring failure; This refers to the characteristic frequency of bearing outer ring failure. The characteristic frequency of bearing rolling element failure; These are the characteristic frequencies of bearing cage failure. The server can use the amplitude or energy at each of these characteristic frequencies in the envelope spectrum as key features to indicate the damage state of the corresponding components. , is the first frequency conversion Subharmonic frequency ( ), For the signal spectrum at frequency The amplitude at that point; For the signal spectrum at frequency The absolute value of the amplitude at that point; This is the frequency conversion amplitude. The server can extract 1 times the frequency conversion amplitude. and twice the frequency amplitude As a key feature, it is used to evaluate rotor imbalance ( Significantly increased) and misalignment ( The degree of relative increase. This refers to the set of frequencies within the sidebands on either side of the carrier frequency (such as the gear meshing frequency). Typically, the frequencies of the sidebands are... , The gear meshing frequency, ; This refers to the sideband energy. Sideband energy is used to quantify the intensity of gear faults or modulation phenomena; when a local fault occurs in the gear, the sideband energy will be significantly enhanced.
[0075] The solution provided in this application, targeting rotating machinery, employs kurtosis, peak factor, envelope spectrum fault characteristic frequency, rotational frequency amplitude, and sideband energy as key features. Kurtosis and peak factor are used to capture the periodic impacts generated by early bearing faults; envelope spectrum fault characteristic frequencies can directly locate specific faulty components such as the inner and outer rings of the bearing; rotational frequency amplitude reflects the degree of rotor imbalance or misalignment; and sideband energy reveals gear modulation faults. These features characterize the physical nature of mechanical faults from multiple perspectives—time domain, frequency domain, and envelope domain—significantly improving the accuracy of fault type identification compared to methods relying solely on vibration amplitude. Furthermore, the feature extraction computation is low, enabling efficient online diagnosis.
[0076] For example, the server analyzes a motor vibration signal and calculates that the kurtosis is 3.6 and the peak factor is 6.2. The envelope spectrum shows obvious BPFO (Ball Pass Frequency on Outer race) and its second harmonic, so it is determined to be an early failure of the bearing outer race.
[0077] 304. The server generates fault prompt words based on key features, inputs the fault prompt words into the large language model, and compares the key features with the physical features of the fault through the large language model to obtain the equipment fault detection report.
[0078] In this embodiment, after extracting key features, the server needs to compare these values with pre-established knowledge rules to determine whether the device has a fault and the specific type of fault. These knowledge rules are called fault physical features, which describe the physical characteristics that the sensing signal should exhibit when different types of faults occur (e.g., how much kurtosis should exceed, which frequency band the spectral centroid should be in, etc.). Through rule matching, the server transforms the abnormal signals obtained by data-driven processing into physically interpretable fault diagnosis conclusions, effectively avoiding the problem of false alarms caused by pure neural network models learning non-faulty interference.
[0079] The server can determine whether a device is faulty and the specific type of fault using a large language model. Specifically, based on the key features and device information extracted in step 303, the server generates a fault prompt. The device information indicates the basic attributes of the device and may include, but is not limited to, device model, device type, and device name. The server then inputs the fault prompt into the large language model. The large language model identifies key features based on the fault prompt and compares these key features with fault physical features stored in the knowledge base. If the key features match the fault physical features, the device is determined to be faulty; otherwise, the device is determined not to be faulty.
[0080] The large language model, based on its knowledge base of fault diagnosis knowledge and common-sense reasoning ability, can also generate a structured fault detection report. This report includes at least one of the following: the type of equipment fault, the severity of the fault, and a fault repair plan.
[0081] For example, a fault message might read: "An audio signal has been detected on a 126kV GIS (Gas Insulated Switchgear). Key characteristics are as follows: kurtosis = 12.5, spectral centroid = 22.3kHz, amplitude coefficient of variation = 1.9, RMS value = 0.35mV (normal baseline 0.12mV). Based on the above information, please provide a fault severity assessment, possible root cause analysis, and maintenance recommendations." Upon receiving this message, the large language model will output a fluent diagnostic report, including at least the fault type (e.g., "point discharge"), fault severity information (e.g., "early stage, maintenance recommended within 3 months"), and fault repair plan (e.g., "inspect and polish burrs or protrusions on the high-voltage conductor"). For rotating machinery, the message could be similarly constructed, for example: "Motor vibration signal: kurtosis = 3.6, envelope spectrum shows BPFO = 700Hz and its second harmonic, please diagnose." The large language model would then output suggestions such as "early wear on the bearing outer ring, lubrication or bearing replacement recommended."
[0082] In some embodiments, the physical characteristics of the fault include the physical characteristics of the sensing signals under multiple fault types. Accordingly, the process by which the server determines the fault detection result includes: if the large language model detects that a key feature matches a physical feature corresponding to any fault type, the server determines that the device is faulty and outputs the matched fault type. If the large language model detects that a key feature does not match a physical feature corresponding to any fault type, the server determines that the device is not faulty.
[0083] The physical characteristics of the fault are pre-stored in a knowledge base in the form of structured text. Each record corresponds to a fault type and includes the threshold conditions or logical combinations that the key characteristics under that fault type should meet. For example, for electrical equipment, the following rules can be defined: if the kurtosis is >5 and the spectral centroid is >18kHz, it is determined to be a tip discharge; if the kurtosis is >8, the amplitude variation coefficient is >1.5, and the effective value rise is >20%, it is determined to be a free metal particle discharge; if the spectral centroid is between 10kHz and 15kHz and the kurtosis is between 3 and 5, it may be a surface discharge. For rotating machinery, the following can be defined: if BPFI and its harmonics appear in the envelope spectrum and the kurtosis is >3.5, it is determined to be a bearing inner ring fault; if the amplitude of 1 times the rotational frequency exceeds 3 times the baseline and the amplitude of 2 times the rotational frequency is less than 0.3 times the amplitude of 1 times the rotational frequency, it is determined to be a rotor imbalance. The server compares the key characteristic values calculated in step 303 with each rule in the knowledge base in turn. When all conditions are met, it is considered that the rule has been "hit". If any rule is matched, the server determines that the device is faulty and outputs the fault type corresponding to that rule; if none of the rules are matched, the server still determines that the device is not faulty, even if the reconstruction error is large (which may be caused by unknown interference, sudden environmental changes or new faults), and marks the signal segment as "unknown anomaly" for manual review.
[0084] The solution provided in this application achieves rapid and clear fault determination by pre-associating the physical characteristics of a fault with a specific fault type and employing a "hit-and-output" matching rule. When a key feature satisfies the physical characteristics corresponding to any fault type, the fault type is directly output; otherwise, no fault is determined. This rule-based verification method avoids the computational overhead and ambiguity caused by complex probability models. Furthermore, because the rules are based on explicit physical mechanisms, the determination results have high interpretability and reliability. Compared to classifiers that require a large number of labeled samples, this method only requires pre-defined rules, significantly reducing deployment costs and improving detection efficiency.
[0085] In some embodiments, in addition to the "hit-and-output" matching rule described above, the server can also employ a weighted scoring method for judgment. Specifically, the server calculates a matching score for the degree of matching between each key feature and the corresponding threshold in the physical features of the fault (e.g., the larger the feature value exceeds the threshold by a multiple, the higher the score; if it is below the threshold, the score decreases proportionally). Then, the server performs a weighted summation of the matching scores of all features, with the weights pre-set according to the contribution of the feature to the specific fault type, to obtain a comprehensive confidence score. When the comprehensive confidence score exceeds a preset scoring threshold, it is determined to be the fault type. Here, the scoring threshold is a preset numerical limit used to determine whether the comprehensive confidence score meets the minimum requirement for determining a fault. This scoring threshold can be determined based on historical data or expert experience. For example, it can be the mean of the comprehensive scores of historical normal samples plus several times the standard deviation (e.g., mean + 3 times the standard deviation), or the maximum normal score on the training set can be used as the threshold to ensure that the probability of misjudging normal samples as faults is controlled within an acceptable range at a certain confidence level. "Normal samples" refer to sensor signal segments collected from devices confirmed to be fault-free. This method can handle cases where some features meet the threshold but not all of them, and is particularly suitable for scenarios where early signs of faults are vague and feature values are in the critical region. For example, when the kurtosis is 4.8 (threshold 5) and the spectral centroid is 19kHz (threshold 18kHz), the weighted scoring method may still give 60 points, exceeding the 50-point threshold, thus providing an early warning of potential tip discharge risk.
[0086] In other embodiments, the server can also employ fuzzy logic for judgment. Specifically, the server fuzzifies the threshold boundaries of key features, using membership functions (such as trapezoidal or triangular membership functions) to represent the degree to which a feature value belongs to a certain fault mode. For example, a kurtosis of 4.8 can belong to "normal" with a membership degree of 0.6 and to "point discharge" with a membership degree of 0.4. Then, the server synthesizes the membership degrees of multiple features through fuzzy inference, outputs a membership degree vector for each fault type, and takes the fault type corresponding to the maximum value as the final result. This method is particularly suitable for scenarios with a large amount of uncertain interference and unclear signal feature boundaries, and can provide smoother diagnostic results.
[0087] In other embodiments, the server can also employ Bayesian inference for judgment. Specifically, the server pre-calculates the probability distribution of key features for each fault type (or learns it using historical data). After obtaining real-time feature values, the server calculates the posterior probability that the feature combination belongs to each fault type, selects the fault type with the highest posterior probability as the result, and outputs this probability as the confidence level. This method requires a certain amount of labeled data, but once established, it has a high diagnostic accuracy and can naturally output probability values for subsequent multi-sensor fusion.
[0088] In other embodiments, the server can also organize the fault physical rules into a decision tree or rule chain to achieve step-by-step judgment. A decision tree is a tree-structured decision model consisting of internal nodes, branches, and leaf nodes. Each internal node represents a judgment condition for a key feature (e.g., "whether the kurtosis is greater than 5"), each branch represents the output of the judgment result (e.g., "yes" or "no"), and each leaf node represents the final decision result (e.g., "normal", "point discharge", "free particle discharge", etc.). The construction of a decision tree can be manually defined based on expert experience or automatically learned from historical data. During inference, the server starts from the root node and traverses downwards along the corresponding branches according to the value of the key feature until it reaches a leaf node. The fault type marked by that leaf node is the detection result. A rule chain (also called a decision rule set or ordered rule list) is an equivalent expression of a decision tree, consisting of a series of "IF-THEN" rules ordered by priority. Each rule's condition part is a logical combination of multiple key characteristic conditions (e.g., "IF kurtosis > 5 AND spectral centroid > 18kHz THEN tip discharge"), and the conclusion part is the corresponding fault type or normal state. The rule chain can be ordered based on rule priority (e.g., based on fault severity, frequency of occurrence, or rule coverage). The server matches rules sequentially from highest to lowest priority. Once a rule's condition is met, its conclusion is immediately output, and subsequent rules are not matched. For example, for partial discharge diagnosis of electrical equipment, the following rule chain can be constructed: Rule 1 (highest priority): IF kurtosis > 8 AND spectral centroid > 20kHz AND amplitude variation coefficient > 2.0 THEN Free particle discharge; Rule 2: IF kurtosis > 5 AND spectral centroid > 18kHz AND amplitude variation coefficient > 1.2 THEN Tip discharge; Rule 3: IF spectral centroid > 10kHz AND spectral centroid ≤ 18kHz AND kurtosis > 3 THEN Surface discharge; Rule 4 (lowest priority): ELSE Normal.
[0089] The formal organization of decision trees and rule chains makes the fault diagnosis process clear and interpretable. During reasoning, it is only necessary to follow a path or match several rules in sequence. This approach can greatly speed up the reasoning process and avoid traversing irrelevant rules. It is especially suitable for edge computing terminals with limited computing resources.
[0090] Based on the identified fault type, the server can further classify the fault severity into several levels, such as warning, attention, alarm, and emergency, according to the degree to which key features exceed thresholds (e.g., kurtosis is 2 or 3 times the threshold) and the trend of feature changes (e.g., whether the effective value is continuously rising or remaining stable at a high level). For example, for free particle discharge, when the effective value increases by 30% and the kurtosis is between 5 and 10, it is classified as "attention"; when the effective value increases by more than 100% and the kurtosis is greater than 15, it is classified as "emergency." Different maintenance suggestions are output for different levels.
[0091] In some embodiments, under the same major category of fault types, servers can be further subdivided into subtypes by combining multiple key features. For example, for electrical equipment, tip discharge can be further divided into subcategories such as high-voltage conductor tip, grounding electrode tip, and floating potential. These can be further distinguished by the specific numerical range of the spectral centroid (e.g., 20kHz~30kHz vs. 30kHz~40kHz) and the difference in amplitude variation coefficient. For example, high-voltage conductor tip discharge: the spectral centroid is usually located in the 20kHz~30kHz range, with a moderate amplitude variation coefficient (1.2~1.8). This is because the discharge pulse generated by the high-voltage conductor tip is steeper, with rich high-frequency components, but limited by high-frequency attenuation in the propagation path, preventing the spectral centroid from becoming too high; at the same time, the discharge pulse has a certain regularity, and the variation coefficient is relatively stable. Grounding electrode tip discharge: the spectral centroid is usually located in the 30kHz~40kHz range or even higher, with a higher amplitude variation coefficient (>1.8). Grounding electrode tips are often located near the equipment casing, with a short propagation distance, and the high-frequency components are preserved more completely, thus the spectral centroid shifts upward more significantly; moreover, grounding electrode discharge is affected by the grounding state, resulting in larger pulse amplitude fluctuations and a significantly higher variation coefficient. Floating potential discharge: The spectral centroid is relatively low (10kHz~20kHz), and the amplitude variation coefficient is moderate or low (1.0~1.5). Floating potential discharge is usually caused by poor contact of metal components. The discharge energy is relatively large, but the pulse rise edge is relatively flat, and there are fewer high-frequency components. At the same time, the discharge mode is relatively stable, and the variation coefficient is not too large.
[0092] In the specific judgment, the server first determines whether it belongs to the category of tip discharge based on characteristics such as kurtosis and spectral centroid. Then, it further judges the range of the spectral centroid and the range of amplitude variation coefficient: if the spectral centroid falls between 20kHz and 30kHz and the variation coefficient is between 1.2 and 1.8, it is judged as high-voltage conductor tip discharge; if the spectral centroid is greater than 30kHz and the variation coefficient is greater than 1.8, it is judged as grounding electrode tip discharge; if the spectral centroid is between 10kHz and 20kHz and the variation coefficient is between 1.0 and 1.5, it is judged as floating potential discharge.
[0093] For rotating machinery, bearing failures can be subdivided into inner ring, outer ring, rolling element, and cage failures. Each type of failure corresponds to a different envelope spectrum characteristic frequency, and the trend of frequency amplitude variation also differs. The server can distinguish the specific faulty component by extracting the energy peaks at these characteristic frequencies in the envelope spectrum and combining this with the degree of matching between the energy and the theoretical frequency. The specific subdivision method is as follows: Bearing inner ring fault: The envelope spectrum will show BPFI (bearing inner ring fault characteristic frequency) and its harmonics, with significant energy peaks at the characteristic frequency. Because the inner ring rotates with the shaft, the relative position of the fault point and the load area changes periodically, and BPFI is usually accompanied by rotational frequency sideband modulation (i.e., , (This refers to the rotational frequency of the rotating machinery), and the amplitude of the rotational frequency may be slightly increased. When the server detects a significant peak at BPFI in the envelope spectrum, and sidebands on both sides are separated by the rotational frequency, it determines that there is a fault in the inner ring of the bearing.
[0094] Bearing outer ring fault: The envelope spectrum will show the BPFO (bearing outer ring fault characteristic frequency) and its harmonics. Since the outer ring is stationary, the relative position of the fault point and the load area is relatively constant. Therefore, the peak value at the BPFO is obvious, and the sideband modulation is weak, while the rotational frequency amplitude remains essentially unchanged. When the server detects a significant peak value at the BPFO in the envelope spectrum without significant sideband modulation, it determines that the bearing outer ring is faulty.
[0095] Rolling element failure: The envelope spectrum will show the BSF (rolling element fault characteristic frequency) and its harmonics. As the rolling element rotates, the fault point periodically enters and leaves the load area; therefore, the peak value at the BSF may fluctuate significantly and is often accompanied by cage frequency modulation (i.e., BSF ± m·FTF). When the server detects a significant peak value at the BSF in the envelope spectrum, accompanied by FTF sideband modulation, it determines it to be a rolling element failure.
[0096] Cage failure: The envelope spectrum will show the FTF (cage failure characteristic frequency) and its lower harmonics. The characteristic frequency is usually low (generally 0.3 to 0.5 times the rotational frequency). In cage failure, vibration energy is usually concentrated in the low-frequency range, with a significant peak at the FTF, often accompanied by impact components, and the kurtosis may also increase. When the server detects a significant peak at the FTF in the envelope spectrum, it determines that it is a cage failure.
[0097] In specific judgment, the server first determines whether a bearing fault exists based on kurtosis (greater than 5-8) and peak factor (greater than 6). Then, it performs envelope demodulation on the first sensor signal to obtain the envelope spectrum, and extracts the amplitude or energy at each theoretical characteristic frequency (BPFI, BPFO, BSF, FTF) in the envelope spectrum. If the amplitude at a certain characteristic frequency exceeds a preset threshold (e.g., 3 times the mean of the surrounding noise floor), it is determined that the component is faulty. If multiple characteristic frequencies exceed the threshold simultaneously, the server can take the fault type corresponding to the characteristic frequency with the highest energy as the primary conclusion and output a secondary fault indication (e.g., "Mainly a fault in the inner race of the bearing, with suspected minor damage to the outer race"). Simultaneously, the server can also input the energy values at each characteristic frequency as key features into a large language model, from which the model outputs a comprehensive diagnostic result.
[0098] The server can also analyze the changing trends of key features over time by combining the results of multiple historical detections, rather than relying on the key feature values of a single detection. For example, if the kurtosis slowly increases from 4 to 6 over a week, even if the current kurtosis does not exceed the threshold of 5, the upward trend is obvious, which can provide an early warning of potential faults. Trend features can be extracted using methods such as moving averages, exponential smoothing, or linear regression, serving as "changing trend conditions" in the physical characteristics of faults.
[0099] In some embodiments, when a key feature simultaneously matches multiple fault physical features (e.g., satisfying both the tip discharge condition and the particle discharge condition), the server can use a priority method (presetting the priority of various fault types, such as prioritizing insulation faults over mechanical faults), a confidence comparison method (comparing the historical accuracy or current matching score of each rule), or an evidence combination method (such as DS evidence theory) to obtain a unique conclusion. For example, fault types with high priority or high confidence can be used as the faults corresponding to the key feature, and a list of multiple candidate faults can be output for manual review or further analysis in subsequent steps (such as large language models).
[0100] Among them, evidence combination methods (such as DS evidence theory) are a fusion method for handling uncertain information. They combine evidence from multiple rules or sensors to obtain a comprehensive confidence distribution, thereby enabling decision-making under conflicting or incomplete evidence. In this embodiment, the server first constructs a basic probability assignment (BPA1) for the probability of each fault type under a single key feature (such as kurtosis) based on the matching result with the fault physics rule. For example, if kurtosis = 12.5, according to the rule "kurtosis > 5 and spectral centroid > 18kHz → tip discharge," but since the spectral centroid has not yet participated in the judgment, the server can set: BPA1 (tip discharge) = 0.6 (kurtosis strongly supports), BPA1 (free particle discharge) = 0.3 (kurtosis may also partially support particle discharge), BPA1 (normal) = 0.1 (kurtosis is slightly high, but theoretically the kurtosis of a normal signal is about 3). Then, the server constructs a second basic probability assignment based on the second key feature (e.g., spectral centroid = 22kHz), denoted as BPA2: BPA2 (point discharge) = 0.8 (spectral centroid strongly supports point discharge), BPA2 (free particle discharge) = 0.15, BPA2 (normal) = 0.05. Next, the server constructs a third basic probability assignment based on the third key feature (e.g., amplitude variation coefficient = 1.8), denoted as BPA3: BPA3 (point discharge) = 0.5, BPA3 (free particle discharge) = 0.4 (particle discharge is also often accompanied by a high variation coefficient), BPA3 (normal) = 0.1. The server merges BPA1, BPA2, and BPA3 to obtain the comprehensive confidence level for each fault type. Finally, the server selects the fault type with the highest comprehensive confidence level that exceeds a preset threshold as the detection conclusion. If there are multiple fault types corresponding to the highest confidence level (i.e., multiple faults listed together) or the highest confidence level is below the threshold, the server can output "uncertain" or a candidate fault list for manual review or further analysis by a large language model.
[0101] For example, after fusing BPA1, BPA2, and BPA3, the calculated overall confidence level is 0.85 for tip discharge, 0.12 for free particle discharge, and 0.03 for normal discharge. Since tip discharge has the highest confidence level (0.85) and exceeds the preset threshold (e.g., 0.6), the server ultimately determines it to be tip discharge. If the overall confidence level is below the threshold (e.g., the highest is only 0.4), the server outputs "Candidate Fault: Tip Discharge (0.4), Free Particle Discharge (0.35), Uncertain (0.25)", for further evaluation by manual review or large language model.
[0102] During long-term operation, once a fault detection result is manually confirmed to be correct, the server can use the correspondence between the key features and the fault type as a new sample to fine-tune the thresholds or weights of the fault's physical features (such as updating the Bayesian prior probability or adjusting the fuzzy membership function parameters). This gives the system self-learning capabilities, continuously improving diagnostic accuracy.
[0103] This application provides a fault detection method that uses a signal reconstruction model to perform unsupervised reconstruction of sensor signals in operation and calculates the reconstruction error. This method can sensitively capture any abnormal signals that deviate from the normal pattern, avoiding missed or false alarms caused by changes in the noise environment. Furthermore, feature extraction and judgment based on fault physical characteristics are triggered only when the reconstruction error exceeds a threshold. This retains the high sensitivity of deep learning models to complex anomalies while eliminating non-fault interference false alarms common in purely data-driven methods through physical mechanism constraints. This dual mechanism of "anomaly perception + physical verification" significantly improves the accuracy and reliability of fault detection. At the same time, since subsequent calculations are only performed when an anomaly is suspected, the overall detection efficiency is greatly improved compared to methods that run complex rules or models throughout the process.
[0104] In some embodiments, the first sensing signal includes multiple physical types of sensing signals, such as simultaneously acquiring vibration signals, ultrasonic signals, and ultra-high frequency signals. The server can independently execute steps 301 to 303 for each physical type of signal to obtain their respective key features. Then, the server fuses the key features corresponding to the multiple physical types of sensing signals to generate fault prompt words for subsequent fault detection. Accordingly, step 304 may include: the server concatenating the key features corresponding to the multiple physical types of sensing signals to obtain fault prompt words; then, the server inputs the fault prompt words into a large language model, and the large language model compares the key features of each sensing signal with the fault physical features of the corresponding physical type to obtain a fault detection report for the device.
[0105] The solution provided in this application directly concatenates key features corresponding to multiple physical types of sensor signals into fault indication words, and automatically compares the key features of each physical type with the corresponding physical fault features using a large language model. This avoids subjective biases and difficulties in scenario adaptation caused by manually designing complex fusion rules. The large language model has powerful semantic understanding and cross-feature association reasoning capabilities, enabling adaptive fusion based on the physical characteristics of different types of sensor signals. It fully explores the complementary information between various physical quantities, thereby generating more accurate and comprehensive fault detection reports without human intervention. This approach not only significantly reduces system deployment and maintenance costs and improves the automation level and scenario generalization capability of diagnosis, but also effectively solves the problem of excessive reliance on prior knowledge and manual rules in traditional multi-sensor fusion methods, further improving the accuracy and reliability of fault detection.
[0106] For high-voltage GIS equipment, ultrasonic sensors, ultra-high frequency sensors, and high-frequency current transformers can be deployed simultaneously to collect acoustic emission signals, electromagnetic wave signals, and pulse current signals, respectively. For rotating machinery, vibration acceleration signals, acoustic emission signals, and rotational speed signals can be collected simultaneously. Each physical signal has different sensitivities, propagation paths, and anti-interference capabilities for the same fault. Multi-physical quantity fusion can significantly improve the robustness and accuracy of diagnosis. The server can connect the key features of each physical type into a long vector according to a preset order and convert its values into natural language descriptions, constructing part of the fault indication words.
[0107] Before feature concatenation, the server can assign a weight to key features for each physical type. The weight is dynamically calculated based on the key feature's signal-to-noise ratio (SNR) under the current operating conditions or its historical diagnostic contribution. For example, when strong electromagnetic interference exists on-site, the SNR of Ultra High Frequency (UHF) signals decreases, and their weight automatically decreases; while Acoustic Emission (AE) signals have strong anti-interference capabilities, and their weight increases. The weighted features are then concatenated, enabling the large language model to distinguish the credibility of different physical quantities.
[0108] For the same fault type, sensor signals from different physical types may give conflicting initial judgments (e.g., AE is judged as tip discharge, while UHF is judged as free particles). The large language model can use the "Comparative Analysis" command in the prompt to require the model to output the support level of each piece of evidence and give a final decision. For example, adding the prompt: "The conclusions of the sensor signals from different physical types are inconsistent. Please analyze comprehensively and explain the reasons based on the fault physical mechanism and typical feature weights." The model will output the reasoning process, such as "Because the phase convergence characteristics are more obvious in the UHF signal, while the AE signal may be affected by mechanical vibration, the UHF conclusion is accepted."
[0109] The server can dynamically adjust the template and level of detail of the prompts based on equipment type, fault history, and current operating conditions. For example, for newly commissioned equipment, the prompts require the model to output more detailed analysis steps; for older equipment, the focus is on comparing historical feature trends. Large language models can receive these meta-instructions to achieve personalized diagnostics.
[0110] For example, for a newly commissioned GIS device (commissioned for less than one month, with no fault records or maintenance history), the server can construct the following prompt: "This is a newly commissioned GIS device. Please conduct a detailed analysis of the key characteristics of the following multi-physical quantity sensor signals, explaining your reasoning process step by step, including the meaning of each physical quantity characteristic, the comparison results with the fault physical rules, and the mutual corroboration relationship between the evidence. Finally, provide a diagnostic conclusion and subsequent monitoring recommendations. Ultrasonic characteristics: kurtosis 12.5, spectral centroid 22.3kHz, RMS value 0.35mV; UHF characteristics: pulse amplitude 8.2mV, phase convergence 0.85; High-frequency current characteristics: apparent discharge 115pC, discharge repetition rate 320 times / second." Then, the large language model will output a detailed analysis report in steps, such as: Step 1: Analyze the ultrasonic characteristics—the kurtosis of 12.5 is much greater than the normal value of 3, and the spectral centroid of 22.3kHz exceeds the 18kHz threshold, indicating the presence of a high-frequency pulse signal, possibly a tip discharge or free particle discharge. Step 2: Analyze the UHF characteristics—the phase convergence of 0.85, concentrated in the positive half-cycle, is consistent with the phase characteristics of tip discharge, while free particle discharge typically has random phases. Step 3: Analyze the high-frequency current characteristics—the discharge quantity is 115pC, and the rise time is 8ns, which is typical for tip discharge. Step 4: Comprehensive judgment—all three physical quantities point to tip discharge, and the phase characteristics of UHF further confirm that it is tip discharge rather than free particle discharge. Conclusion: Tip discharge exists in the high-voltage conductor, with a high confidence level (92%). Recommendation: Due to the new commissioning of the equipment, it is recommended to check the surface of the high-voltage conductor for burrs during the first power outage maintenance. For an old piece of equipment (operating for over 10 years, historical data shows that the kurtosis has slowly increased from 4 to 8 over the past year, and the RMS value has gradually increased from 0.12mV to 0.30mV), the server can construct the following prompt: "This is an old GIS device that has been operating for 10 years. Please focus on the following key characteristics and their comparison with historical trends, analyze the severity and development trend of the current fault, and output diagnostic conclusions and maintenance suggestions. Current characteristics: Kurtosis 12.5 (4.2 in the same period last year), spectral centroid 22.3kHz (12.5kHz in the same period last year), RMS value 0.35mV (0.15mV in the same period last year), historical trend shows a gradual deterioration over the past year. Historical highest kurtosis 8.0 (last quarter)." The kurtosis has now surged to 12.5. The large language model then outputs a report emphasizing trend comparisons and severity assessments, for example: "Comparison with historical data: Over the past year, the kurtosis has increased from 4.2 to 12.5, and the effective value has increased from 0.15mV to 0.35mV, indicating a continued increase in discharge intensity and a recent accelerated deterioration. The current spectral centroid is 22.3kHz, significantly higher than the 12.5kHz of the same period last year, suggesting that the discharge type may have evolved from early surface discharge to a more severe tip discharge. Diagnostic conclusion: Tip discharge, moderate to high severity. It is recommended to schedule a power outage for maintenance within one month, inspecting the high-voltage conductor and equalizing shield, while strengthening real-time monitoring. If the discharge exceeds 300pC, immediate power outage is necessary." By using this personalized diagnostic approach that dynamically adjusts prompts, the large language model can provide differentiated analysis depth and maintenance suggestions based on the device's lifecycle stage and operational history, making the diagnostic results more aligned with actual maintenance needs.
[0111] Figure 4 This is a block diagram of a fault detection device according to an embodiment of this application. The fault detection device is used to perform the steps of the above-described fault detection method, see [link to relevant documentation]. Figure 4 The fault detection device includes: The reconstruction module 401 is used to reconstruct the first sensing signal collected by the device in the operating state based on the signal reconstruction model to obtain the second sensing signal. The signal reconstruction model is trained with the sensing signal of the device when there is no fault as the training sample and with the goal of minimizing the reconstruction error. The first determining module 402 is used to determine the reconstruction error of the signal reconstruction model based on the first sensing signal and the second sensing signal. The reconstruction error is used to indicate the difference between the first sensing signal and the second sensing signal. The feature extraction module 403 is used to extract features from the first sensing signal when the reconstruction error is greater than the error threshold, and obtain key features, which are the physical characteristics related to the fault in the sensing signal of the device. The second determining module 404 is used to determine the fault detection result of the equipment based on key features and fault physical features, wherein the fault physical features are used to indicate the physical characteristics of the sensing signal when the equipment fails.
[0112] In some embodiments, the device is an electrical device, and key features include at least one of the following: The kurtosis of the first sensing signal, which is used to indicate the degree of peaks in the amplitude distribution of the first sensing signal; The spectral centroid of the first sensing signal, which is used to indicate the centroid frequency of the spectrum of the first sensing signal; The amplitude variation coefficient of the first sensing signal is used to indicate the degree of dispersion of the amplitude of the first sensing signal. The effective value of the first sensing signal is used to indicate the overall energy level of the first sensing signal.
[0113] In some embodiments, the device is rotating machinery, and key features include at least one of the following: The kurtosis of the first sensing signal, which is used to indicate the degree of peaks in the amplitude distribution of the first sensing signal; The peak factor of the first sensing signal, which is used to indicate the impact characteristics of the first sensing signal; The envelope spectrum fault characteristic frequency of the first sensing signal is used to indicate the degree of energy concentration of a specific frequency component in the spectrum of the first sensing signal after envelope demodulation. The frequency conversion amplitude of the first sensing signal is used to indicate the magnitude of the frequency conversion and its harmonic components in the spectrum of the first sensing signal. The sideband energy of the first sensing signal is used to indicate the degree of energy concentration in the sidebands on both sides of the carrier frequency in the spectrum of the first sensing signal.
[0114] In some embodiments, the second determining module 404 is configured to generate fault prompt words based on the key features; input the fault prompt words into a large language model; compare the key features with the fault physical features through the large language model to obtain a fault detection report of the device; the fault detection report includes at least one of the following: fault type, fault severity information, and fault repair plan.
[0115] In some embodiments, the physical characteristics of the fault include the physical characteristics of the sensing signals under multiple fault types; the second determining module 404 is used to determine that the device has a fault and output the fault type when the large language model detects that the key feature matches the physical characteristics corresponding to any fault type; and to determine that the device does not have a fault when the large language model detects that the key feature does not match the physical characteristics corresponding to any fault type.
[0116] In some embodiments, the first sensing signal includes sensing signals of multiple physical types; The second determining module 404 concatenates the key features corresponding to the multiple physical types of sensor signals to obtain the fault prompt words; inputs the fault prompt words into a large language model, and compares the key features of each sensor signal with the fault physical features of the corresponding physical type through the large language model to obtain a fault detection report of the device.
[0117] In some embodiments, the second determining module 404 is further configured to determine that the device is fault-free if the reconstruction error is not greater than an error threshold.
[0118] In some embodiments, the apparatus further includes: The third determination module is used to determine the error threshold based on the numerical distribution of historical reconstruction errors of the equipment in a fault-free state.
[0119] In some embodiments, the signal reconstruction model is a one-dimensional convolutional denoising autoencoder, which includes an encoder and a decoder. The encoder is used to extract features of the input signal, and the decoder is used to reconstruct the input signal based on the features output by the encoder. The reconstruction module 401 is used to input the first sensing signal into the encoder, extract the signal feature vector through the convolution and pooling operations of the one-dimensional convolutional layer in the encoder, input the signal feature vector into the decoder, and reconstruct the second sensing signal through the upsampling and convolution operations of the one-dimensional transposed convolutional layer in the decoder.
[0120] In some embodiments, the sensing signal of the device when it is fault-free includes at least one of the following: Sensor signals collected by the equipment under normal operating conditions with power on and confirmed to be fault-free; The sensor signal segments are automatically filtered and extracted from the collected sensor signals. The sensor signal segments have a preset duration and do not contain amplitude pulses exceeding the preset value.
[0121] In some embodiments, the apparatus further includes: The training module is used to pre-train the initial reconstruction model with the third sensor signal collected by the device under power failure as training sample, aiming to minimize the reconstruction error, and obtain the intermediate reconstruction model. The third sensor signal is a pure environmental noise sample. The fourth sensor signal collected by the device under normal operation with power and no faults is used as training sample. With the goal of minimizing the reconstruction error, the parameters of the intermediate reconstruction model are adjusted to obtain the signal reconstruction model.
[0122] This application provides a fault detection device that performs unsupervised reconstruction of sensor signals in operation using a signal reconstruction model and calculates the reconstruction error. This enables the device to sensitively capture any abnormal signals that deviate from the normal pattern, avoiding missed or false alarms caused by changes in the noise environment. Furthermore, feature extraction and judgment based on the physical characteristics of the fault are triggered only when the reconstruction error exceeds a threshold. This retains the high sensitivity of deep learning models to complex anomalies while eliminating non-fault interference false alarms common in purely data-driven methods through physical mechanism constraints. This dual mechanism of "anomaly perception + physical verification" significantly improves the accuracy and reliability of fault detection. At the same time, since subsequent calculations are only performed when an anomaly is suspected, the overall detection efficiency is greatly improved compared to methods that run complex rules or models throughout the process.
[0123] It should be noted that the fault detection device provided in the above embodiments is only illustrated by the division of the above functional modules when detecting whether a device is faulty. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the fault detection device and the fault detection method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0124] This application also provides a computer device. Figure 5 This is a schematic diagram of a computer device 500 according to an embodiment of this application. The computer device 500 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 501 and one or more memories 502. The memory 502 stores at least one computer program, which is loaded and executed by the processor 501 to implement the fault detection methods provided in the various method embodiments described above. Of course, the computer device may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device may also include other components for implementing device functions, which will not be elaborated here.
[0125] In the embodiments of this application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can act as the execution subject to implement the technical solutions provided in the embodiments of this application. When the computer device is configured as a server, the server can act as the execution subject to implement the technical solutions provided in the embodiments of this application. Alternatively, the technical solutions provided in this application can be implemented through the interaction between the terminal and the server. The embodiments of this application do not limit this.
[0126] This application also provides a computer-readable storage medium storing at least one computer program. This computer program is loaded and executed by a processor of a computer device to implement the operations performed by the computer device in the fault detection method of the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0127] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.
[0128] This application also provides a computer program product, including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the fault detection method provided in the various optional implementations described above.
[0129] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0130] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A fault detection method, characterized in that, The method includes: Based on the signal reconstruction model, the first sensing signal collected by the device in operation is reconstructed to obtain the second sensing signal. The signal reconstruction model is trained with the sensing signal of the device when it is fault-free as the training sample and with the goal of minimizing the reconstruction error. Based on the first sensing signal and the second sensing signal, the reconstruction error of the signal reconstruction model is determined, and the reconstruction error is used to indicate the difference between the first sensing signal and the second sensing signal. If the reconstruction error is greater than the error threshold, feature extraction is performed on the first sensing signal to obtain key features, which are the fault-related physical characteristics in the sensing signal of the device. Based on the key features and physical fault characteristics, the fault detection result of the device is determined, wherein the physical fault characteristics are used to indicate the physical characteristics of the sensing signal when the device fails.
2. The method according to claim 1, characterized in that, The process of determining the fault detection result of the device based on the key features and physical characteristics of the fault includes: Based on the aforementioned key features, fault prompt words are generated; The fault prompt words are input into a large language model, and the large language model compares the key features with the physical features of the fault to obtain a fault detection report of the device. The fault detection report includes at least one of the following: the fault type of the device, the severity of the fault, and the fault repair plan.
3. The method according to claim 2, characterized in that, The physical characteristics of the fault include the physical characteristics of the sensing signals under multiple fault types; The process of comparing the key features with the physical features of the fault using the large language model to obtain a fault detection report for the device includes: If the large language model detects that the key feature matches the physical feature corresponding to any fault type, it determines that the device is faulty and outputs the fault type. If the large language model detects that the key feature does not match any physical feature corresponding to any fault type, it is determined that the device is not faulty.
4. The method according to claim 2, characterized in that, The first sensing signal includes sensing signals of multiple physical types; The generation of fault prompt words based on the key features includes: The key features corresponding to the multiple physical types of sensor signals are spliced together to obtain the fault prompt words; The step of inputting the fault prompt words into a large language model, comparing the key features with the physical features of the fault through the large language model, and obtaining a fault detection report for the device includes: The fault prompt words are input into a large language model, which compares the key features of each sensor signal with the corresponding physical fault physical features to obtain a fault detection report for the device.
5. The method according to claim 1, characterized in that, The method further includes: If the reconstruction error is not greater than the error threshold, the device is determined to be fault-free.
6. The method according to claim 1, characterized in that, The method further includes: The error threshold is determined based on the numerical distribution of the historical reconstruction error of the device under fault-free conditions.
7. The method according to claim 1, characterized in that, The signal reconstruction model is a one-dimensional convolutional denoising autoencoder, which includes an encoder and a decoder. The encoder is used to extract features of the input signal, and the decoder is used to reconstruct the input signal based on the features output by the encoder. The signal reconstruction model reconstructs the first sensing signal collected by the device during operation to obtain the second sensing signal, including: The first sensing signal is input into the encoder, and the signal feature vector is extracted by the convolution and pooling operations of the one-dimensional convolutional layer in the encoder. The signal feature vector is input into the decoder, and the second sensing signal is reconstructed by upsampling and convolution operations of the one-dimensional transposed convolutional layer in the decoder.
8. The method according to claim 1, characterized in that, The sensing signal of the device when it is fault-free includes at least one of the following: The sensor signals collected by the device under normal operating conditions that are powered on and confirmed to be fault-free; The sensor signal segments are automatically filtered and extracted from the collected sensor signals. The sensor signal segments have a preset duration and do not contain amplitude pulses exceeding a preset value.
9. The method according to claim 1, characterized in that, The method further includes: The third sensor signal collected by the device in the power outage state is used as a training sample. With the goal of minimizing the reconstruction error, the initial reconstruction model is pre-trained to obtain an intermediate reconstruction model. The third sensor signal is a pure environmental noise sample. The fourth sensor signal collected by the device under normal operating conditions with power and no faults is used as a training sample. With the goal of minimizing the reconstruction error, the parameters of the intermediate reconstruction model are adjusted to obtain the signal reconstruction model.
10. A fault detection device, characterized in that, The device includes: The reconstruction module is used to reconstruct the first sensing signal collected by the device in the operating state based on the signal reconstruction model to obtain the second sensing signal. The signal reconstruction model is trained with the sensing signal of the device when it is fault-free as the training sample and with the goal of minimizing the reconstruction error. The first determining module is used to determine the reconstruction error of the signal reconstruction model based on the first sensing signal and the second sensing signal, wherein the reconstruction error is used to indicate the difference between the first sensing signal and the second sensing signal; The feature extraction module is used to extract features from the first sensing signal when the reconstruction error is greater than the error threshold, and obtain key features, wherein the key features are the physical characteristics related to the fault in the sensing signal of the device. The second determining module is used to determine the fault detection result of the device based on the key features and the physical characteristics of the fault, wherein the physical characteristics of the fault are used to indicate the physical characteristics of the sensing signal when the device fails.
11. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as the fault detection method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one computer program, which is used to perform the fault detection method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fault detection method as described in any one of claims 1 to 9.