Household appliance multi-mode fault detection method and system based on double-graph attention network

By fusing dual-graph attention networks with multimodal features, the problems of misjudgment, response lag, and poor adaptability in the fault detection of household appliances are solved, enabling accurate fault identification and early warning of household appliances, and adapting to the detection needs of different brands and complex categories of appliances.

CN121834459APending Publication Date: 2026-04-10HUNAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN UNIV
Filing Date
2025-12-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for detecting faults in household appliances are susceptible to environmental interference and misjudgment, have slow response times, rely on expert experience, have poor adaptability, are difficult to adapt across brands, and have limited monitoring dimensions, making them unable to meet the testing needs of complex categories of appliances.

Method used

A fault detection model is constructed by using a dual-graph attention network and a multimodal fusion feature extraction method. Multi-dimensional data is acquired through sensors, and a fault detection model is constructed by using a dual-graph attention network and a semi-supervised classification network combined with a multimodal Transformer layer for deep feature fusion.

Benefits of technology

It achieves accurate identification and early warning of household appliance faults, adapts to new types of faults and multiple types of appliances, reduces model training costs, and improves detection accuracy and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834459A_ABST
    Figure CN121834459A_ABST
Patent Text Reader

Abstract

The invention discloses a household electrical appliance multi-modal fault detection method based on a double-graph attention network, and the method comprises the steps: collecting the multi-dimensional data of a household electrical appliance through sensing layer multi-sensor (temperature, current and voltage, vibration and thermal imaging) and log reading, extracting the time sequence features through a double-graph attention network, combining with the depth features of a multi-modal Transform layer, and carrying out the recognition of the fault of the household electrical appliance. Carrying out anomaly detection by using a self-encoder; according to the method, a multi-dimensional data acquisition module covers all working condition stages such as starting, operation and standby of the household electrical appliance, a data association value is mined in combination with a deep feature fusion technology, a high-precision fault detection and pre-judgment model is constructed, and accurate identification and early warning of potential faults of the household electrical appliance are realized. The core target is to solve the defects of a traditional detection method in a targeted manner and provide safer, more intelligent and more worry-saving household appliance use experience for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning and smart home fault diagnosis technology, and more specifically, relates to a method and system for multimodal fault detection of home appliances based on a dual-graph attention network. Background Technology

[0002] As people's living standards improve, household appliances (such as refrigerators, air conditioners, washing machines, and water heaters) have become indispensable in family life. However, these appliances are prone to malfunctions over time due to factors such as aging parts, changes in the operating environment, and improper operation.

[0003] Existing methods for detecting faults in household appliances include the following: The first is threshold-based detection, which involves pre-setting normal threshold ranges for key operating parameters of the household appliance (such as temperature, voltage, and current). When a parameter exceeds the threshold, it is considered a fault. This method is widely used in early, simple household appliance monitoring equipment due to its simplicity and low cost. The second is expert system-based detection, which relies on domain experts' experience summarizing the fault mechanisms of household appliances to build a fault rule base. Fault identification is achieved by matching the device's operating status data with the fault feature information in the rule base in real time. The third is traditional machine learning-based detection, which manually extracts shallow statistical features (such as mean, variance, and peak value) from the household appliance's operating data and uses classic algorithms such as support vector machines, decision trees, and random forests to build a fault classification model for preliminary fault identification. The fourth is single-sensor data-based detection, which collects operating data in only one dimension using a single sensing device such as a temperature sensor or voltage sensor, and uses this as the sole basis for fault judgment. This is commonly seen in low-cost, single-function household appliance monitoring scenarios.

[0004] However, the aforementioned existing methods for detecting faults in household appliances all have some significant drawbacks: 1. Existing threshold-based detection methods are susceptible to external interference such as fluctuations in ambient temperature and humidity and instantaneous changes in power grid voltage, which can lead to misjudgments. They cannot effectively distinguish between parameter fluctuations during normal operation and fault abnormal signals. Furthermore, they can only trigger alarms after a fault occurs, and cannot identify fault precursors in advance. The response is significantly delayed, which can easily lead to the escalation of the fault. 2. Existing detection methods based on expert systems are highly dependent on expert experience. The construction of the rule base requires a lot of time and manpower, and the update and iteration efficiency of the rule base is low. It has poor adaptability when facing the failure modes of new household appliances or complex failure scenarios under complex working conditions, making it difficult to achieve technical expansion. 3. Existing detection methods based on traditional machine learning are not targeted enough in extracting data features. Manually designed shallow features cannot accurately depict the deep evolution of complex faults, and the generalization ability of the model is weak. When faced with the differences in operating characteristics of different brands and models of home appliances, it is necessary to retrain the model. The adaptation of cross-brand and cross-model home appliances is difficult and costly. 4. Existing detection methods based on single sensor data have limited monitoring dimensions and cannot comprehensively capture the complete operational status information of household appliances. They are prone to missed fault detection due to missing key data, and in particular, they cannot meet the fault detection needs of complex categories of household appliances such as refrigerators and air conditioners that work in concert with multiple components. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a method and system for multimodal fault detection of household appliances based on a dual-graph attention network. The aim is to solve the technical problems of existing threshold-based detection methods, which are susceptible to external interference such as fluctuations in ambient temperature and humidity and instantaneous changes in grid voltage, leading to misjudgments; inability to effectively distinguish between parameter fluctuations during normal operation and abnormal fault signals; and the inability to trigger alarms only after a fault occurs, failing to identify fault precursors in advance, exhibiting significant response lag, and easily leading to fault escalation. Furthermore, existing expert-system-based detection methods are highly dependent on expert experience, requiring significant time and manpower to build rule bases, and suffer from low efficiency in updating and iterating rule bases. This is particularly problematic when facing fault modes or complex operating conditions of new household appliances. The existing detection methods, based on traditional machine learning, suffer from poor adaptability to complex fault scenarios, making it difficult to extend the technology. Furthermore, they lack specificity in extracting data features, and manually designed shallow features cannot accurately depict the deep evolution of complex faults. The models also have weak generalization capabilities, requiring retraining to address the differences in operating characteristics across different brands and models of home appliances. This leads to significant challenges and high costs associated with adapting to different brands and models. Additionally, existing detection methods based on single sensor data have limited monitoring dimensions, failing to comprehensively capture the complete operational status information of home appliances. This can result in missed fault detection due to missing key data, and they are particularly inadequate for fault detection needs of complex categories of home appliances, such as refrigerators and air conditioners, which require multiple components to work collaboratively.

[0006] To achieve the above objectives, according to one aspect of the present invention, a method for multimodal fault detection of household appliances based on a dual-graph attention network is provided, comprising the following steps: (1) Acquire real-time operating data of household appliances through sensors; (2) Input the real-time operating data of the household appliance obtained in step (1) into the pre-trained fault detection model to obtain the preliminary fault classification result of the household appliance.

[0007] Preferably, the real-time operating data of the household appliances includes the triaxial acceleration timing signal of the household appliances collected by the vibration sensor, the current timing signal and voltage timing signal of the household appliances collected by the electrical signal sensor, the 128×128 resolution thermal image of the household appliances collected by the thermal imaging device, and the log text data generated by the daily operation of the household appliances. Preferably, the fault detection model adopts an architecture that integrates a dual-graph attention network and a semi-supervised classification network, and includes an input layer, a hidden layer, and an output layer, with the following structure: The first layer is the input layer, which takes in the multimodal raw data generated during the operation of household appliances. This includes the triaxial acceleration time-series signal of the household appliances collected by the vibration sensor, the current time-series signal and voltage time-series signal of the household appliances collected by the electrical signal sensor, the 128×128 resolution thermal image of the household appliances collected by the thermal imaging device, and the log text data generated by the household appliances during daily operation. This layer filters and normalizes the multimodal data to obtain a 256×256 vibration time-frequency diagram, a time-series vector of a one-dimensional electrical time-series signal aligned to a 1000Hz sampling rate, a 128×128 thermal image, and a 64-dimensional log semantic vector, which together serve as the processed multimodal data. The second layer is a hidden layer, whose input is the multimodal data processed by the first layer. This layer first uses a specific feature extraction sublayer to extract features from the multimodal data to obtain modal features. Then, it uses a sensor association graph attention sublayer and a time series graph attention sublayer to calculate the association graph and time series signal fusion features for each feature data. After that, it uses a multimodal Transformer fusion sublayer to fuse the weight information of all feature data to obtain fusion node features with a dimension of 256. Finally, it uses a semi-supervised classification joint sublayer to integrate the fusion node features to obtain a fusion feature vector with a dimension of 1×256, a fault probability distribution, and an initial risk value, and outputs them. The third layer is the output layer, whose inputs are the fused feature vector output from the second layer, the fault probability distribution, and the initial risk value. This layer first uses an autoencoder to quantize the reconstruction error of the fused feature vector to obtain a scalar anomaly score in the 0-1 interval. Then, the fault probability distribution is subjected to probability calibration and classification correction using the Platt scaling algorithm to obtain a precise fault probability distribution in 1×11 dimensions. Afterward, the initial risk value is subjected to working condition weighting and risk classification to obtain a scalar risk value R in the 0-1 interval. Preferably, the process of filtering and normalizing the triaxial acceleration time-series signal of a household appliance to obtain a vibration time-frequency diagram with a dimension of 256×256 is as follows: First, the original triaxial acceleration time-series signal is bandpass filtered; then, the composite acceleration time-series signal obtained after filtering is normalized by min-max to map the signal amplitude to the [0,1] interval; finally, the normalized composite acceleration time-series signal is subjected to a short-time Fourier transform (STFT), with a time window length of 256 sampling points and a window shift of 1 sampling point, and the frequency axis dimension is determined to be 256, so as to convert the one-dimensional composite acceleration time-series signal into a two-dimensional time-frequency matrix. Subsequently, if the size of the transformed time-frequency matrix is ​​not 256×256, the time-frequency matrix is ​​normalized by bilinear interpolation, and finally a vibration time-frequency diagram with a dimension of 256×256 is obtained and output. The process of filtering and normalizing the current and voltage timing signals of household appliances to obtain a timing vector of a one-dimensional electrical timing signal aligned to a 1000Hz sampling rate is as follows: First, the current and voltage timing signals are low-pass filtered respectively; then, the filtered current and voltage timing signals are resampled; finally, the resampled current and voltage timing signals are normalized respectively; and finally, the normalized current and voltage timing signals are concatenated to obtain a timing vector of a one-dimensional electrical timing signal aligned to a 1000Hz sampling rate. The process of filtering and normalizing a 128×128 resolution thermal image of a household appliance to obtain a 128×128 resolution thermal image is as follows: First, Gaussian filtering is used to denoise the resolution thermal image; then, the denoised 128×128 thermal image is normalized to finally obtain a 128×128 resolution thermal image. The process of filtering and normalizing the log text data generated by household appliances during daily operation to obtain a 64-dimensional log semantic vector is as follows: First, interference information in the log text data is removed, retaining only the valid text content; then, the log text data after removing interference information is standardized; next, the standardized log text data is downsampled; then, a lightweight text encoder is used to convert individual log texts in the downsampled log text data into initial semantic vectors; finally, each encoded initial semantic vector is dimensionally compressed and normalized, and all adjusted initial semantic vectors constitute a 64-dimensional log semantic vector.

[0008] Preferably, the process of quantizing the reconstruction error of the fused feature vector to obtain the scalar anomaly score in the 0-1 interval is as follows: First, samples marked as normal operation are screened, and the fused feature vectors of these samples are used to count all reconstruction errors. The 95th percentile of all reconstruction errors is taken as the anomaly judgment threshold. Finally, the reconstruction error between the fused feature vector of all samples and the reconstruction output is calculated, and the reconstruction error is divided by the anomaly judgment threshold to obtain the scalar anomaly score. The risk values ​​are weighted by operating conditions and classified by risk level to obtain scalar risk values ​​in the 0-1 range. This process is as follows: First, extract the equipment service life from the risk value. Average runtime over the past 7 days Current operating load rate Ambient temperature and humidity deviation values These are four categories of key operating condition indicators; then, the basic weight of each category of key operating condition indicator is obtained, and the operating condition weight coefficient is obtained based on the basic weight. : ; in, Weighted by the equipment's service life, and has ; The weight is the average runtime over the past 7 days, and there is ; This is the weight of the current operating load rate of household appliances, which is equal to the operating load rate; The weight of the environmental temperature and humidity deviation value is directly taken as the normalized value of the deviation of the environmental temperature and humidity relative to the standard value. Subsequently, based on the working condition weighting coefficient Calculate the weighted scalar risk value : ; in, The initial risk value output by the hidden layer; Subsequently, the judgment Is it greater than 1 or less than 0? If it is the former, set the anomaly score to 1; if it is the latter, set the anomaly score to 0, to ensure that the weighted risk value is still in the 0-1 range.

[0009] Preferably, the fault detection model is established in the following way: (2-1) Obtain real-time data of household appliances during operation, filter and normalize the obtained real-time data to obtain the processed real-time data as a dataset, and divide the dataset into training set, validation set and test set in a ratio of 7:2:1. (2-2) For each sample in the training set obtained in step (2-1), the sample is input into the first layer of the fault detection model to obtain the vibration time-frequency map with a dimension of 256×256, the time-series vector of the one-dimensional electrical time-series signal aligned to a sampling rate of 1000Hz, the thermal image with a dimension of 128×128, and the log semantic vector with a dimension of 64. (2-3) For each sample in the training set obtained in step (2-1), the vibration time-frequency map with a dimension of 256×256, the time-series vector of the one-dimensional electrical time-series signal aligned to a sampling rate of 1000Hz, the thermal image with a dimension of 128×128, and the 64-dimensional log semantic vector corresponding to the sample obtained in step (2-2) are input into the special feature extraction sublayer in the hidden layer of the fault detection model to obtain the special features corresponding to the sample. (2-4) For each sample in the training set obtained in step (2-1), the specific features of each sample obtained in step (2-3) are input into the sensor association graph attention sublayer in the hidden layer to obtain the association graph corresponding to the sample, and the specific features of the sample are input into the time series graph attention sublayer in the hidden layer for feature fusion to obtain the time series signal fusion features corresponding to the sample. (2-5) For each sample in the training set obtained in step (2-1), the association graph and time-series signal fusion features corresponding to the sample obtained in step (2-4) and the special features corresponding to the sample obtained in step (2-3) are input into the multimodal Transformer fusion sublayer in the hidden layer to perform deep aggregation of cross-modal features, so as to obtain the fusion feature vector with a dimension of 1×256 corresponding to the sample. (2-6) For each sample in the training set obtained in step (2-1), the fusion feature vector with a dimension of 1×256 corresponding to the sample obtained in step (2-5) is input into the semi-supervised classification joint sub-layer to obtain the fault probability distribution and initial risk value with a dimension of 1×11 corresponding to the sample. (2-7) For each sample in the training set obtained in step (2-1), based on the initial risk value obtained in step (2-6), obtain the scalar risk value in the 0-1 interval corresponding to that sample. ; (2-8) For each sample in the training set obtained in step (2-1), the scalar risk value of the corresponding 0-1 interval obtained in step (2-7) is calculated. The cross-entropy loss for the fault classification corresponding to the sample is obtained from the fault probability distribution obtained in steps (2-6). : ; in The th sample is the th in the true failure probability distribution. One element, For the model, the fault probability distribution with dimension 1×11 is the first... One element; (2-9) For each sample in the training set obtained in step (2-1), the cross-entropy loss corresponding to the fault classification of that sample is obtained in step (2-8). Obtain the total loss corresponding to this sample. : ; in Cross-entropy loss for fault classification, This is a penalty item for abnormal scores; (2-10) For each sample in the training set obtained in step (2-1), based on the total loss corresponding to that sample obtained in step (2-11), the fault detection model is iteratively trained using gradient descent until the fault detection model reaches the preset number of iterations, or the total loss is reached. The fluctuation range is less than The model is trained until the total loss of the validation set does not increase significantly, or the accuracy of fault classification is higher than 90% and the F1 score of anomaly detection is greater than 88%, and the optimal parameters of the fault detection model at this time are obtained, thus obtaining the initially trained fault detection model. (2-11) Use the test set obtained in step (2-1) to test the fault detection model initially trained in step (2-10) until the detection accuracy reaches the optimal level, so as to obtain the final trained fault detection model.

[0010] Preferably, step (2-4) specifically involves: First, for the 256×256 vibration time-frequency map in the sample, the specific features of the vibration time-frequency map are flattened and mapped to 128 dimensions through a 1×1 convolutional layer to obtain a unified 128-dimensional vibration time-frequency map time-series vector. For the electrical time-series signal in the sample, the specific features of the electrical time-series signal are mapped to 128 dimensions through a 1×1 convolutional layer to obtain a unified 128-dimensional electrical time-series signal time-series vector. For the thermal image in the sample, it is flattened into a 1×128-dimensional vector through global average pooling. For the log semantic vector in the sample, the log semantic vector is extended into a 1×128-dimensional vector through linear mapping. Then, obtain the sample's first... time series vectors and the time series vectors The degree of correlation between them: ; in ∈[1, total number of time vectors of the samples], ∈[1, the total number of time-series vectors of the samples], and have ≠ , For learnable weight matrix, For the first The time series vector and the first The result of concatenating several time-series vectors Here, d is the activation function, and d is the dimension of the time-series vector. Subsequently, only all specific feature pairs with a correlation greater than 0.6 are retained, and all retained specific feature pairs constitute the correlation graph corresponding to the sample. Then, the vibration time-frequency diagram time-series vector and the electrical time-series signal time-series vector, both unified to 128 dimensions, corresponding to this sample are concatenated to form a 256-dimensional time-series feature vector. This time-series feature vector is then divided into... Individual time-series feature fragments , ..., And calculate each sub-temporal feature segment. With the Individual time-series feature fragments Time weight : ; in, For the first The dimension of each sub-temporal feature segment For the first The time weight matrix of each sub-temporal feature segment For the first The first sub-temporal feature fragment and the first Vector concatenation of sub-temporal feature segments, t∈[1, ]; Then, according to the first The first sub-temporal feature fragment and the first Temporal weights of individual time-series feature segments For the first Individual time-series feature fragments Update to obtain the updated version. Individual time-series feature fragments : ; in, For the first Bias of individual time-series feature segments Use the Sigmoid activation function; Finally, a 1×1 convolutional layer is used to update the first... Individual time-series feature fragments Mapped to 128 dimensions, all mapped sub-temporal feature fragments constitute the temporal signal fusion features corresponding to the sample.

[0011] Preferably, step (2-5) specifically involves: First, for the thermal image in the sample, the specific features corresponding to the thermal image are flattened into a 1×128-dimensional vector using global average pooling. For the electrical time-series signal in the sample, the specific features corresponding to the electrical time-series signal are compressed into a 1×128-dimensional vector through a 1×1 convolutional layer. For the vibration time-frequency map in the sample, the specific features corresponding to the vibration time-frequency map are transformed from a two-dimensional matrix into a 1×128-dimensional vector through convolutional and pooling layers. For the log semantic vector in the sample, the specific features corresponding to the log semantic vector are extended into a 1×128-dimensional vector through linear mapping. For the association graph generated by the sample in step (2-4), the association graph is mapped into a 1×128-dimensional vector through a linear layer. ; Then, using the 128-dimensional vector corresponding to the specific features of the thermal image. query vector The vector corresponding to the specific features of the time-series vector of the electrical time-series signal. as a key vector Sum value vector Attention weights are calculated and fused to obtain the fused features of thermal images and electrical time-series signals. : ; in The dimension of the key vector; Then, the 128-dimensional vector corresponding to the vibration time-frequency map fusion feature is used. as query vector The vector corresponding to the log semantic features as a key vector Sum value vector Attention weights are calculated and fused to obtain the fused features of the vibration time-frequency map and log semantic vector. ; Subsequently, the fusion features of the obtained thermal image and electrical time-series signal are analyzed. Fusion features of vibration time-frequency graph and log semantic vector and the association graph obtained in steps (2-3) Fusion vector of temporal features Perform cross-attention fusion to generate intermediate features. : ; ; in For vector concatenation, This represents the concatenated vector. () represents a linear transformation function; Finally, the intermediate fusion features With learnable weight matrix Weighted integration is performed to obtain a fused feature vector with a dimension of 1×256 corresponding to the sample.

[0012] Preferably, step (2-6) specifically involves: First, the fused feature vector of the samples is input into the pre-trained autoencoder to obtain anomaly scores: ; in This represents the anomaly detection threshold, and its value ranges from 0, 9 to 1. This represents the first feature vector corresponding to the sample. One element, The first sample reconstruction output represents the... There are elements, where x represents a sample; Then, the fused feature vector and anomaly score of the samples are input into the fully connected layer to obtain a fault probability distribution with a dimension of 1×11; Subsequently, the initial risk value of the sample is calculated using the anomaly score and the failure probability distribution. ; in This represents the maximum value in a fault probability distribution with dimension 1×11.

[0013] According to another aspect of the present invention, a method for multimodal fault detection of household appliances based on a dual-graph attention network is provided, comprising: The first module is used to acquire real-time operating data of home appliances through sensors; The second module is used to input the real-time operating data of the household appliances obtained by the first module into the pre-trained fault detection model in order to obtain the preliminary fault classification results of the household appliances.

[0014] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. Because the present invention adopts step (2), it utilizes dual-graph attention network and multimodal fusion feature extraction, which can comprehensively collect multi-dimensional data of household appliance operation and deeply mine the correlation features between parameters, accurately distinguish normal parameter fluctuations from fault precursor signals, and thus solve the technical problems of easy misjudgment, delayed response, and inability to identify fault precursors in advance in the existing detection methods based on threshold judgment. 2. Because the present invention uses steps (2-3) to (2-5), it extracts features by acquiring multimodal data of household appliances. It can quickly adapt to new faults and multiple types of household appliances without relying on expert experience to build a fixed rule base. Therefore, it can solve the technical problems of existing detection methods based on expert systems, such as reliance on experience, high cost of rule base construction, poor adaptability, and difficulty in expansion. 3. This invention employs steps (2-4) to (2-5), which utilize a dual-graph attention network and multimodal fusion feature extraction. It automatically mines deep features of complex faults through deep learning and achieves dynamic model optimization through incremental training. It can adapt to different brands and models of home appliances without retraining. Therefore, it can solve the technical problems of insufficient shallow feature characterization, weak generalization ability, difficulty in cross-brand and cross-model adaptation, and high cost of existing detection methods based on traditional machine learning. 4. Because the present invention adopts step (1), it covers multiple dimensions of data such as electrical, mechanical, and thermal distribution through multiple types of sensors and log reading modules, and constructs a complete operation data matrix. Therefore, it can solve the technical problems of existing detection methods based on single sensor data, such as single monitoring dimensions, easy to miss reports due to missing data, and inability to meet the detection needs of complex categories of household appliances. Attached Figure Description

[0015] Figure 1 This is a flowchart of the multimodal fault detection method for household appliances based on dual-graph attention network of the present invention; Figure 2 This is a network framework diagram of the fault detection model used in this invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0017] The basic idea of ​​this invention is to collect multi-dimensional data from home appliances through a multi-sensor layer (temperature, current, voltage, vibration, thermal imaging) and log reading. It utilizes a dual-graph attention network to extract temporal features, combines this with a multi-modal Transformer layer to fuse deep features, and employs an autoencoder for anomaly detection, providing a multi-modal fault detection method for home appliances based on a dual-graph attention network. This method covers all operating stages of home appliances, including startup, operation, and standby, through a multi-dimensional data acquisition module. It combines deep feature fusion technology to mine the value of data correlations, constructing a high-precision fault detection and prediction model. This enables accurate identification and early warning of potential faults in home appliances. The core objective is to specifically address various pain points of traditional technologies, providing users with a safer, smarter, and more worry-free home appliance experience.

[0018] like Figure 1 As shown, this invention provides a method for multimodal fault detection of household appliances based on a dual-graph attention network, comprising the following steps: (1) Acquire real-time operating data of household appliances through sensors; Specifically, the real-time operating data of household appliances includes triaxial acceleration timing signals collected by vibration sensors, current timing signals and voltage timing signals collected by electrical signal sensors, 128×128 resolution thermal images of household appliances collected by thermal imaging equipment, and log text data generated by the daily operation of household appliances.

[0019] (2) Input the real-time operating data of the household appliance obtained in step (1) into the pre-trained fault detection model to obtain the preliminary fault classification result of the household appliance; like Figure 2 As shown, the fault detection model of this invention adopts an architecture design that integrates a dual-graph attention network and a semi-supervised classification network. The overall structure consists of a three-layer core network structure, including an input layer, a hidden layer, and an output layer, as follows: The first layer is the input layer, which takes in the multimodal raw data generated during the operation of household appliances. This includes the triaxial acceleration time-series signal of the household appliances collected by the vibration sensor, the current time-series signal and voltage time-series signal of the household appliances collected by the electrical signal sensor, the 128×128 resolution thermal image of the household appliances collected by the thermal imaging device, and the log text data generated by the household appliances during daily operation. This layer filters and normalizes the multimodal data to obtain a 256×256 vibration time-frequency map, a time-series vector of a one-dimensional electrical time-series signal aligned to a 1000Hz sampling rate, a 128×128 thermal image, and a 64-dimensional log semantic vector, which together serve as the processed multimodal data.

[0020] Specifically, the process of filtering and normalizing the triaxial acceleration time-series signal of a household appliance to obtain a vibration time-frequency diagram with a dimension of 256×256 involves the following steps: First, bandpass filtering is performed on the original triaxial acceleration (X, Y, Z axes) time-series signal (the filtering frequency band is set to 0.5Hz-100Hz); then, min-max normalization is performed on the filtered composite acceleration time-series signal to map the signal amplitude to the [0,1] interval; finally, a short-time Fourier transform (SFT) is performed on the normalized composite acceleration time-series signal. The time-fourier transformation (STFT) sets the time window length to 256 sampling points, the window shift to 1 sampling point, and the frequency axis dimension to 256 to convert the one-dimensional composite acceleration time-series signal into a two-dimensional time-frequency matrix (each element in the time-frequency matrix corresponds to the vibration energy density at the time-frequency point). Subsequently, if the size of the transformed time-frequency matrix is ​​not 256×256, the time-frequency matrix is ​​normalized using bilinear interpolation to finally obtain and output a vibration time-frequency diagram with a dimension of 256×256.

[0021] The process of filtering and normalizing the current and voltage timing signals of household appliances to obtain a one-dimensional electrical timing vector aligned to a 1000Hz sampling rate is as follows: First, the current and voltage timing signals are low-pass filtered (the filter cutoff frequency is set to 400Hz, implemented using a fourth-order Chebyshev Type I filter). Then, the filtered current and voltage timing signals are resampled (i.e., the current and voltage timing signals are aligned to a 1000Hz sampling rate using linear interpolation, i.e., 1000 data points are collected per second to ensure consistent time resolution of electrical signals across different devices and time periods). Finally, the resampled current and voltage timing signals are normalized (eliminating the dimensional differences between current and voltage signals and mapping the data distribution to a standard normal distribution range). Finally, the normalized current and voltage timing signals are concatenated to obtain a one-dimensional electrical timing vector aligned to a 1000Hz sampling rate.

[0022] The process of filtering and normalizing a 128×128 resolution thermal image of a household appliance to obtain a 128×128 resolution thermal image is as follows: First, Gaussian filtering is used to denoise the resolution thermal image (eliminating sensor noise, environmental thermal radiation interference, and other issues); then, the denoised 128×128 thermal image is normalized to finally obtain a 128×128 resolution thermal image.

[0023] The process of filtering and normalizing the log text data generated by household appliances during daily operation to obtain a 64-dimensional log semantic vector is as follows: First, interference information (including garbled characters, special symbols, whitespace characters, etc.) is removed from the log text data, retaining only the valid text content (including device operating status descriptions, start / stop commands, error codes, runtime, etc.). Then, the log text data after removing interference information is standardized (i.e., timestamps of different formats are converted into a unified format, and error codes are semantically mapped to eliminate ambiguity). Next, the standardized log text data is downsampled (retaining logs at key time nodes to reduce data redundancy). Then, a lightweight text encoder is used to convert individual log texts in the downsampled log text data into initial semantic vectors. Finally, each encoded initial semantic vector is dimensionally compressed and normalized, and all adjusted initial semantic vectors constitute a 64-dimensional log semantic vector.

[0024] The second layer is a hidden layer, whose input is the multimodal data processed by the first layer. This layer first uses a specialized feature extraction sublayer to extract features from the multimodal data to obtain modal features. Then, it uses a sensor association graph attention sublayer and a time series graph attention sublayer to calculate the association graph and time series signal fusion features for each feature data. After that, it uses a multimodal Transformer fusion sublayer to fuse the weight information of all feature data to obtain fusion node features with a dimension of 256. Finally, it uses a semi-supervised classification joint sublayer to integrate the fusion node features to obtain a fusion feature vector with a dimension of 1×256, a fault probability distribution, and an initial risk value, and outputs them.

[0025] The third layer is the output layer, whose inputs are the fused feature vector output from the second layer, the fault probability distribution, and the initial risk value. This layer first uses an autoencoder to quantize the reconstruction error of the fused feature vector to obtain a scalar anomaly score in the 0-1 interval. Then, the fault probability distribution is subjected to probability calibration and classification correction using the Platt scaling algorithm to obtain a precise fault probability distribution in 1×11 dimensions. After that, the initial risk value is subjected to working condition weighting and risk classification to obtain a scalar risk value R in the 0-1 interval.

[0026] The third layer performs reconstruction error quantization on the fused feature vector to obtain a scalar anomaly score in the 0-1 interval. Specifically, the process is as follows: First, samples marked as normal operation are selected. Using the fused feature vectors of these samples, all reconstruction errors are statistically analyzed, and the 95th percentile of all reconstruction errors is taken as the anomaly judgment threshold. Finally, the reconstruction error between the fused feature vectors of all samples and the reconstruction output is calculated, and then the reconstruction error is divided by the anomaly judgment threshold to obtain the scalar anomaly score.

[0027] The third layer performs operating condition weighting and risk classification on the risk values ​​to obtain scalar risk values ​​in the 0-1 range. This process specifically involves: first, extracting the equipment service life from the risk values. Average runtime over the past 7 days Current operating load rate Ambient temperature and humidity deviation values These are four categories of key operating condition indicators; then, the basic weight of each category of key operating condition indicator is obtained, and the operating condition weight coefficient is obtained based on the basic weight. : ; in, Weighted by the equipment's service life, and has ; The weight is the average runtime over the past 7 days, and there is ; This is the weight of the current operating load rate of household appliances, which is equal to the operating load rate; The weight of the environmental temperature and humidity deviation value is directly taken as the normalized value of the deviation of the environmental temperature and humidity relative to the standard value. Subsequently, based on the working condition weighting coefficient Calculate the weighted scalar risk value : ; in, The initial risk value output by the hidden layer; Subsequently, the judgment Is it greater than 1 or less than 0? If it is the former, set the anomaly score to 1; if it is the latter, set the anomaly score to 0, to ensure that the weighted risk value is still in the 0-1 range.

[0028] Furthermore, the fault detection model of the present invention is established in the following manner: (2-1) Obtain real-time data of household appliances during operation, filter and normalize the obtained real-time data to obtain the processed real-time data as a dataset, and divide the dataset into training set, validation set and test set in a ratio of 7:2:1. Specifically, the filtering and normalization processes in this step are exactly the same as those in the first layer described above, and will not be repeated here.

[0029] (2-2) For each sample in the training set obtained in step (2-1), the sample is input into the first layer of the fault detection model to obtain the vibration time-frequency map with a dimension of 256×256, the time-series vector of the one-dimensional electrical time-series signal aligned to a sampling rate of 1000Hz, the thermal image with a dimension of 128×128, and the log semantic vector with a dimension of 64. Specifically, this step is exactly the same as the input layer processing described above, and will not be repeated here.

[0030] The advantage of steps (2-1) to (2-2) above is that the data processing logic in actual applications is directly reused, the data format is unified, no additional model adaptation is required, and the generalization ability of the model for different household appliances is increased.

[0031] (2-3) For each sample in the training set obtained in step (2-1), the vibration time-frequency map with a dimension of 256×256, the time-series vector of the one-dimensional electrical time-series signal aligned to a sampling rate of 1000Hz, the thermal image with a dimension of 128×128, and the 64-dimensional log semantic vector corresponding to the sample obtained in step (2-2) are input into the special feature extraction sublayer in the hidden layer of the fault detection model to obtain the special features corresponding to the sample. Specifically, to ensure the accuracy of feature extraction, cross-entropy loss and mean squared error are initially set as the basic loss indicators for time-frequency plots and time-series vector features, respectively.

[0032] (2-4) For each sample in the training set obtained in step (2-1), the specific features of each sample obtained in step (2-3) are input into the sensor association graph attention sublayer in the hidden layer to obtain the association graph corresponding to the sample, and the specific features of the sample are input into the time series graph attention sublayer in the hidden layer for feature fusion to obtain the time series signal fusion features corresponding to the sample. Specifically, the above process is as follows: First, for the 256×256 vibration time-frequency map in the sample, the specific features of the vibration time-frequency map are flattened and mapped to 128 dimensions through a 1×1 convolutional layer to obtain a unified 128-dimensional vibration time-frequency map time-series vector; for the electrical time-series signal in the sample, the specific features of the electrical time-series signal are mapped to 128 dimensions through a 1×1 convolutional layer to obtain a unified 128-dimensional electrical time-series signal time-series vector; for the thermal image in the sample, it is flattened into a 1×128-dimensional vector through global average pooling; for the log semantic vector in the sample, the log semantic vector is extended into a 1×128-dimensional vector through linear mapping. Then, obtain the sample's first... time series vectors and the time series vectors The degree of correlation between them: ; in ∈[1, total number of time vectors of the samples], ∈[1, the total number of time-series vectors of the samples], and have ≠ , The learnable weight matrix (initially randomly generated, optimized during training). For the first The time series vector and the first The result of concatenating several time-series vectors Here, d is the activation function, and d is the dimension of the time series vector (here). ); Subsequently, only all specific feature pairs with a correlation greater than 0.6 are retained, and all retained specific feature pairs constitute the correlation graph corresponding to the sample. Then, the vibration time-frequency diagram time-series vector and the electrical time-series signal time-series vector, both unified to 128 dimensions, corresponding to this sample are concatenated to form a 256-dimensional time-series feature vector. This time-series feature vector is then divided into... (Generally take) (Round up if not divisible) Sub-time series feature segments , ..., And calculate each sub-temporal feature segment. With the Individual time-series feature fragments Time weight : ; in, For the first The dimension of each sub-temporal feature segment For the first The time weight matrix of each sub-temporal feature segment For the first The first sub-temporal feature fragment and the first Vector concatenation of sub-temporal feature segments, t∈[1, ].

[0033] Then, according to the first The first sub-temporal feature fragment and the first Temporal weights of individual time-series feature segments For the first Individual time-series feature fragments Update to obtain the updated version. Individual time-series feature fragments : ; in, For the first Bias of individual time-series feature segments Use the Sigmoid activation function; Finally, a 1×1 convolutional layer is used to update the first... Individual time-series feature fragments Mapped to 128 dimensions, all mapped sub-temporal feature fragments constitute the temporal signal fusion feature corresponding to the sample. (2-5) For each sample in the training set obtained in step (2-1), the association graph and time-series signal fusion features corresponding to the sample obtained in step (2-4) and the special features corresponding to the sample obtained in step (2-3) are input into the multimodal Transformer fusion sublayer in the hidden layer to perform deep aggregation of cross-modal features, so as to obtain the fusion feature vector with a dimension of 1×256 corresponding to the sample.

[0034] The above process is as follows: First, for the thermal image in the sample, the specific features corresponding to the thermal image are flattened into a 1×128-dimensional vector through global average pooling. For the electrical time-series signal in the sample, the specific features corresponding to the electrical time-series signal are compressed into a 1×128-dimensional vector through a 1×1 convolutional layer. For the vibration time-frequency map in the sample, the specific features corresponding to the vibration time-frequency map are transformed from a two-dimensional matrix into a 1×128-dimensional vector through convolutional and pooling layers. For the log semantic vector in the sample, the specific features corresponding to the log semantic vector are extended into a 1×128-dimensional vector through linear mapping. For the association graph generated by the sample in step (2-4), the association graph is mapped into a 1×128-dimensional vector through a linear layer. ; Then, using the 128-dimensional vector corresponding to the specific features of the thermal image. query vector The vector corresponding to the specific features of the time-series vector of the electrical time-series signal. as a key vector Sum value vector Attention weights are calculated and fused to obtain the fused features of thermal images and electrical time-series signals. : ; in The dimension is the key vector, which is the dimension in this invention. ; Then, the 128-dimensional vector corresponding to the vibration time-frequency map fusion feature is used. as query vector The vector corresponding to the log semantic features as a key vector Sum value vector Attention weights are calculated and fused to obtain the fused features of the vibration time-frequency map and log semantic vector. (The calculation formula is exactly the same as the one in the previous paragraph).

[0035] Subsequently, the fusion features of the obtained thermal image and electrical time-series signal are analyzed. Fusion features of vibration time-frequency graph and log semantic vector and the association graph obtained in steps (2-3) Fusion vector of temporal features Perform cross-attention fusion to generate intermediate features. : ; ; in For vector concatenation, This represents the concatenated vector. () is a linear transformation function.

[0036] Finally, the intermediate fusion features With a learnable weight matrix (whose initial values ​​are randomly generated and optimized during training). Weighted integration is performed to obtain a fused feature vector with a dimension of 1×256 corresponding to the sample.

[0037] The advantage of steps (2-3) to (2-5) above is that features are extracted from multimodal data, the correlation between features is found, and useful information is fused hierarchically after unifying the feature dimensions, which improves the feature capture capability of the model.

[0038] (2-6) For each sample in the training set obtained in step (2-1), the fusion feature vector with a dimension of 1×256 corresponding to the sample obtained in step (2-5) is input into the semi-supervised classification joint sub-layer to obtain the fault probability distribution and initial risk value with a dimension of 1×11 corresponding to the sample.

[0039] The above processing procedure is as follows: First, the fused feature vector of the samples is input into the pre-trained autoencoder to obtain anomaly scores: ; in This represents the anomaly detection threshold, and its value ranges from 0.9 to 1, preferably 0.95. This represents the first feature vector corresponding to the sample. One element, The first sample reconstruction output represents the... There are elements, where x represents a sample; Then, the fused feature vector and anomaly score of the sample are input into the fully connected layer to obtain a fault probability distribution with a dimension of 1×11 (achieving preliminary classification of fault types). Subsequently, the initial risk value of the sample is calculated using the anomaly score and the failure probability distribution. ; in The maximum value in the fault probability distribution with dimension 1×11 is represented by the initial quantitative index reflecting the potential risk of the fault, which is generated by balancing the degree of anomaly and the fault confidence through weight allocation.

[0040] (2-7) For each sample in the training set obtained in step (2-1), based on the initial risk value obtained in step (2-6), obtain the scalar risk value in the 0-1 interval corresponding to that sample. ; Specifically, this step is exactly the same as the processing procedure of the output layer described above, and will not be repeated here.

[0041] (2-8) For each sample in the training set obtained in step (2-1), the scalar risk value of the corresponding 0-1 interval obtained in step (2-7) is calculated. The cross-entropy loss for the fault classification corresponding to the sample is obtained from the fault probability distribution obtained in steps (2-6). ; Specifically, this step uses the following formula: ; in, The th sample is the th in the true failure probability distribution. One element, For the model, the fault probability distribution with dimension 1×11 is the first... One element; (2-9) For each sample in the training set obtained in step (2-1), the cross-entropy loss corresponding to the fault classification of that sample is obtained in step (2-8). Obtain the total loss corresponding to this sample. : Specifically, the total loss equals: ; in The cross-entropy loss for fault classification is used to optimize the accuracy of the classification results. This is a penalty item for abnormal scores.

[0042] (2-10) For each sample in the training set obtained in step (2-1), based on the total loss corresponding to that sample obtained in step (2-11), the fault detection model is iteratively trained using gradient descent until the fault detection model reaches a preset number of iterations (100 times in this invention), or the total loss... The fluctuation range is less than The model is trained until the total loss of the validation set does not increase significantly, or the accuracy of fault classification is higher than 90% and the F1 score of anomaly detection is greater than 88%, and the optimal parameters of the fault detection model at this time are obtained, thus obtaining the initially trained fault detection model. (2-11) Use the test set obtained in step (2-1) to test the fault detection model initially trained in step (2-10) until the detection accuracy reaches the optimal level, so as to obtain the final trained fault detection model.

[0043] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for multimodal fault detection of household appliances based on a dual-graph attention network, characterized in that, Includes the following steps: (1) Acquire real-time operating data of household appliances through sensors; (2) Input the real-time operating data of the household appliance obtained in step (1) into the pre-trained fault detection model to obtain the preliminary fault classification result of the household appliance.

2. The method for multimodal fault detection of household appliances based on dual-graph attention network according to claim 1, characterized in that, Real-time operating data of household appliances includes triaxial acceleration timing signals collected by vibration sensors, current timing signals and voltage timing signals collected by electrical signal sensors, 128×128 resolution thermal images of household appliances collected by thermal imaging equipment, and log text data generated by daily operation of household appliances.

3. The method for multimodal fault detection of household appliances based on dual-graph attention network according to claim 1 or 2, characterized in that, The fault detection model adopts an architecture that integrates a dual-graph attention network and a semi-supervised classification network, and includes an input layer, a hidden layer, and an output layer, with the following structure: The first layer is the input layer, which takes in the multimodal raw data generated during the operation of household appliances. This includes the triaxial acceleration time-series signal of the household appliances collected by the vibration sensor, the current time-series signal and voltage time-series signal of the household appliances collected by the electrical signal sensor, the 128×128 resolution thermal image of the household appliances collected by the thermal imaging device, and the log text data generated by the household appliances during daily operation. This layer filters and normalizes the multimodal data to obtain a 256×256 vibration time-frequency diagram, a time-series vector of a one-dimensional electrical time-series signal aligned to a 1000Hz sampling rate, a 128×128 thermal image, and a 64-dimensional log semantic vector, which together serve as the processed multimodal data. The second layer is a hidden layer, whose input is the multimodal data processed by the first layer. This layer first uses a specific feature extraction sublayer to extract features from the multimodal data to obtain modal features; then, it uses a sensor correlation graph attention sublayer and a time series graph attention sublayer to calculate the correlation graph and time series signal fusion features of each feature data. Subsequently, the weight information of all feature data is fused using a multimodal Transformer fusion sublayer to obtain fused node features with a dimension of 256; finally, the fused node features are integrated using a semi-supervised classification joint sublayer to obtain a fused feature vector with a dimension of 1×256, a fault probability distribution, and an initial risk value, which are then output. The third layer is the output layer, whose inputs are the fused feature vector output by the second layer, the fault probability distribution, and the initial risk value. This layer first uses an autoencoder to perform reconstruction error quantization on the fused feature vector to obtain a scalar anomaly score in the 0-1 interval. Then, the Platt scaling algorithm is used to perform probability calibration and classification correction on the fault probability distribution to obtain a precise fault probability distribution in 1×11 dimensions. Subsequently, the initial risk value is subjected to working condition weighting and risk classification to obtain a scalar risk value R in the 0-1 interval.

4. The method for multimodal fault detection of household appliances based on a dual-graph attention network according to any one of claims 1 to 3, characterized in that, The process of filtering and normalizing the triaxial acceleration time-series signal of household appliances to obtain a vibration time-frequency diagram with a dimension of 256×256 is as follows: First, the original triaxial acceleration time-series signal is bandpass filtered; then, the composite acceleration time-series signal obtained after filtering is normalized by min-max to map the signal amplitude to the [0,1] interval. Finally, a short-time Fourier transform (STFT) is performed on the normalized composite acceleration time-series signal. The time window length is set to 256 sampling points, the window shift is 1 sampling point, and the frequency axis dimension is determined to be 256. This converts the one-dimensional composite acceleration time-series signal into a two-dimensional time-frequency matrix. Subsequently, if the size of the transformed time-frequency matrix is ​​not 256×256, the time-frequency matrix is ​​normalized using bilinear interpolation. Finally, a vibration time-frequency diagram with a dimension of 256×256 is obtained and output. The process of filtering and normalizing the current and voltage timing signals of household appliances to obtain a timing vector of a one-dimensional electrical timing signal aligned to a 1000Hz sampling rate is as follows: First, the current and voltage timing signals are low-pass filtered respectively; then, the filtered current and voltage timing signals are resampled; finally, the resampled current and voltage timing signals are normalized respectively; and finally, the normalized current and voltage timing signals are concatenated to obtain a timing vector of a one-dimensional electrical timing signal aligned to a 1000Hz sampling rate. The process of filtering and normalizing a 128×128 resolution thermal image of a household appliance to obtain a 128×128 resolution thermal image is as follows: First, Gaussian filtering is used to denoise the resolution thermal image; then, the denoised 128×128 thermal image is normalized to finally obtain a 128×128 resolution thermal image. The process of filtering and normalizing the log text data generated by household appliances during daily operation to obtain a 64-dimensional log semantic vector is as follows: First, interference information in the log text data is removed, retaining only the valid text content; then, the log text data after removing interference information is standardized; next, the standardized log text data is downsampled; then, a lightweight text encoder is used to convert individual log texts in the downsampled log text data into initial semantic vectors; finally, each encoded initial semantic vector is dimensionally compressed and normalized, and all adjusted initial semantic vectors constitute a 64-dimensional log semantic vector.

5. The method for multimodal fault detection of household appliances based on dual-graph attention network according to claim 4, characterized in that, The process of quantizing the reconstruction error of the fused feature vector to obtain the scalar anomaly score in the 0-1 interval is as follows: First, samples marked as normal operation are selected, and the fused feature vectors of these samples are used to count all reconstruction errors. The 95th percentile of all reconstruction errors is taken as the anomaly judgment threshold. Finally, the reconstruction error between the fused feature vector of all samples and the reconstruction output is calculated, and the reconstruction error is divided by the anomaly judgment threshold to obtain the scalar anomaly score. The risk values ​​are weighted by operating conditions and classified by risk level to obtain scalar risk values ​​in the 0-1 range. This process is as follows: First, extract the equipment service life from the risk value. Average runtime over the past 7 days Current operating load rate Ambient temperature and humidity deviation values These are four categories of key operating condition indicators; then, the basic weight of each category of key operating condition indicator is obtained, and the operating condition weight coefficient is obtained based on the basic weight. : ; in, Weighted by the equipment's service life, and has ; The weight is the average runtime over the past 7 days, and there is ; This is the weight of the current operating load rate of household appliances, which is equal to the operating load rate; The weight of the environmental temperature and humidity deviation value is directly taken as the normalized value of the deviation of the environmental temperature and humidity relative to the standard value. Subsequently, based on the working condition weighting coefficient Calculate the weighted scalar risk value : ; in, The initial risk value output by the hidden layer; Subsequently, the judgment Is it greater than 1 or less than 0? If it is the former, set the anomaly score to 1; if it is the latter, set the anomaly score to 0, to ensure that the weighted risk value is still in the 0-1 range.

6. The method for multimodal fault detection of household appliances based on dual-graph attention network according to claim 5, characterized in that, The fault detection model is established in the following way: (2-1) Obtain real-time data of household appliances during operation, filter and normalize the obtained real-time data to obtain the processed real-time data as a dataset, and divide the dataset into training set, validation set and test set in a ratio of 7:2:

1. (2-2) For each sample in the training set obtained in step (2-1), the sample is input into the first layer of the fault detection model to obtain the vibration time-frequency map with a dimension of 256×256, the time-series vector of the one-dimensional electrical time-series signal aligned to a sampling rate of 1000Hz, the thermal image with a dimension of 128×128, and the log semantic vector with a dimension of 64. (2-3) For each sample in the training set obtained in step (2-1), the vibration time-frequency map with a dimension of 256×256, the time-series vector of the one-dimensional electrical time-series signal aligned to a sampling rate of 1000Hz, the thermal image with a dimension of 128×128, and the 64-dimensional log semantic vector corresponding to the sample obtained in step (2-2) are input into the special feature extraction sublayer in the hidden layer of the fault detection model to obtain the special features corresponding to the sample. (2-4) For each sample in the training set obtained in step (2-1), the specific features of each sample obtained in step (2-3) are input into the sensor association graph attention sublayer in the hidden layer to obtain the association graph corresponding to the sample, and the specific features of the sample are input into the time series graph attention sublayer in the hidden layer for feature fusion to obtain the time series signal fusion features corresponding to the sample. (2-5) For each sample in the training set obtained in step (2-1), the association graph and time-series signal fusion features corresponding to the sample obtained in step (2-4) and the special features corresponding to the sample obtained in step (2-3) are input into the multimodal Transformer fusion sublayer in the hidden layer to perform deep aggregation of cross-modal features, so as to obtain the fusion feature vector with a dimension of 1×256 corresponding to the sample. (2-6) For each sample in the training set obtained in step (2-1), the fusion feature vector with a dimension of 1×256 corresponding to the sample obtained in step (2-5) is input into the semi-supervised classification joint sub-layer to obtain the fault probability distribution and initial risk value with a dimension of 1×11 corresponding to the sample. (2-7) For each sample in the training set obtained in step (2-1), based on the initial risk value obtained in step (2-6), obtain the scalar risk value in the 0-1 interval corresponding to that sample. ; (2-8) For each sample in the training set obtained in step (2-1), the scalar risk value of the corresponding 0-1 interval obtained in step (2-7) is calculated. The cross-entropy loss for the fault classification corresponding to the sample is obtained from the fault probability distribution obtained in steps (2-6). : ; in The th sample is the th in the true failure probability distribution. One element, For the model, the fault probability distribution with dimension 1×11 is the first... One element; (2-9) For each sample in the training set obtained in step (2-1), the cross-entropy loss corresponding to the fault classification of that sample is obtained in step (2-8). Obtain the total loss corresponding to this sample. : ; in Cross-entropy loss for fault classification, This is a penalty item for abnormal scores; (2-10) For each sample in the training set obtained in step (2-1), based on the total loss corresponding to that sample obtained in step (2-11), the fault detection model is iteratively trained using gradient descent until the fault detection model reaches the preset number of iterations, or the total loss is reached. The fluctuation range is less than The model is trained until the total loss of the validation set does not increase significantly, or the accuracy of fault classification is higher than 90% and the F1 score of anomaly detection is greater than 88%, and the optimal parameters of the fault detection model at this time are obtained, thus obtaining the initially trained fault detection model. (2-11) Use the test set obtained in step (2-1) to test the fault detection model initially trained in step (2-10) until the detection accuracy reaches the optimal level, so as to obtain the final trained fault detection model.

7. The method for multimodal fault detection of household appliances based on dual-graph attention network according to claim 6, characterized in that, Steps (2-4) are as follows: First, for the 256×256 vibration time-frequency map in the sample, the specific features of the vibration time-frequency map are flattened and mapped to 128 dimensions through a 1×1 convolutional layer to obtain a unified 128-dimensional vibration time-frequency map time-series vector. For the electrical time-series signal in the sample, the specific features of the electrical time-series signal are mapped to 128 dimensions through a 1×1 convolutional layer to obtain a unified 128-dimensional electrical time-series signal time-series vector. For the thermal image in the sample, it is flattened into a 1×128-dimensional vector through global average pooling. For the log semantic vector in the sample, the log semantic vector is extended into a 1×128-dimensional vector through linear mapping. Then, obtain the sample's first... time series vectors and the time series vectors The degree of correlation between them: ; in ∈[1, total number of time vectors of the samples], ∈[1, the total number of time-series vectors of the samples], and have ≠ , For learnable weight matrix, For the first The time series vector and the first The result of concatenating several time-series vectors Here, d is the activation function, and d is the dimension of the time-series vector. Subsequently, only all specific feature pairs with a correlation greater than 0.6 are retained, and all retained specific feature pairs constitute the correlation graph corresponding to the sample. Then, the vibration time-frequency diagram time-series vector and the electrical time-series signal time-series vector, both unified to 128 dimensions, corresponding to this sample are concatenated to form a 256-dimensional time-series feature vector. This time-series feature vector is then divided into... Individual time-series feature fragments , ..., And calculate each sub-temporal feature segment. With the Individual time-series feature fragments Time weight : ; in, For the first The dimension of each sub-temporal feature segment For the first The time weight matrix of each sub-temporal feature segment For the first The first sub-temporal feature fragment and the first Vector concatenation of sub-temporal feature segments, t∈[1, ]; Then, according to the first The first sub-temporal feature fragment and the first Temporal weights of individual time-series feature segments For the first Individual time-series feature fragments Update to obtain the updated version. Individual time-series feature fragments : ; in, For the first Bias of individual time-series feature segments Use the Sigmoid activation function; Finally, a 1×1 convolutional layer is used to update the first... Individual time-series feature fragments Mapped to 128 dimensions, all mapped sub-temporal feature fragments constitute the temporal signal fusion features corresponding to the sample.

8. The method for multimodal fault detection of household appliances based on dual-graph attention network according to claim 7, characterized in that, Steps (2-5) are as follows: First, for the thermal image in the sample, the specific features corresponding to the thermal image are flattened into a 1×128-dimensional vector using global average pooling. For the electrical time-series signal in the sample, the specific features corresponding to the electrical time-series signal are compressed into a 1×128-dimensional vector through a 1×1 convolutional layer. For the vibration time-frequency map in the sample, the specific features corresponding to the vibration time-frequency map are transformed from a two-dimensional matrix into a 1×128-dimensional vector through convolutional and pooling layers. For the log semantic vector in the sample, the specific features corresponding to the log semantic vector are extended into a 1×128-dimensional vector through linear mapping. For the association graph generated by the sample in step (2-4), the association graph is mapped into a 1×128-dimensional vector through a linear layer. ; Then, using the 128-dimensional vector corresponding to the specific features of the thermal image. query vector The vector corresponding to the specific features of the time-series vector of the electrical time-series signal. as a key vector Sum value vector Attention weights are calculated and fused to obtain the fused features of thermal images and electrical time-series signals. : ; in The dimension of the key vector; Then, the 128-dimensional vector corresponding to the vibration time-frequency map fusion feature is used. as query vector The vector corresponding to the log semantic features as a key vector Sum value vector Attention weights are calculated and fused to obtain the fused features of the vibration time-frequency map and log semantic vector. ; Subsequently, the fusion features of the obtained thermal image and electrical time-series signal are analyzed. Fusion features of vibration time-frequency graph and log semantic vector and the association graph obtained in steps (2-3) Fusion vector of temporal features Perform cross-attention fusion to generate intermediate features. : ; ; in For vector concatenation, This represents the concatenated vector. () represents a linear transformation function; Finally, the intermediate fusion features With learnable weight matrix Weighted integration is performed to obtain a fused feature vector with a dimension of 1×256 corresponding to the sample.

9. The method for multimodal fault detection of household appliances based on dual-graph attention network according to claim 8, characterized in that, Steps (2-6) are as follows: First, the fused feature vector of the samples is input into the pre-trained autoencoder to obtain anomaly scores: ; in This represents the anomaly detection threshold, and its value ranges from 0, 9 to 1. This represents the first feature vector corresponding to the sample. One element, The first sample reconstruction output represents the... There are elements, where x represents a sample; Then, the fused feature vector and anomaly score of the samples are input into the fully connected layer to obtain a fault probability distribution with a dimension of 1×11; Subsequently, the initial risk value of the sample is calculated using the anomaly score and the failure probability distribution. ; in This represents the maximum value in a fault probability distribution with dimension 1×11.

10. A method for multimodal fault detection of household appliances based on a dual-graph attention network, characterized in that, include: The first module is used to acquire real-time operating data of home appliances through sensors; The second module is used to input the real-time operating data of the household appliances obtained by the first module into the pre-trained fault detection model in order to obtain the preliminary fault classification results of the household appliances.