Fault Diagnosis Method and Equipment for Washing and Screening Vibrating Screen Based on Multimodal Fusion
By extracting features and transforming probability distributions from multimodal data of washing and screening vibrating screen equipment, and combining them with evidence theory, the problem of collaborative analysis of multi-source heterogeneous data was solved, thereby improving the credibility and interpretability of fault diagnosis results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CCTEG COAL IND PLANNING INSTITUTE CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-26
AI Technical Summary
In the operation of washing and screening vibrating screen equipment, it is difficult to conduct collaborative analysis of multi-source heterogeneous data, there is a lack of unified representation of features between modes, multi-modal fusion lacks credibility and interpretability, and the model's generalization ability and stability are insufficient, resulting in one-sided and unreliable fault diagnosis results.
By extracting features from source data of different modalities and converting them into probability distributions, performing cross-modal semantic alignment in the probability space, and fusing them using evidence theory methods, fault probability distribution and confidence information are obtained. Uncertainty factors are introduced to improve the credibility and interpretability of diagnostic results.
It improves the accuracy and robustness of fault prediction results, ensures the reliability and interpretability of fault prediction results, solves the problems of low information utilization and one-sided results, and enhances the stability of the model in noisy environments.
Smart Images

Figure CN121524893B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent diagnostic technology for washing and screening vibrating screen equipment, and in particular to a fault diagnosis method and equipment for washing and screening vibrating screen equipment based on multimodal fusion. Background Technology
[0002] The washing and screening vibrating screen equipment has a complex structure and operates in a harsh and complex environment. It is affected by extreme factors such as high load, strong impact, high dust and high humidity, resulting in the following prominent problems in the monitoring of its operating status and fault diagnosis:
[0003] I. Difficulty in Collaborative Analysis of Multi-Source Heterogeneous Data: During operation, washing and screening vibrating screen equipment generates various types of data, including vibration, temperature, current, power, images, and text. The sampling frequency, timing structure, noise characteristics, and data dimensions differ significantly between different modes. Traditional methods typically analyze only a single signal, resulting in low information utilization, incomplete diagnostic results, and an inability to fully reflect the true operating condition of the equipment.
[0004] Second, there is a lack of unified representation for features between modalities: Traditional deep learning diagnostic models often embed features of each modality as fixed point vectors, ignoring the differences in statistical distribution and uncertainty between data, making it difficult to model the dynamic changes under complex coal mine conditions, resulting in insufficient robustness of the model in different scenarios.
[0005] Third, multimodal fusion lacks credibility and interpretability: Existing multimodal fusion methods generally rely on static weights or simple concatenation, without considering the conflicts and confidence differences between modal information. When some sensors are abnormal or the acquisition noise is large, the model is prone to "erroneous amplification" phenomenon, lacking a quantitative interpretation mechanism for the credibility of the results and the modal contribution.
[0006] IV. Insufficient Model Generalization Ability and Stability: The operating conditions of washing and screening vibrating screen equipment are complex and volatile. Traditional models show significant performance degradation under conditions of small sample size, high noise, and environmental disturbances. At the same time, deep models have many parameters and rely on a large amount of labeled data, making it difficult to achieve rapid deployment and adaptive updates in coal mines. Summary of the Invention
[0007] At least one aspect and advantage of this application will be set forth in part in the description which follows, or may be apparent from the description, or may be obtained by practicing the subject matter of this disclosure.
[0008] According to the first aspect of this application, a fault diagnosis method for a washing and screening vibrating screen based on multimodal fusion is provided, the method comprising:
[0009] Feature extraction is performed on source data of different modalities to obtain data features corresponding to different modalities;
[0010] The data features corresponding to different modalities are transformed to obtain the probability distributions corresponding to different modalities;
[0011] Semantic alignment of probability distributions corresponding to different modalities across the probability space is performed to obtain the correlation between probability distributions corresponding to different modalities;
[0012] The probability distributions of different modes are fused based on evidence theory methods to obtain the fault probability distributions and confidence information associated with different modes;
[0013] The data of different modes include vibration signal data, temperature signal, power signal, image data, and text data associated with the washing and screening vibrating screen equipment.
[0014] According to one embodiment of this application, before fusing the probability distributions of different modes, the features of the same mode under different noise are forced to align, and the features of the same mode under different noise are balanced and compressed.
[0015] According to one embodiment of this application, when performing balanced feature compression, the hyperparameter value ranges from 0.005 to 0.05.
[0016] According to one embodiment of this application, when the source data is a vibration signal, the corresponding data features are obtained by processing the vibration signal after short-time Fourier transform by a multi-scale convolutional neural network;
[0017] When the source data is a temperature or power signal, the corresponding data features are obtained based on a hybrid recurrent neural network;
[0018] When the source data is image data, the corresponding data features are obtained based on a convolutional neural network;
[0019] When the source data is text data, the corresponding data features are obtained based on the Transformer network.
[0020] According to one embodiment of this application, the data features corresponding to the vibration signal include a first periodic trend feature and a second periodic trend feature;
[0021] The sampling period corresponding to the first periodic trend is greater than the sampling period corresponding to the second periodic trend.
[0022] According to one embodiment of this application, the probability distribution corresponding to different modes is a Gaussian distribution obtained based on the deterministic eigenvector, mean vector and covariance matrix of each mode.
[0023] According to one embodiment of this application, when performing cross-modal semantic alignment of probability distributions corresponding to different modalities in probability space, the selected temperature hyperparameter ranges from 0.05 to 0.3.
[0024] According to one embodiment of this application, the sampling frequency of the vibration signal data is 8-12kHz, and the sampling frequency of the temperature signal and power signal is 5-15Hz.
[0025] According to a second aspect of this application, a fault diagnosis device for a washing and screening vibrating screen based on multimodal fusion is provided, the device comprising:
[0026] The data feature acquisition module is used to extract features from source data of different modalities to obtain data features corresponding to different modalities.
[0027] The probability distribution determination module is used to transform the data features corresponding to different modes to obtain the probability distributions corresponding to different modes.
[0028] The modal probability association module is used to perform cross-modal semantic alignment of probability distributions corresponding to different modalities in the probability space, and obtain the association between probability distributions corresponding to different modalities;
[0029] The modal probability fusion module is used to fuse the probability distributions of different modes based on evidence theory methods to obtain the fault probability distribution and confidence information associated with different modes.
[0030] The data of different modes include vibration signal data, temperature signal, power signal, image data, and text data associated with the washing and screening vibrating screen equipment.
[0031] According to a third aspect of this application, an electronic device is provided, comprising a memory and a processor, the memory being coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the steps of the method described above.
[0032] According to a fourth aspect of this application, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the steps of the above-described method.
[0033] This application embodiment extracts features from source data of different modalities, obtains data features corresponding to different modalities, and converts them into probability distributions. The probability distributions corresponding to different modalities are then semantically aligned across modalities in the probability space. After obtaining the correlation between the probability distributions corresponding to different modalities, evidence theory methods are used to fuse them to obtain fault probability distributions and confidence information associated with different modalities. This method of fault diagnosis using multimodal data solves the problems of low information utilization and one-sided results caused by using single data for fault prediction. On the other hand, by introducing uncertainty factors through probability distribution, it solves the problem of low reliability of fault prediction results caused by unreliable data in multimodal data fusion, ensuring the credibility and interpretability of fault prediction results and improving the accuracy and robustness of fault prediction results. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0035] Figure 1 This is a flowchart illustrating a fault diagnosis method for a washing and screening vibrating screen based on multimodal fusion, provided in one embodiment of this application.
[0036] Figure 2 A flowchart illustrating a fault diagnosis method for a washing and screening vibrating screen based on multimodal fusion, provided for another embodiment of this application;
[0037] Figure 3 A schematic diagram of the two-dimensional time spectrum of the vibration signal in a fault diagnosis method for a washing and screening vibrating screen equipment based on multimodal fusion provided in another embodiment of this application;
[0038] Figure 4 A schematic diagram of the data characteristics of vibration signals in a fault diagnosis method for washing and screening vibrating screen equipment based on multimodal fusion, provided in another embodiment of this application;
[0039] Figure 5 The output result of a fault diagnosis method for a washing and screening vibrating screen equipment based on multimodal fusion provided in another embodiment of this application;
[0040] Figure 6 This is a block diagram of a fault diagnosis device for a washing and screening vibrating screen based on multimodal fusion, provided as an embodiment of this application. Detailed Implementation
[0041] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0042] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0043] It should be understood that although this application may use the terms first, second, third, etc., to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0044] According to one embodiment of this application, a fault diagnosis method for washing and screening vibrating screen equipment based on multimodal fusion is provided, such as... Figure 1 As shown, the method includes steps S101 to S104.
[0045] Step S101: Extract features from the source data of different modes to obtain the data features corresponding to different modes; the source data of different modes includes vibration signal data, temperature signal, power signal, image data and text data associated with the washing and screening vibrating screen equipment.
[0046] It should be noted that the source data of different modalities refers to the raw data obtained after data acquisition and synchronization using data acquisition devices (such as sensors).
[0047] Because source data from different modalities have different structures and physical meanings, different feature extractors are needed to extract the core features that best characterize the device's state from the raw data. For example, vibration signal data contains high-frequency and low-frequency components, requiring the capture of features at different time scales; a multi-scale convolutional neural network (CNN) can be used for this purpose. Temperature and power signals are time-series data; a hybrid recurrent neural network (such as LSTM and GRU) capable of handling long-term dependencies can be used for this purpose. For image data, a CNN with spatial channel attention mechanism can be used, while a Transformer network can be used for text data. It should be noted that the above feature extractors are exemplary algorithm models for feature extraction from source data of different modalities. In practical applications, other models can be used depending on business needs; these are not listed here. By using feature extractors with different algorithm models, deep and automated feature learning of different modalities can be achieved, thereby improving feature extraction efficiency.
[0048] Step S102: Transform the data features corresponding to different modes to obtain the probability distributions corresponding to different modes.
[0049] It should be noted that the data features in this embodiment refer to the deterministic feature vectors obtained by processing the source data of different modules using a feature extractor, which are essentially point estimations.
[0050] It should be noted that probability distribution refers to representing the features of different modalities in the form of a range of values. Compared with the method of representing data reliability with fixed values, it can improve the reliability of data in noisy environments and provide reliable support data for subsequent fusion. This embodiment transforms the data features corresponding to different modalities through a probability distribution layer.
[0051] In practice, the probability distribution can be Gaussian, Laplace, exponential, etc.
[0052] This embodiment uses a Gaussian distribution, which is described by the mean of possible feature values and the variance that quantifies the uncertainty of the feature. It has the function of explicitly describing the average trend and variance change of the data. Since the probability distribution can quantify the reliability of the feature, this can avoid the problem of fault diagnosis errors caused by over-reliance on unreliable features in the subsequent fusion process.
[0053] Step S103: Perform cross-modal semantic alignment of the probability distributions corresponding to different modalities in the probability space to obtain the correlation between the probability distributions corresponding to different modalities.
[0054] The probability space in this embodiment refers to an abstract, mathematical semantic representation space in which the features of each modality are mapped to a probability distribution. Through this mapping, the features (probability distributions) of different modalities under the same working condition can be represented by the same semantic vector to represent the same equipment state, providing reliable data support for the subsequent fusion of features of different modalities under the same working condition and improving cross-modal collaboration capabilities.
[0055] In practice, the probability distributions of different modalities are used as input. A probability space is learned during model training through an introduced probabilistic contrastive learning mechanism. In this probability space, positive samples (i.e., different modal distributions describing the same equipment state) are close together, while negative samples (i.e., different modal distributions describing different equipment states) are far apart. Since the probability distribution learns the semantic features of the operating conditions, even if the original data for the same operating condition changes, its semantics remain unchanged, and its positions in the aligned space are similar, providing reliable data support for the subsequent accurate identification of faults in the washing and screening vibrating screen equipment.
[0056] Step S104: The probability distributions of different modes are fused based on evidence theory methods to obtain the fault probability distribution and confidence information associated with different modes.
[0057] It should be noted that evidence theory refers to a mathematical framework for dealing with uncertainty. Its basic unit is the Basic Probability Assignment (BPA). Essentially, it converts the probability of each modality into a BPA function, uses the fusion algorithm of evidence theory to fuse the BPA functions of different modalities, and obtains a fused global BPA function. Based on the global BPA function, the belief and plausibility of each fault hypothesis are calculated, thereby obtaining the fault probability distribution and confidence information.
[0058] This embodiment uses a fusion algorithm based on DS combination rules, which first transforms the probability distribution and uncertainty of each mode into a basic probability assignment function for the fault hypothesis set, and then uses DS combination rules to make decision output.
[0059] It should be noted that the fusion of probability distributions of different modalities can be achieved using any fusion algorithm based on evidence theory, such as the DS combination rule, Yager combination rule, Dubois-Prade combination rule, weighted evidence combination method, discounted DS combination rule, PCR6 rule, etc. There is no limitation here, and the specific choice can be made according to the applicability of fault diagnosis.
[0060] This application embodiment extracts features from source data of different modalities, obtains data features corresponding to different modalities, and converts them into probability distributions. The probability distributions corresponding to different modalities are then semantically aligned across modalities in the probability space. After obtaining the correlation between the probability distributions corresponding to different modalities, evidence theory methods are used to fuse them to obtain fault probability distributions and confidence information associated with different modalities. This method of fault diagnosis using multimodal data solves the problems of low information utilization and one-sided results caused by using single data for fault prediction. On the other hand, by introducing uncertainty factors through probability distribution, it solves the problem of low reliability of fault prediction results caused by unreliable data in multimodal data fusion, ensuring the credibility and interpretability of fault prediction results and improving the accuracy and robustness of fault prediction results.
[0061] In some embodiments, before fusing the probability distributions of different modes, the features of the same mode under different noise conditions are forced to align, and the features of the same mode under different noise conditions are balanced and compressed. The forced alignment and balanced feature compression processes can enhance the stability of the same mode in noisy environments and prevent feature distortion.
[0062] It should be noted that the purpose of forced distribution alignment is to minimize the difference in feature distribution of the same mode under different perturbations, so that it can resist the influence of noise. In this way, even if different noises are introduced into the sensor data, the probability distribution obtained by the data feature transformation can still remain consistent.
[0063] It should be noted that the purpose of balanced feature compression is to compress redundant information and noise in the input data as much as possible while retaining task-related information (key features related to fault diagnosis).
[0064] In practice, forced distribution alignment can be achieved using algorithms such as KL divergence, JS divergence, maximum mean difference, and Wasserstein distance; balanced feature compression can be achieved using algorithms such as information bottleneck, variational autoencoder, principal component analysis, and sparse coding, which will not be listed here.
[0065] To prevent multimodal features from being distorted in noisy environments, this embodiment introduces a distribution consistency and information bottleneck mechanism. The specific loss function is as follows:
[0066] ;
[0067] ;
[0068] In the formula, This is the distributed consistency loss, which measures the consistency of the feature distribution of the same modality under different noise or data augmentation conditions. and These are samples of the same modality data that have undergone different data augmentations (or are subject to different noise interferences). for and Expectations It is the KL divergence, used to measure the difference between two distributions; For mode m; Indicates the given raw data hour, The conditional probability distribution; Indicates the given raw data hour, The conditional probability distribution;
[0069] Losses due to information bottlenecks; It is a prior distribution. It is a probability distribution With fault labels The mutual information between them is approximated in practice by the cross-entropy loss of the classification task; It is a hyperparameter that controls compressive strength; Used to make the characteristic distribution Approximately simple prior distribution (e.g., standard normal distribution).
[0070] In the above embodiment of balanced feature compression, the hyperparameters The value range is set to 0.005-0.05. This ensures sufficient fault information is retained while effectively compressing redundant information and noise, thereby improving the model's robustness in noisy environments. This is because if the hyperparameter is too small (less than 0.005), the compression term... The gradient signal is weak, and the feature distribution learned by the model is... It will be very complicated, with If the differences are significant, the learned features may contain too much redundant information and noise; if the hyperparameter is too large (greater than 0.05), the compression term will be insufficient. The gradient signal will be too strong, and the characteristic distribution will be too weak. All of them will be optimized to be very close to the prior distribution. At this point, excessive feature compression leads to the discarding of some information useful for fault diagnosis, reducing the distinguishability of feature distributions for different equipment states and different fault types, making them similar and unable to effectively distinguish between different states, thus reducing the accuracy of subsequent diagnosis.
[0071] In some embodiments, when the source data is a vibration signal, the corresponding data features are obtained by processing the vibration signal after short-time Fourier transform by a multi-scale convolutional neural network.
[0072] When the source data is a temperature or power signal, the corresponding data features are obtained based on a hybrid recurrent neural network;
[0073] When the source data is image data, the corresponding data features are obtained based on a convolutional neural network;
[0074] When the source data is text data, the corresponding data features are obtained based on the Transformer network.
[0075] It should be noted that since source data from different modalities are generally processed using different feature extractors, preprocessing is usually required for the data from different models in order to ensure that the data meets the input requirements of the respective feature extractors.
[0076] It should be noted that the source data for different modalities are time-aligned data. That is, after acquiring the data uploaded by different sensors, time alignment is performed first, and then feature extraction is performed using the feature extractor corresponding to the modal data.
[0077] It should be noted that, in order to ensure that the data meets the requirements of subsequent feature processing, the source data (which has been time-synchronized) needs to be preprocessed.
[0078] The specific steps are as follows:
[0079] For the original vibration signal that has been synchronized in time By applying the short-time Fourier transform, it is converted from a one-dimensional time series containing high-frequency impulse characteristics into a two-dimensional time-spectrum graph. To retain both time and frequency information; ;in, Represents the time spectrum over time t and frequency f. For window functions, The original vibration signal, The kernel of the Fourier transform. It is an imaginary number;
[0080] Low-frequency temperature and power signals are subjected to exponential smoothing or moving average filtering to eliminate noise and are then aligned with the time points of the vibration signal through interpolation; the exponential smoothing formula is: ;in, It is the smoothed estimate. This is a smoothed estimate from the previous time step. This is the original temperature measurement value. is the smoothing factor, a hyperparameter between 0 and 1;
[0081] Image and text data are classified according to their recording time. It is then indexed and matched with the synchronized timing signal.
[0082] This embodiment employs different feature extraction methods for source data of different modalities to ensure accurate representation of the features of each modality. Specifically:
[0083] Vibration signal feature extraction: Time-frequency spectrum processing using a multi-scale convolutional neural network (MSCNN) .
[0084] In the formula, Represents the vibration mode 1 Feature map of the layer; Indicates the convolution operation; Indicates input data, i.e., the spectrum. ; and It is the first Layer convolution kernels and biases, using convolution kernels of different scales to simultaneously capture long-term trends and short-term shock features.
[0085] Temperature and power signal feature extraction: Processing temperature signals using a hybrid recurrent neural network (LSTM and GRU). With power signal The core formula for the LSTM unit is as follows:
[0086] ;
[0087] In the formula, The input vector (temperature or power mode data) for the current time step t. , , These are the forget gate, input gate, and output gate of the LSTM unit, respectively. , , These represent the memory units at time steps. The cell state at the previous time step t-1, and the candidate cell state; , , , These represent the weight matrices corresponding to the forget gate, input gate, candidate state, and output gate, respectively. , , , These represent the bias vectors corresponding to the forget gate, input gate, candidate state, and output gate, respectively. For time step The hidden state (i.e., the output of the L unit). It is the sigmoid function. It is the Hadamard product.
[0088] Image Feature Extraction: Image Processing Using Convolutional Neural Networks (CNNs) and Spatial-Channel Attention Mechanisms Spatial attention is used to generate a weight map that highlights key regions in the image; channel attention is used to learn the importance weights of different feature channels.
[0089] Text Feature Extraction: Processing Text Data Using a Small Transformer Network Its core attention mechanism is as follows: ;in, These are query, key, and value matrices, respectively. It is the dimension of the key vector. For querying the dot product of the key.
[0090] In some embodiments, the data features corresponding to the vibration signal include a first periodic trend feature and a second periodic trend feature.
[0091] The sampling period corresponding to the first periodic trend is greater than the sampling period corresponding to the second periodic trend.
[0092] It should be noted that the multi-scale convolutional neural network in this embodiment includes branches or convolutional layers with convolutional kernels of different scales, the purpose of which is to enable it to extract features at different time scales from the input data simultaneously. Since convolutional kernels of different scales have different receptive fields and information capture capabilities, such as large-scale convolutional kernels having a wider receptive field than small-scale convolutional kernels and being able to capture information over a longer time span, the first periodic trend feature is the long-period feature extracted using the large-scale convolutional kernel, and the second periodic trend feature is the short-period feature extracted using the small-scale convolutional kernel.
[0093] The first periodic trend feature in this embodiment is used to characterize the overall operating status of the equipment, such as the overall vibration energy level change of the equipment and the slow trend of load-related changes.
[0094] The second periodic trend feature in this embodiment is used to characterize the instantaneous faults of the equipment, such as extracting the brief impact signal generated by the fault point with precise location.
[0095] In specific implementation, the branches or convolutional layers with convolutional kernels of different scales in the multi-scale convolutional neural network of this embodiment can be parallel or serial structures. For example, the multi-scale convolutional neural network sequentially includes an input layer, a convolutional structure, a feature fusion layer, a fully connected layer, and an output layer. The convolutional structure includes multiple parallel branches, each branch using a convolutional kernel of a different size (e.g., the first branch uses an 8*8 convolutional kernel, and the second branch uses a 64*64 convolutional kernel). As another example, the multi-scale convolutional neural network sequentially includes an input layer, a convolutional structure, a feature fusion layer, a fully connected layer, and an output layer. The convolutional structure includes multiple sequentially connected sub-convolutional structures, each sub-convolutional structure sequentially including a small-scale convolutional kernel (e.g., all 3*3) and a pooling layer.
[0096] In some embodiments, the probability distribution corresponding to different modes is a Gaussian distribution obtained based on the deterministic eigenvector, mean vector and covariance matrix of each mode.
[0097] In this embodiment, the deterministic feature vector refers to the feature representation of modal data, specifically a fixed-dimensional vector extracted by a feature extractor, namely a multi-scale convolutional neural network, a hybrid recurrent neural network, a convolutional neural network, and a Transformer network.
[0098] In this embodiment, the mean vector and covariance matrix are probabilistic parameters of the deterministic eigenvector, used to define a Gaussian distribution. The mean vector refers to the feature center after probabilistic transformation of the deterministic eigenvector, i.e., the central trend or the most likely value; the covariance matrix is used to quantify the possible fluctuation range and direction of the mean vector.
[0099] This embodiment will define the deterministic feature vector of each mode. Converting to a probability distribution form aims to reveal the uncertainty of the modeling features. Specifically, multimodal probability distribution modeling is performed according to the following formula:
[0100] (1) Use a feature extractor to extract the deterministic feature vector of each mode m. ;
[0101] (2) For each mode Deterministic feature vectors The multimodal probability distribution is modeled as a Gaussian distribution, and the process is as follows: ;in, For random variables The probability density function, Let be the probability density function of a multivariate Gaussian distribution, where Indicates a Gaussian distribution. This represents the mean vector of a Gaussian distribution. express The correlation between the dimensions and the variance of each dimension itself.
[0102] In this embodiment, the covariance matrix It is usually simplified to a diagonal matrix, that is, assuming that the feature dimensions are independent, its diagonal elements This represents the variance or uncertainty of each feature dimension. In this way, the output of each modality is no longer a fixed point, but a distribution, providing a measure of the reliability of the information for subsequent fusion.
[0103] In some embodiments, when performing cross-modal semantic alignment of probability distributions corresponding to different modes in the probability space, the selected temperature hyperparameter ranges from 0.05 to 0.3.
[0104] Because the original feature spaces of data from different modalities are different—for example, vibration signals are numerical sequences, images are pixel matrices, and text is character sequences—they lack direct comparability, and directly fusing them is meaningless for fault diagnosis. Therefore, this embodiment uses forced semantic alignment to convert data from different modalities into probability distributions, enabling similar representations of the same device state within the same probability space, while producing significantly different representations for different device states.
[0105] This embodiment uses cross-modal semantic alignment based on probabilistic contrastive learning. That is, by introducing a probabilistic contrastive learning mechanism, different modalities are compared and trained in the probability distribution space to ensure semantic alignment between modalities.
[0106] The probabilistic contrastive learning loss function is:
[0107] In the formula, For probabilistic contrastive learning loss, The number of positive sample pairs. Indicates a positive sample pair. and The probability distribution of two different modes describing the state of the same device at the same point in time. and They represent the first One and A Gaussian distribution with multiple modes It is a set of positive sample pairs (i.e., different modal distribution pairs from the same point in time that describe the state of the same device). yes and The similarity between two Gaussian distributions is usually expressed as the negative of the 2-Wasserstein distance.
[0108] Therefore, similarity can be defined as: ; It is a temperature hyperparameter used to adjust the sensitivity of contrast loss.
[0109] Temperature overparameters The smaller the value, the higher the similarity. The larger the value, the more obvious the differences between similarities; temperature hyperparameter The higher the value (greater than 0.3), the higher the similarity. The smaller the value, the less obvious the differences between similarities. In this embodiment, the temperature hyperparameter range of 0.05-0.3 can just right highlight the discriminative similarity differences in multimodal data, neither being overly sensitive to cause training oscillations nor too insensitive to cause coarse alignment.
[0110] when When smaller (e.g.) →0), the loss function is more sensitive to difficult samples, which amplifies the similarity difference between positive and negative samples; when When smaller (e.g.) If the similarity difference is greater than 1, the difference is reduced, the features learned by the model have weak discriminative power, and the model performance deteriorates. Therefore, this embodiment will... Setting the value to a range of 0.05-0.3 ensures that the model can learn consistent representations between modalities with appropriate granularity during cross-modal semantic alignment. This means that the distributions of different modalities under the same device state can be sufficiently close, while the distributions of different device states can be sufficiently far apart, thus ensuring stable training and strong generalization ability.
[0111] In some embodiments, the sampling frequency of the vibration signal data is 8-12kHz, and the sampling frequency of the temperature signal and power signal is 5-15Hz.
[0112] Typical mechanical failures in washing and screening vibrating screen equipment, such as bearing peeling, gear tooth breakage, and component loosening, generate brief, high-frequency impact vibrations. These impacts excite high-frequency resonances in the equipment structure, with frequency components typically reaching several kilohertz. An 8kHz sampling frequency can cover most of the impact responses and structural resonances caused by early failures, while a 12kHz sampling frequency provides a margin for capturing higher-frequency fault characteristics, while also avoiding the generation of unnecessary massive amounts of data due to excessively high sampling rates, thus reducing the burden on storage and computing.
[0113] Since temperature rise and power anomalies in equipment are typically slow-changing processes, the temperature and power signals of the washing and screening vibrating screen are low-frequency, slowly varying signals, with effective fault information concentrated in extremely low frequency bands. Therefore, if the sampling frequency of the temperature and power signals is too high (e.g., 100Hz), a large number of highly redundant data points will be generated. Because adjacent points have almost identical values, this increases storage and computational burden. Simultaneously, high-frequency sampling introduces more sensor noise, making the model prone to learning noise rather than the true trend, reducing robustness. Furthermore, high sampling frequencies place higher demands on hardware, increasing the cost of fault diagnosis for the washing and screening vibrating screen. If the sampling frequency of the temperature and power signals is too low (e.g., 1Hz), key transient features will be missed, and the limited number of data points will affect subsequent feature extraction, failing to achieve early warning and accurate diagnosis. This embodiment sets the sampling frequency of the temperature and power signals to 5-15Hz. This is sufficient to depict their changing trends and capture abnormal rise or fall slopes, while also facilitating implementation in industrial sensors and control systems, providing higher time resolution for observing dynamic processes.
[0114] The principles and working process of this application will be further explained below with reference to another embodiment.
[0115] like Figure 2 As shown, the method specifically includes steps S1-S8.
[0116] S1: Data acquisition and synchronization;
[0117] First, the following four types of modal data are simultaneously collected using multiple sensors installed on the washing and screening vibrating screen equipment:
[0118] Vibration signal A triaxial accelerometer can be used for data acquisition, with a sampling frequency of... To capture signals with high-frequency impact and structural fault characteristics , ,in, These are the signals representing the acceleration over time along the x, y, and z axes, respectively. Yes The column vector obtained by transposing;
[0119] Temperature signal With power signal Temperature sensors (such as PT100) and power transmitters can be used to collect data at a sampling frequency of [frequency missing]. It is used to monitor the thermal stability and load status of equipment;
[0120] Image data Explosion-proof industrial cameras can be used to acquire visual images at a fixed frame rate (e.g., 30fps) for detecting appearance anomalies (e.g., cracks, wear).
[0121] Text data It involves extracting time-related text information from inspection records, maintenance logs, and control system alarms;
[0122] Secondly, a unified high-precision timestamp is applied to the above four types of modal data. .
[0123] For signals with different sampling rates, dynamic time warping is used for time alignment to ensure that all modal data are strictly synchronized on the time axis t, forming a unified multimodal data sample as shown in Table 1. , .
[0124] Table 1 - Sample of some multimodal data
[0125]
[0126] S2: Data Preprocessing and Feature Standardization
[0127] The collected multimodal data is preprocessed to ensure it meets the requirements for subsequent processing. The specific steps are as follows:
[0128] The original vibration signals acquired synchronously Applying the short-time Fourier transform, it is converted from a one-dimensional time series into a two-dimensional time-spectrum graph. A portion of the two-dimensional time-spectrum graph is shown below. Figure 3 As shown, this is done to simultaneously retain time and frequency information:
[0129] ;in, Represents the time spectrum over time t and frequency f. For window functions, The original vibration signal, The kernel of the Fourier transform. It is an imaginary number;
[0130] Low-frequency temperature and power signals are subjected to exponential smoothing or moving average filtering to eliminate noise, and then interpolated to align them with the time points of the vibration signal. Exponential smoothing formula: ;in, It is the smoothed estimate. This is a smoothed estimate from the previous time step. This is the original temperature measurement value. is the smoothing factor, a hyperparameter between 0 and 1;
[0131] Image and text data are sorted according to their recording time. It is then indexed and matched with the synchronized timing signal.
[0132] Preprocessed data for each mode (in (representing modes) using a Gaussian distribution. This is used to characterize its inherent uncertainty and stability.
[0133] In the formula, the mean The average trend and variance of the data This represents the volatility or uncertainty of the data. The parameters of the Gaussian distribution are shown in Table 2.
[0134] Table 2 - Some Gaussian Distribution Parameters
[0135]
[0136] S3: Modal Feature Extraction and Modeling
[0137] Different feature extraction methods are used to ensure that the features of each modality are accurately represented:
[0138] Vibration signal: Time spectrum processed by multi-scale convolutional neural network (MSCNN) .
[0139] ;in, This represents the convolution operation. and It is the first Layer convolution kernels and biases, using convolution kernels of different scales to simultaneously capture long-term trends and short-term shock features, such as... Figure 4 As shown.
[0140] Temperature and power signal feature extraction: Processing temperature signals using a hybrid recurrent neural network (LSTM and GRU). With power signal The core formula for the LSTM unit is as follows:
[0141] ;
[0142] In the formula, For the current time step The input vector (temperature or power mode data); , , These are the forget gate, input gate, and output gate of the LSTM unit, respectively. , , These represent the memory units at time steps. The cell state at the previous time step Cell state, candidate cell state; , , , These represent the weight matrices corresponding to the forget gate, input gate, candidate state, and output gate, respectively. , , , These represent the bias vectors corresponding to the forget gate, input gate, candidate state, and output gate, respectively. This represents the hidden state at time step t (i.e., the output of the LSTM unit). It is the sigmoid function. It is the Hadamard product.
[0143] Image Feature Extraction: Image Processing Using Convolutional Neural Networks (CNNs) and Spatial-Channel Attention Mechanisms .
[0144] Text Feature Extraction: Processing Text Data Using a Small Transformer Network .
[0145] S4: Multimodal probability distribution modeling:
[0146] The deterministic feature vector of each mode extracted from S3 This is converted to a probability distribution form to show the uncertainty of the modeled features. For each mode m, its features are modeled as a Gaussian distribution:
[0147] ;in, It is the mean vector of the feature distribution, representing the central tendency of the feature; It is the covariance matrix (usually simplified to a diagonal matrix, i.e., assuming independence between feature dimensions), whose diagonal elements This represents the variance or uncertainty of each feature dimension. In this way, the output of each modality is no longer a fixed point, but a distribution, providing a measure of the reliability of the information for subsequent fusion.
[0148] S5: Cross-modal semantic alignment based on probabilistic contrastive learning:
[0149] A probability-based contrastive learning mechanism is introduced to perform comparative training on different modalities in a probability distribution space, ensuring semantic alignment between modalities. The probabilistic contrastive learning loss function is:
[0150] ;in, It is a set of positive sample pairs (i.e., different modal distribution pairs from the same point in time that describe the state of the same device).
[0151] It is a similarity measure between two Gaussian distributions, usually using the negative value of the 2-Wasserstein distance;
[0152] ;
[0153] Similarity can be defined as: ;in, It is a temperature hyperparameter used to adjust the sensitivity of contrast loss.
[0154] S6: Robust learning mechanism:
[0155] To prevent multimodal features from being distorted in noisy environments, a distribution consistency and information bottleneck mechanism is introduced. The specific loss function is:
[0156] ;in, and These are samples of the same modality data that have undergone different data augmentations (or are subject to different noise interferences). It is the KL divergence, used to measure the difference between two distributions;
[0157] ;in, It is a prior distribution. It is a feature With fault labels The mutual information between them is approximated in practice by the cross-entropy loss of the classification task; It is a hyperparameter that controls compressive strength.
[0158] S7: Multimodal Fusion and Trustworthy Decision Making
[0159] This application employs Dempster–Shafer evidence theory to fuse the probability distribution outputs of each modality into a comprehensive decision with confidence assessment.
[0160] Basic Probability Assignment (BPA): Assigning the probability distribution of each mode. Its uncertainty is transformed into a set of fault assumptions. BPA function .
[0161] DS Combination Rule: BPA functions of all modes To integrate, among which, The BPA function represents the k-th mode. .
[0162] ;in, It is a set of fault assumptions any subset of; The basic probability assignment function for the k-th mode. The elements of the domain, i.e., the set of fault assumptions. any subset; It is a normalization constant, called the conflict factor, and is calculated using the following formula: ;in, The magnitude of the value reflects the degree of conflict between intermodal evidence.
[0163] This embodiment has 5 modes (vibration signal, temperature signal, power signal, image data, and text data), namely .
[0164] Decision output: Based on the fused BPA function Making a decision typically involves calculating the elements of each set of failure hypotheses. Reliability and similarity. Among them, reliability... It is The minimum level of trust required for authenticity; similarity Yes This represents the highest level of confidence (not false). The final fault diagnosis result is the hypothesis with the highest confidence or Pignistic probability. .
[0165] The results of the DS evidence theory fusion include:
[0166] Diagnostic decision: A fault exists;
[0167] Confidence level: 0.714;
[0168] Conflict score: 0.997;
[0169] Modal weights: vibration 0.639, temperature 0.228, power 0.134.
[0170] S8: Output and Feedback. This embodiment provides a structured diagnostic report, specifically including:
[0171] Fault type ;
[0172] Confidence ;
[0173] Confidence interval ;
[0174] Contribution assessment of each modality: This is calculated by analyzing the weight of the BPA of each modality in the fusion process or its positive impact on the final reliability.
[0175] Intermodal conflict information: i.e., the value of the conflict factor. , The larger the value, the greater the contradiction in the evidence provided by different sensors, and the reliability of the results should be treated with caution. The output is as follows: Figure 5 The fault diagnosis results are shown.
[0176] One embodiment of this application provides a fault diagnosis device for a washing and screening vibrating screen based on multimodal fusion, such as... Figure 6 As shown, the device 60 includes: a data feature acquisition module 601, a probability distribution determination module 602, a modal probability association module 603, and a modal probability fusion module 604.
[0177] The data feature acquisition module 601 is used to extract features from source data of different modalities to obtain data features corresponding to different modalities.
[0178] The probability distribution determination module 602 is used to transform the data features corresponding to different modes to obtain the probability distributions corresponding to different modes;
[0179] The modal probability association module 603 is used to perform cross-modal semantic alignment of the probability distributions corresponding to different modalities in the probability space to obtain the association between the probability distributions corresponding to different modalities.
[0180] The modal probability fusion module 604 is used to fuse the probability distributions of different modes based on evidence theory methods to obtain the fault probability distribution and confidence information associated with different modes.
[0181] The data of different modes include vibration signal data, temperature signal, power signal, image data, and text data associated with the washing and screening vibrating screen equipment.
[0182] This application embodiment extracts features from source data of different modalities, obtains data features corresponding to different modalities, and converts them into probability distributions. The probability distributions corresponding to different modalities are then semantically aligned across modalities in the probability space. After obtaining the correlation between the probability distributions corresponding to different modalities, evidence theory methods are used to fuse them to obtain fault probability distributions and confidence information associated with different modalities. This method of fault diagnosis using multimodal data solves the problems of low information utilization and one-sided results caused by using single data for fault prediction. On the other hand, by introducing uncertainty factors through probability distribution, it solves the problem of low reliability of fault prediction results caused by unreliable data in multimodal data fusion, ensuring the credibility and interpretability of fault prediction results and improving the accuracy and robustness of fault prediction results.
[0183] Furthermore, before fusing the probability distributions of different modalities, the features of the same modality under different noise levels are forced to align their distributions, and the features of the same modality under different noise levels are subjected to balanced feature compression.
[0184] Furthermore, when performing balanced feature compression, the hyperparameter values range from 0.005 to 0.05.
[0185] Furthermore, when the source data is a vibration signal, the corresponding data features are obtained by processing the vibration signal after short-time Fourier transform using a multi-scale convolutional neural network.
[0186] When the source data is a temperature or power signal, the corresponding data features are obtained based on a hybrid recurrent neural network;
[0187] When the source data is image data, the corresponding data features are obtained based on a convolutional neural network;
[0188] When the source data is text data, the corresponding data features are obtained based on the Transformer network.
[0189] Furthermore, the data characteristics corresponding to the vibration signal include first-cycle trend characteristics and second-cycle trend characteristics;
[0190] The sampling period corresponding to the first periodic trend is greater than the sampling period corresponding to the second periodic trend.
[0191] Furthermore, the probability distributions corresponding to different modes are Gaussian distributions obtained based on the deterministic eigenvectors, mean vectors, and covariance matrices of each mode.
[0192] Furthermore, when performing cross-modal semantic alignment of the probability distributions corresponding to different modes in the probability space, the selected temperature hyperparameter ranges from 0.05 to 0.3.
[0193] Furthermore, the sampling frequency of the vibration signal data is 8-12kHz, and the sampling frequency of the temperature signal and power signal is 5-15Hz.
[0194] The fault diagnosis device for washing and screening vibrating screen equipment based on multimodal fusion in this embodiment can execute the fault diagnosis method for washing and screening vibrating screen equipment based on multimodal fusion shown in the embodiment of this application. The implementation principle is similar, and will not be described again here.
[0195] Another embodiment of this application provides an electronic device, including: a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above method.
[0196] Specifically, the processor can be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0197] Specifically, the processor connects to the memory via a bus, which may include a path for transmitting information. The bus can be a PCI bus or an EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc.
[0198] The memory may be ROM or other types of static storage devices that can store static information and instructions, RAM or other types of dynamic storage devices that can store information and instructions, or EEPROM, CD-ROM or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto.
[0199] Optionally, the memory stores the code of a computer program that executes the scheme of this application, and the execution is controlled by a processor. The processor executes the application code stored in the memory to implement the operation of the apparatus provided in the above embodiments.
[0200] Another embodiment of this application provides a computer-readable storage medium storing computer-executable instructions for performing the methods provided in the above embodiments.
[0201] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0202] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0203] The above is a detailed description of the preferred embodiments of this application. However, this application is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
[0204] The above is a detailed description of the preferred embodiments of this application. However, this application is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A multi-modal fusion-based fault diagnosis method for a washing and screening vibrating screen device, characterized by, include: Feature extraction is performed on source data of different modalities to obtain data features corresponding to different modalities; The data features corresponding to different modalities are transformed to obtain the probability distributions corresponding to different modalities; Semantic alignment of probability distributions corresponding to different modalities is performed across the probability space to obtain the correlation between probability distributions corresponding to different modalities. The probability distributions of different modalities are used as input, and the probability space is learned during the model training process through the introduced probability contrastive learning mechanism. In this probability space, positive samples are close to each other and negative samples are far apart. Positive samples describe different modal distributions of the same device state, and negative samples describe different modal distributions of different device states. The probability distributions of different modes are fused based on evidence theory methods to obtain the fault probability distributions and confidence information associated with different modes; The data of different modes include vibration signal data, temperature signal, power signal, image data and text data associated with the washing and screening vibrating screen equipment; Before fusing the probability distributions of different modalities, the features of the same modality under different noise levels are forced to align their distributions, and the features of the same modality under different noise levels are subjected to balanced feature compression. raw vibration signal for which the time synchronization has been completed Applying a short-time Fourier transform converts it from a one-dimensional time series containing high-frequency impact features into a two-dimensional time-frequency spectrogram while preserving both time and frequency information, denotes the time spectrum in time t and frequency f; Vibration signal feature extraction: through a multi-scale convolutional neural network to process two-dimensional time-frequency spectrograms ; wherein, denotes the vibration mode of the layer feature map; denotes the convolution operation; denotes the input data, i.e. the two-dimensional time-frequency spectrum ; and are the convolution kernel and bias of the layer, using different scales of the convolution kernel to capture both long-period trends and short-time impact features; The data characteristics corresponding to the vibration signal include the first period trend characteristics and the second period trend characteristics; The sampling period corresponding to the first periodic trend is greater than the sampling period corresponding to the second periodic trend. The first-cycle trend feature is used to characterize the overall operating status of the equipment, and the second-cycle trend feature is used to characterize the instantaneous faults of the equipment. Low-frequency temperature and power signals are subjected to exponential smoothing or moving average filtering to eliminate noise and are aligned with the time points of the vibration signal by interpolation. Temperature and power signal feature extraction: processing temperature signals using a hybrid recurrent neural network. With power signal ; Image and text data are classified according to their recording time. It is indexed and matched with the synchronized timing signal; Image Feature Extraction: Image Processing Using Convolutional Neural Networks and Spatial-Channel Attention Mechanisms Spatial attention is used to generate a weight map that highlights key regions in the image; channel attention is used to learn the importance weights of different feature channels. Text Feature Extraction: Processing Text Data Using a Small Transformer Network ; The probability distributions corresponding to different modes are Gaussian distributions obtained based on the deterministic eigenvectors, mean vectors, and covariance matrices of each mode; Specifically, the multimodal probability distribution modeling should be completed according to the following formula: (1) Use a feature extractor to extract the deterministic feature vector of each mode m. ; (2) For each mode Deterministic feature vectors The multimodal probability distribution is modeled as a Gaussian distribution, and the process is as follows: ;in, For random variables The probability density function, Let be the probability density function of a multivariate Gaussian distribution, where Indicates a Gaussian distribution. This represents the mean vector of a Gaussian distribution. express The correlation between the dimensions and the variance of each dimension itself.
2. The method as described in claim 1, characterized in that, When performing balanced feature compression, the hyperparameter values range from 0.005 to 0.
05.
3. The method as described in claim 1, characterized in that, When the source data is a vibration signal, the corresponding data features are obtained by processing the vibration signal after short-time Fourier transform using a multi-scale convolutional neural network. When the source data is a temperature or power signal, the corresponding data features are obtained based on a hybrid recurrent neural network; When the source data is image data, the corresponding data features are obtained based on a convolutional neural network; When the source data is text data, the corresponding data features are obtained based on the Transformer network.
4. The method as described in claim 1, characterized in that, When performing cross-modal semantic alignment of probability distributions corresponding to different modes in the probability space, the selected temperature hyperparameter ranges from 0.05 to 0.
3.
5. The method as described in claim 1, characterized in that, The sampling frequency for vibration signal data is 8-12kHz, and the sampling frequency for temperature and power signals is 5-15Hz.
6. An electronic device, characterized in that, The method includes a memory and a processor, the memory being coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the method according to any one of claims 1-5.
7. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Multi-source data fusion fault diagnosis method and system based on DS evidence theory
CN119557766A
Brain-like multi-modal semantic probabilistic alignment and integration measurement method and system
CN121093962A