Industrial equipment state prediction method based on multi-modal sensing fusion
By timestamp alignment and preprocessing the multimodal sensor data of the rotary tiller, building a Transformer fusion network, and generating the device state vector, the problem of insufficient modal information fusion in the existing technology is solved, and high-precision state prediction and reliable early warning of the rotary tiller are achieved, reducing maintenance costs.
Patent Information
- Application Number
- CN202510789274.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing industrial equipment status prediction technologies have problems such as low modal information fusion, single prediction features, and insufficient fault warning accuracy. It is difficult to accurately, stably and comprehensively predict and warn the dynamic working conditions of complex equipment. Especially in agricultural equipment such as rotary tillers, single modality or simple fusion technology cannot identify early fault hazards, resulting in insufficient warning timeliness and increased maintenance costs.
By collecting multimodal sensor data of a rotary tiller under a unified time base, performing timestamp alignment and preprocessing, and using feature extraction methods to generate fused feature tensors, a dual-path Transformer fusion network was constructed. A gated recurrent unit timing prediction model with an attention mechanism was combined to output the device state vector, and maintenance warning information was generated based on the threshold.
It achieves high-precision prediction and reliable early warning of the rotary tiller equipment status, improves prediction accuracy and early warning reliability, reduces maintenance costs, and ensures the safety and reliability of equipment operation.
Smart Images

Figure CN120744347A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent monitoring and state prediction of industrial equipment, and in particular to an industrial equipment state prediction method based on multimodal sensor fusion. Background Art
[0002] Industrial equipment status prediction technology is an important support for realizing refined management of modern industry and intelligent maintenance of equipment. At present, industrial equipment status prediction is mainly achieved through physical models, data-driven methods and their hybrid strategies. Among them, data-driven methods are widely used because they do not require in-depth mechanism modeling and have strong flexibility and adaptability. However, the prediction method of single modal data has the problems of limited data dimension and insufficient feature expression. Therefore, multimodal fusion technology has gradually become an emerging research direction to improve the accuracy of equipment status prediction.
[0003] CN118070491A discloses a method for predicting the health status of industrial equipment based on a hybrid approach, which predicts the power degradation trend by fitting an empirical function. However, its single mode (output power) data cannot fully characterize the overall operating status of the equipment under complex working conditions.
[0004] CN116599034A discloses a method for load prediction and control optimization of agricultural equipment. Although it takes load data and meteorological factors into account, it lacks multimodal fusion analysis of mechanical and physical signals such as vibration and temperature during the equipment's own operation, resulting in the prediction results being unable to accurately reflect the equipment's actual operating status and potential failure trends.
[0005] In summary, the existing industrial equipment status prediction technology has the problems of low modal information fusion, single prediction features, and insufficient fault warning accuracy. It is difficult to accurately, stably and comprehensively predict and warn the dynamic working conditions of complex equipment; especially in the field of agricultural equipment, such as rotary tillers and other mechanical equipment, there are multiple modes of coupling such as vibration, temperature rise, and electrical fluctuations in actual operation. Single modality or simple fusion technology is difficult to accurately identify early fault hazards, resulting in insufficient warning timeliness, increased maintenance costs, and damaged equipment reliability and life.
[0006] Therefore, how to effectively integrate multimodal data of equipment and use advanced machine learning methods to fully extract and integrate the characteristic information of each modality to improve the accuracy of state prediction and the reliability of early warning has become a technical problem that needs to be solved urgently in this field. Summary of the Invention
[0007] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract of the specification and the title of the invention of this application to avoid blurring the purpose of this section, the abstract of the specification and the title of the invention, and such simplifications or omissions cannot be used to limit the scope of the invention.
[0008] In view of the above existing problems, the present invention is proposed.
[0009] To solve the above technical problems, the present invention provides the following technical solutions: collecting multimodal sensor data of a rotary tiller within a fixed time window, and performing time stamp alignment on the multimodal sensor data based on a unified time reference;
[0010] performing preprocessing on the aligned multimodal sensing data to obtain a standardized data matrix;
[0011] Using a feature extraction method, the standardized data matrix is subjected to time-frequency wavelet packet decomposition, Mel spectrum extraction, and exponential sliding average calculation, and each modal feature vector is output;
[0012] Construct a dual-path Transformer fusion network consisting of a modality self-attention branch and a graph attention branch based on sensor space topology, fuse the feature vectors of each modality, and generate a fused feature tensor;
[0013] Inputting the fused feature tensor into a gated recurrent unit timing prediction model based on an attention mechanism, outputting a device state vector with a predicted time span, wherein the device state vector represents at least a health index and a failure probability;
[0014] The device state vector is compared with the health threshold and the fault threshold respectively. When any component exceeds the corresponding threshold, maintenance warning information is generated and reported.
[0015] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, the multimodal sensor data of the rotary tiller is collected, including: a three-axis vibration acceleration sensor installed on the gearbox housing and the frame, an infrared temperature and K-type thermocouple combination sensor arranged near the main shaft bearing, a MEMS microphone array located on the side panel of the machine body, a Hall speed encoder embedded in the power output shaft, a three-phase current-voltage integrated sensor module placed at the input end of the electronically controlled hydraulic pump, and a high-precision air pressure-humidity combination sensor installed in the machine cover.
[0016] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, the multimodal sensor data includes mechanical vibration acceleration data, infrared temperature measurement values and thermocouple temperature measurement values, sound waveform sequence, speed pulse count and instantaneous speed, phase current and phase voltage and power factor, ambient temperature, humidity and pressure.
[0017] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, timestamp alignment of the multimodal sensor data based on a unified time reference includes:
[0018] Using the GPS-PPS pulse signal as the external master clock, the local clock of each acquisition node is calibrated once a day with a maximum of 0.1ms in frequency and phase.
[0019] The data streams with different sampling frequencies are resampled to the reference frequency of 1kHz using linear interpolation;
[0020] The missing time segments are marked as missing and inpainted using local spline interpolation with a sliding window width of 5s.
[0021] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, preprocessing is performed on the aligned multimodal sensor data to obtain a standardized data matrix, including:
[0022] Apply fourth-order Butterworth bandpass filtering to the mechanical vibration acceleration data, speed pulse counts, and instantaneous speed and sound waveform sequences;
[0023] Outliers were removed by using the isolation forest algorithm, and the removal threshold was 0.6 times the theoretical upper limit of the average path length;
[0024] Apply the 3σ principle to the phase current signal, phase voltage signal and power to eliminate pulse spikes;
[0025] Bidirectional exponential sliding smoothing is used for infrared temperature measurement values, thermocouple temperature measurement values, ambient temperature, humidity and pressure, and the smoothing factor is defined as 0.3;
[0026] All channels are Z-score normalized, and a two-dimensional standardized data matrix is constructed according to the sensor number-time step method.
[0027] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, outputting each modal feature vector includes:
[0028] The vibration signal is decomposed into five layers of Daubechies-4 wavelet packets to extract 63-dimensional features including energy entropy, kurtosis and main frequency of each sub-band.
[0029] A Mel spectrum with a frame length of 25ms and a frame shift of 10ms is generated for the acoustic emission signal. The number of filter banks is 64. The logarithmic amplitude is taken and global average pooling is performed to obtain a 64-dimensional acoustic feature.
[0030] An exponential sliding average with a window length of 100ms is used for the three-phase current and voltage to output real-time RMS, peak-to-peak value, and phase angle 7-dimensional features.
[0031] The residuals are retained after the second-order polynomial trend decomposition of the environmental parameters as the three-dimensional interference compensation features.
[0032] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, generating the fusion feature tensor includes:
[0033] The modal self-attention branch uses 8-head attention, defines the key value dimension as 64, and the output dimension as 256;
[0034] The sensor space topology is constructed as an undirected weighted graph based on the physical coordinates and functional correlation of the sensors. The edge weight is the weighted sum of the inverse of the normalized Euclidean distance and the mutual information score.
[0035] The graph attention branch uses a two-layer graph attention network with 4 attention heads in each layer and a node embedding dimension of 128;
[0036] The outputs of the modal self-attention branch and the graph attention branch are cascaded according to the channel dimension and then linearly mapped and layer normalized to obtain a 384-dimensional fused feature tensor.
[0037] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, outputting the equipment state vector with the prediction time span includes:
[0038] Input the fused feature tensor into a two-layer gated recurrent unit with 256 hidden units in each layer, and introduce an adaptive attention gate between the two layers;
[0039] The prediction time span is a configurable discrete value within 5 to 30 minutes, and the optimal prediction time span is selected based on the least squares fitting residual;
[0040] The normalized health index HI and failure probability FP are obtained using the Sigmoid activation function.
[0041] As a preferred solution of the industrial equipment state prediction method based on multimodal sensor fusion described in the present invention, the equipment state vector is compared with the health threshold and the fault threshold respectively. When any component exceeds the corresponding threshold, maintenance warning information is generated and reported, including:
[0042] Health threshold HI thAdaptive update based on 15% of historical 30-day normal operating data;
[0043] Fault threshold FP th Defined as 0.6, when the failure probability FP exceeds the failure threshold FP for three consecutive prediction cycles th It is judged as a fault when
[0044] When HI <HI th or FP>FP th When an abnormality occurs, a JSON structured warning message containing the device number, abnormal modal weight ranking and recommended downtime window is generated;
[0045] The early warning information is pushed to the maintenance management platform via the MQTT protocol, and alerts are issued through multiple channels such as SMS and email.
[0046] The beneficial effects of the present invention are as follows: the present invention ensures the consistency and coordination of data by collecting and aligning the multimodal sensor data of the rotary tiller under a unified time reference; significantly enhances the accuracy and stability of feature expression through effective preprocessing and multidimensional feature extraction; fully captures the correlation information within and between modalities by constructing a dual-path Transformer fusion network of self-attention and graph attention, and improves the equipment status feature fusion capability; further adopts a gated recurrent unit based on the attention mechanism to achieve accurate state prediction and reliably evaluate the equipment health status and fault trend; finally, through threshold judgment and early warning mechanism, the prediction results are timely converted into effective maintenance decisions, thereby improving the safety and reliability of equipment operation and reducing maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:
[0048] Figure 1 Schematic diagram of the process of industrial equipment state prediction method based on multimodal sensor fusion shown in the present invention. DETAILED DESCRIPTION
[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0050] Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without making any creative work should fall within the scope of protection of the present invention.
[0051] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0052] According to an embodiment of the present invention, Figure 1 The flowchart shown is a method for predicting the state of industrial equipment based on multimodal sensor fusion, which specifically includes the following steps:
[0053] S1. Collect multimodal sensor data of the rotary tiller within a fixed time window and align the timestamps of the multimodal sensor data based on a unified time reference. The following should be noted in this step:
[0054] The triaxial vibration acceleration sensor installed on the gearbox housing and the frame is defined to have a sampling frequency of ≥6kHz;
[0055] The infrared temperature and K-type thermocouple combination sensor is placed near the spindle bearing, and the sampling frequency is defined as 2Hz;
[0056] The MEMS microphone array located on the side panel of the fuselage is used to collect acoustic emission signals in the range of 20Hz-20kHz;
[0057] The Hall effect speed encoder embedded in the power output shaft has a defined resolution of 2048P / R;
[0058] The three-phase current-voltage integrated sensor module placed at the input of the electronically controlled hydraulic pump has a defined sampling frequency of 10kHz.
[0059] A high-precision pressure-humidity combination sensor installed in the hood is used to record environmental operating parameters.
[0060] Exemplarily, the multimodal sensing data includes mechanical vibration acceleration data, infrared temperature measurement values and thermocouple temperature measurement values, sound waveform sequence, speed pulse count and instantaneous speed, phase current and phase voltage and power factor, ambient temperature, humidity and pressure.
[0061] Furthermore, timestamp alignment of multimodal sensor data is performed based on a unified time base, including:
[0062] Using the GPS-PPS pulse signal as the external master clock, the local clock of each acquisition node is calibrated once a day with a maximum of 0.1ms in frequency and phase.
[0063] The data streams with different sampling frequencies are resampled to the reference frequency of 1kHz using linear interpolation;
[0064] The missing time segments are marked as missing and inpainted using local spline interpolation with a sliding window width of 5s.
[0065] It should be noted that by collecting the multimodal sensor data of the rotary tiller within a fixed time window and aligning the timestamps of the multimodal sensor data based on a unified time reference, a high degree of synchronization of different types of sensor data in the time dimension is ensured, efficient fusion and collaborative analysis of cross-modal information is achieved, and the problems of information inconsistency or error accumulation caused by time misalignment between different modal data are eliminated, thereby effectively improving the accuracy and reliability of subsequent state analysis.
[0066] S2. Preprocess the aligned multimodal sensor data to obtain a standardized data matrix. Note that:
[0067] Apply a fourth-order Butterworth bandpass filter to the mechanical vibration acceleration data, speed pulse count, instantaneous speed, and sound waveform sequence, for example, with a passband of 20 Hz to 6 kHz, and use a Butterworth low-pass filter of the same order (cutoff 500 Hz) to suppress high-frequency electromagnetic noise;
[0068] The isolation forest algorithm is called on all filtered channels, with 256 subtrees and 256 subsamples. A threshold of 0.6 times the theoretical upper limit of the average path length is used to identify abnormal samples and perform linear interpolation replacement.
[0069] The three-phase current, voltage and power factor sequences are identified by the 3σ principle, and the peak segments are locally reconstructed based on the mean values on both sides of the adjacent segments.
[0070] Bidirectional exponential sliding smoothing is used for infrared temperature measurement values, thermocouple temperature measurement values, ambient temperature, humidity and pressure, and the smoothing factor is defined as 0.3;
[0071] Calculate the mean and standard deviation of all samples for each sensor channel i, and use the Z-score formula to normalize all time step data;
[0072] According to the order of sensor number and time step, the matrix row and column indexes are written in sequence to form a two-dimensional standardized data matrix D with the number of rows C (number of sensor channels) and the number of columns T (number of time steps).
[0073] As an example, the mathematical expression of the normalized data matrix D is as follows:
[0074]
[0075]
[0076] Among them, x i,t is the original value of the i-th sensor at the t-th time step, μ i is the mean value of all samples of the i-th sensor, σ i is the standard deviation of the entire sample of the i-th sensor, d i,t is the dimensionless value written into the matrix after normalization, C is the total number of multimodal sensor channels, and T is the total number of discrete time steps in a fixed time window.
[0077] It should be noted that by performing preprocessing on the aligned multimodal sensor data, including filtering, outlier removal, data smoothing and normalization operations, a standardized data matrix is obtained, which ensures the quality and consistency of the input data, effectively removes the influence of environmental noise, equipment interference and sensor errors, and provides a high-quality data foundation for the subsequent accurate extraction of equipment status features, significantly improving the stability and effectiveness of feature extraction.
[0078] S3. Using the feature extraction method, perform time-frequency wavelet packet decomposition, Mel spectrum extraction, and exponential sliding average calculation on the standardized data matrix, and output each modal feature vector. Among them, it should be noted that:
[0079] A five-layer Daubechies-4 wavelet packet decomposition is performed on the standardized vibration signal. The energy entropy, kurtosis and main frequency are calculated in each subband, and a total of 63-dimensional vibration features are obtained.
[0080] The acoustic emission signal is divided into frames with a frame length of 25ms and a frame shift of 10ms, and a Hamming window is added. A spectrogram is generated by passing it through 64 Mel filters. The logarithmic amplitude is taken and global average pooling is performed to output 64-dimensional acoustic features.
[0081] The three-phase current and voltage sequences are subjected to exponential sliding averaging using a 100ms sliding window. The real-time RMS value, peak-to-peak value, and phase angle are calculated within each window, and 7-dimensional electrical parameter features are output.
[0082] Perform second-order polynomial trend decomposition on the temperature, humidity, and pressure series, and retain only the residual part as the 3D environmental compensation feature;
[0083] Write the acquired features in sequence according to the modal number to form the vibration feature vector v (1) , acoustic eigenvector v (2) , electrical parameter eigenvector v (3) and the environmental feature vector v (4) ;
[0084] Furthermore, the mathematical expressions of the modal eigenvectors are obtained as follows:
[0085]
[0086] in, is the kth vibration subband feature (such as energy entropy / kurtosis / dominant frequency), is the logarithmic amplitude pooling coefficient of the kth Mel spectrogram, is the kth electrical parameter statistic (such as effective value, peak-to-peak value / phase angle), is the kth component in the environmental modal residual vector, v (j) is the eigenvector of the jth mode, is a real vector space of dimension n, and [·]T is the vector transpose symbol.
[0087] It should be noted that this embodiment uses a feature extraction method to perform time-frequency wavelet packet decomposition, Mel spectrum extraction, and exponential sliding average calculation on the standardized data matrix, and outputs each modal feature vector, thereby achieving efficient expression of the feature space of different modal data such as vibration signals, acoustic signals, and electrical signals. It can fully reveal the potential information of the equipment's operating status, effectively solve the problem of insufficient expression ability of a single modality or simple feature characterization method, and improve the sensitivity and richness of early fault feature capture.
[0088] S4. Construct a dual-path Transformer fusion network consisting of a modal self-attention branch and a graph attention branch based on sensor space topology, fuse the feature vectors of each modality, and generate a fused feature tensor.
[0089] Among them, the following points need to be explained in this step:
[0090] The four types of feature vectors obtained in step S3, namely vibration, acoustics, electrical parameters and environment, are stacked in sequence as input tensors Where M = 4 is the number of modes, and d is the length of each mode vector;
[0091] Add 8 multi-head self-attention to X, with the key / value dimension of each head set to 64, and generate the modal self-attention output after residual and layer normalization
[0092] Based on the three-dimensional physical coordinates of the sensor and the functional correlation of the measured variables, an undirected weighted graph G = (V, ε) is established, where the node V corresponds to a single sensor and the edge weight
[0093] Among them, p i is the coordinate vector, MI ij is the mutual information score, λ is the trade-off coefficient;
[0094] Set the initial feature of the node to the channel vector of the corresponding sensor in X;
[0095] Through a two-layer 4-head graph attention network, the node embedding dimension is 128, and the graph attention output is obtained
[0096] Concatenate Z and G in the channel dimension:
[0097] Linearly mapped matrix After ReLU activation, add layer normalization to get the fused feature tensor
[0098] For example, the mathematical expression of the fused feature tensor is as follows:
[0099]
[0100] Among them, Z t is the modal self-attention output at the t-th time step, G t is the graph attention output at the t-th time step, ⊕ is the channel dimension cascade operator, W is the 384×384 fully connected weight matrix, LN(·) is the layer normalization operator, F t is the 384-dimensional fused feature vector obtained at time step t.
[0101] It should be noted that by constructing a dual-path Transformer (a sequence model based on the attention mechanism) fusion network including a modal self-attention branch and a graph attention branch based on the sensor spatial topology, the feature vectors of each modality are deeply fused to generate a fused feature tensor, thereby realizing the collaborative learning of intra-modal and inter-modal features. Among them, the self-attention mechanism highlights the dynamic importance of each modal feature, and the graph attention mechanism explicitly utilizes the spatial information of the sensor topology structure, thereby enhancing the overall perception ability of the complex operating status of the equipment, avoiding the problem that simple linear fusion methods easily ignore the correlation between key features, and significantly improving the feature expression ability and prediction accuracy after fusion.
[0102] S5. Input the fused feature tensor into the attention-based gated recurrent unit time series prediction model, and output the device state vector with the predicted time span. The device state vector at least represents the health index and the failure probability.
[0103] Send F to the first layer of GRU in sequence, with 256 hidden units, to get the intermediate state
[0104] by As input, after adaptive attention gate Selective reinforcement of critical moment characteristics;
[0105] Will Input the second layer GRU, the number of hidden units is 256, and a high-order temporal representation is obtained.
[0106] Define discrete candidate sets For each Δ i Calculate the least squares residual
[0107] Pick As the optimal prediction span for this round;
[0108] right exist Perform exponential weighted integration on historical segments to form a comprehensive memory u t ;
[0109] After the fully connected layer W o Then use Sigmoid activation to obtain the normalized device state vector
[0110] As an example, the mathematical expression formula for outputting the device state vector with the predicted time span based on the time series prediction model in this embodiment is as follows:
[0111]
[0112] in, is the optimal prediction span at time t The obtained two-dimensional normalized state vector, the first dimension Represents health index, the second dimension represents the failure probability, σ(·) is the component-wise Sigmoid function to ensure that the output value range is (0,1), K is the number of historical segments into which the prediction window is divided, and e -αk is the exponential memory weight, erf(·) is the Gaussian error function, F t is the 384-dimensional fusion feature vector at time t, ||F τ ||2 is its binary norm, W g With W o are the weight matrices of the adaptive gate and output layer, is the hidden state of the second layer GRU at time tk, b is the bias scalar, ⊙ is the Hadamard product, and π is the error function normalization factor;
[0113] Furthermore, when HI>0.8, the equipment is in excellent condition; when HI<0.4, the equipment performance is degraded; and when FP>0.6, a fault warning is triggered.
[0114] It should be noted that by inputting the fused feature tensor into the gated recurrent unit timing prediction model based on the attention mechanism and outputting the device state vector with the predicted time span, it is possible to reliably predict the state of the device in the future. The gated recurrent unit effectively extracts timing information by capturing the long-term and short-term dependency features of historical data. The attention mechanism further enhances the focus on key state changes, thereby improving the sensitivity and accuracy of the prediction model to the evolution trend of device state.
[0115] S6, compare the device state vector with the health threshold and the fault threshold respectively, and generate maintenance warning information and report it when any component exceeds the corresponding threshold. Follow the steps below to complete threshold determination and warning reporting:
[0116] In the rolling window W 30d The health index sequence {HI τ};
[0117] Calculate the 15th percentile value of the sequence and set it as the current health threshold θ HI , and refresh every 24 hours;
[0118] Define the failure probability threshold θ FP =0.6;
[0119] Create a loop counter C with a length of 3 FP ,like The count is incremented by one, otherwise it is cleared to zero. FP =3, it is judged to enter the fault state;
[0120] like or C FP = If one of the three conditions is met, a maintenance warning will be triggered;
[0121] Call the attention weight of the fusion network Arrange in descending order to get the modal contribution list Rank t ;
[0122] Set the device ID, Rank t 、 and timestamp are encapsulated as a JSON message, as shown below:
[0123]
[0124]
[0125] Call the SMS API and SMTP service to simultaneously push alarm texts to the on-duty engineer's mobile phone and maintenance email.
[0126] It should be noted that this embodiment compares the equipment state vector with the health threshold and the fault threshold respectively, and generates and reports maintenance warning information when any component exceeds the corresponding threshold, thereby achieving effective conversion of the predicted state vector into actual maintenance decisions, ensuring that potential problems of the equipment can be discovered and responded to quickly, avoiding the amplification of fault events and the risk of sudden equipment failure, improving the response speed of maintenance decisions and the safety of equipment operation, and reducing operating and maintenance costs.
[0127] The aforementioned methods for timestamp alignment, preprocessing, and feature extraction of multimodal sensor data can be performed using methods and means in the prior art and will not be described in detail in this example.
[0128] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for predicting industrial equipment status based on multimodal sensor fusion, characterized in that: include: Collecting multimodal sensor data of the rotary tiller within a fixed time window, and aligning timestamps of the multimodal sensor data based on a unified time reference; performing preprocessing on the aligned multimodal sensing data to obtain a standardized data matrix; Using a feature extraction method, the standardized data matrix is subjected to time-frequency wavelet packet decomposition, Mel spectrum extraction, and exponential sliding average calculation, and each modal feature vector is output; Construct a dual-path Transformer fusion network consisting of a modality self-attention branch and a graph attention branch based on sensor space topology, fuse the feature vectors of each modality, and generate a fused feature tensor; Inputting the fused feature tensor into a gated recurrent unit timing prediction model based on an attention mechanism, outputting a device state vector with a predicted time span, wherein the device state vector represents at least a health index and a failure probability; The device state vector is compared with the health threshold and the fault threshold respectively. When any component exceeds the corresponding threshold, maintenance warning information is generated and reported.
2. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 1 is characterized in that: The multimodal sensing data collection system for the rotary tiller includes: a three-axis vibration acceleration sensor installed on the gearbox housing and the frame, an infrared temperature and K-type thermocouple combination sensor arranged near the main shaft bearing, a MEMS microphone array located on the side panel of the machine body, a Hall speed encoder embedded in the power output shaft, a three-phase current-voltage integrated sensing module placed at the input end of the electronically controlled hydraulic pump, and a high-precision air pressure-humidity combination sensor installed in the machine cover.
3. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 2 is characterized in that: The multimodal sensing data includes mechanical vibration acceleration data, infrared temperature measurement values and thermocouple temperature measurement values, sound waveform sequence, speed pulse count and instantaneous speed, phase current and phase voltage and power factor, ambient temperature, humidity and pressure.
4. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 1 or 3, characterized in that: Performing time stamp alignment on the multimodal sensor data based on a unified time reference includes: Using the GPS-PPS pulse signal as the external master clock, the local clock of each acquisition node is calibrated once a day with a maximum of 0.1ms in frequency and phase. The data streams with different sampling frequencies are resampled to the reference frequency of 1kHz using linear interpolation; The missing time segments are marked as missing and inpainted using local spline interpolation with a sliding window width of 5s.
5. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 4 is characterized in that: Preprocessing is performed on the aligned multimodal sensor data to obtain a standardized data matrix, including: Apply fourth-order Butterworth bandpass filtering to the mechanical vibration acceleration data, speed pulse counts, and instantaneous speed and sound waveform sequences; Outliers were removed by using the isolation forest algorithm, and the removal threshold was 0.6 times the theoretical upper limit of the average path length; Apply the 3σ principle to the phase current signal, phase voltage signal and power to eliminate pulse spikes; Bidirectional exponential sliding smoothing is used for infrared temperature measurement values, thermocouple temperature measurement values, ambient temperature, humidity and pressure, and the smoothing factor is defined as 0.3; All channels are Z-score normalized, and a two-dimensional standardized data matrix is constructed according to the sensor number-time step method.
6. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 5 is characterized in that: Outputting each modal eigenvector includes: The vibration signal is decomposed into five layers of Daubechies-4 wavelet packets to extract 63-dimensional features including energy entropy, kurtosis and main frequency of each sub-band. A Mel spectrum with a frame length of 25ms and a frame shift of 10ms is generated for the acoustic emission signal. The number of filter banks is 64. The logarithmic amplitude is taken and global average pooling is performed to obtain a 64-dimensional acoustic feature. An exponential sliding average with a window length of 100ms is used for the three-phase current and voltage to output real-time RMS, peak-to-peak value, and phase angle 7-dimensional features. The residuals are retained after the second-order polynomial trend decomposition of the environmental parameters as the three-dimensional interference compensation features.
7. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 6 is characterized in that: Generating the fused feature tensor includes: The modal self-attention branch uses 8-head attention, defines the key value dimension as 64, and the output dimension as 256; The sensor space topology is constructed as an undirected weighted graph based on the physical coordinates and functional correlation of the sensors. The edge weight is the weighted sum of the inverse of the normalized Euclidean distance and the mutual information score. The graph attention branch uses a two-layer graph attention network with 4 attention heads in each layer and a node embedding dimension of 128; The outputs of the modal self-attention branch and the graph attention branch are cascaded according to the channel dimension and then linearly mapped and layer normalized to obtain a 384-dimensional fused feature tensor.
8. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 7 is characterized in that: Outputs the device state vector with the predicted time span, including: Input the fused feature tensor into a two-layer gated recurrent unit with 256 hidden units in each layer, and introduce an adaptive attention gate between the two layers; The prediction time span is a configurable discrete value within 5 to 30 minutes, and the optimal prediction time span is selected based on the least squares fitting residual; The normalized health index HI and failure probability FP are obtained using the Sigmoid activation function.
9. The method for industrial equipment state prediction based on multimodal sensor fusion according to claim 8, characterized in that: The device state vector is compared with the health threshold and the fault threshold respectively. When any component exceeds the corresponding threshold, maintenance warning information is generated and reported, including: Health threshold HI th Adaptive update based on 15% of historical 30-day normal operating data; Fault threshold FP th Defined as 0.6, when the failure probability FP exceeds the failure threshold FP for three consecutive prediction cycles th It is judged as a fault when When HI <HI th or FP>FP th When an abnormality occurs, a JSON structured warning message containing the device number, abnormal modal weight ranking and recommended downtime window is generated; The early warning information is pushed to the maintenance management platform via the MQTT protocol, and alerts are issued through multiple channels such as SMS and email.
Citation Information
Cited By
Industrial equipment control method based on multi-protocol fusion and related device
CN121433172A