Ship Fault Prediction System with Multi-Source Data Fusion
Through multi-source data fusion and feature engineering technology, combined with improved Kalman filtering algorithm and optimized SVM model, the difficulties of data fusion and feature extraction in ship failure prediction are solved, achieving higher prediction accuracy and reliability.
Patent Information
- Application Number
- CN202411208433.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-08-30
AI Technical Summary
The prior art is difficult to effectively integrate multi-source heterogeneous data in ship failure prediction and extract representative features, resulting in insufficient accuracy and reliability of fault prediction.
A ship failure prediction system with multi-source data fusion is adopted to achieve accurate prediction of ship failures through data acquisition, preprocessing, Kalman filtering algorithm fusion, feature engineering (including VMD decomposition and autoencoder) and optimized SVM models.
It significantly improves the accuracy and reliability of ship failure prediction, reduces operational risks and maintenance costs, can better handle nonlinearities and uncertainties, and adapt to complex ship operating environments.
Smart Images

Figure CN119179950B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ship fault prediction, and particularly to a ship fault prediction system based on multi-source data fusion. Background Art
[0002] The normal operation of ship engine room equipment, especially key components such as engines, transmission devices, and shaft systems in the propulsion system, is crucial for ship safety. In recent years, with the development of sensing technology, data analysis, and artificial intelligence, significant progress has been made in ship fault prediction and diagnosis technology. Traditional methods mainly rely on expert experience and simple threshold judgment, while modern methods more often adopt data-driven intelligent algorithms.
[0003] In terms of fault diagnosis, technologies such as vibration analysis, acoustic emission, and oil analysis are widely used in the detection of engine faults. Researchers have proposed signal processing methods based on wavelet transform, empirical mode decomposition (EMD), etc. to extract fault features. At the same time, machine learning algorithms such as support vector machine (SVM), artificial neural network (ANN), etc. are used for fault mode recognition and classification.
[0004] However, there are still some limitations in the existing technologies. First, a single data source often fails to comprehensively reflect the complex state of the equipment. Second, the correlation between different types of data (such as vibration, temperature, pressure, etc.) has not been fully utilized. Third, existing feature extraction methods may not be able to effectively capture the key information in multi-source heterogeneous data. Finally, in practical applications, the impact of data quality problems (such as noise, missing, etc.) on prediction accuracy has not been well solved.
[0005] Therefore, how to effectively fuse multi-source heterogeneous data, extract more representative features, and on this basis improve the accuracy and reliability of fault prediction has become an important research direction in the current field of ship fault prediction. Summary of the Invention
[0006] In view of this, the present invention proposes a ship fault prediction system based on multi-source data fusion, which effectively integrates various sensor data of ship engine room equipment through multi-source data fusion and feature engineering technology, and uses an improved Kalman filter algorithm and an optimized SVM model to process complex and nonlinear systems, so as to extract more representative features, realize accurate prediction of ship faults, and ultimately improve the safety and reliability of ship operation.
[0007] The technical solution of the present invention is implemented as follows: The present invention provides a ship fault prediction system based on multi-source data fusion, including:
[0008] A data acquisition module for acquiring the original signal data of the engine room equipment related to ship operation;
[0009] A data preprocessing module for aligning raw signal data from different sources. After data alignment, the preprocessed raw signal data is preprocessed according to the preprocessing steps to obtain preprocessed data;
[0010] A data fusion module for performing multi-source fusion on the preprocessed data using an improved Kalman filter algorithm to obtain fused data;
[0011] A feature engineering module for performing multi-dimensional feature extraction, feature selection on the fused data, and fusing the selected features to obtain fused features;
[0012] A fault prediction module for receiving the fused features and predicting ship faults based on the fused features using a pre-trained fault prediction model.
[0013] Based on the above technical solution, preferably, the data preprocessing module includes:
[0014] A data alignment unit for preliminarily checking the raw signal data from different sources, standardizing the timestamps, creating a unified time index, mapping the raw signal data from different sources to the unified time index, and making the raw signal data in a time series format to complete data alignment;
[0015] A data preprocessing unit for preprocessing the raw signal data after data alignment, including data interpolation and outlier filtering, to obtain preprocessed data.
[0016] Based on the above technical solution, preferably, the improved Kalman filter algorithm includes:
[0017] Based on the characteristics of the engine room equipment related to ship operation, a state space model containing key state variables is established. This state space model consists of a state equation and an observation equation, which respectively describe the evolution of the system state and the observation process. Among them, the key state variables include engine speed, engine temperature, fuel consumption rate, exhaust temperature, cooling water temperature, and lubricating oil pressure;
[0018] Obtain the initial state space model parameters and optimize the parameters through an optimization method to obtain an optimal parameter set;
[0019] Perform Kalman filter algorithm iteration on the preprocessed data:
[0020] Prediction step: Predict the current state based on the previous state estimate and control input;
[0021] Unknown input estimation: Estimate the slowly changing and rapidly changing unknown input terms respectively;
[0022] Update step: Combine the observed values and the position input estimate to update the state estimate and the error covariance matrix;
[0023] Adjustment step: Dynamically adjust the process noise covariance matrix and the observation noise covariance matrix according to the innovation sequence;
[0024] Fuse data from different sources by expanding the observation equation, and adjust the observation noise covariance matrix according to the reliability of each data source;
[0025] Output the denoised and fused state estimate value as the fused data.
[0026] On the basis of the above technical solution, preferably, in the state space model:
[0027] State equation:
[0028] x(k + 1) = A(k)x(k) + B(k)u(k) + F(k)d(k) + w(k)
[0029] Observation equation:
[0030] y(k) = C(k)x(k) + v(k)
[0031] Wherein, x(k + 1) is the state vector at time k + 1, x(k) is the state vector at time k, u(k) is the control input vector, d(k) is the unknown input vector, w(k) is the process noise, A(k) is the state transition matrix, B(k) is the control input matrix, F(k) is the unknown input matrix, y(k) is the observation vector, C(k) is the observation matrix, and v(k) is the observation noise;
[0032] In the iteration of the Kalman filter algorithm:
[0033] Prediction step:
[0034]
[0035] P(k|k - 1) = A(k - 1)P(k - 1|k - 1)A(k - 1) T + Q(k - 1)
[0036] Unknown input estimate:
[0037] M(k) = C(k)P(k|k - 1)C(k) T + R(k)
[0038] K d (k) = P(k|k - 1)C(k) T M(k) -1
[0039]
[0040] Update steps:
[0041] K(k) = P(k|k - 1)C(k) T (C(k)(k|k - 1)C(k) T +R(k)) -1
[0042]
[0043] P(k|k) = (I - K(k)C(k))P(k|k - 1)
[0044] Adjustment steps:
[0045]
[0046] S(k) = C(k)P(k|k - 1)C(k) T +R(k)
[0047] Q(k) = λQ(k - 1)+(1 - λ)K(k)ε(k)ε(k) T K(k) T
[0048] R(k) = λR(k - 1)+(1 - λ)(ε(k)ε(k) T -C(k)P(k|k - 1)C(k) T )
[0049] Wherein, is the prior state estimate at time k, is the posterior state estimate at time k - 1, P(k|k - 1) is the prior error covariance matrix at time k, P(k - 1|k - 1) is the posterior error covariance matrix at time k - 1, Q(k - 1) is the process noise covariance matrix, R(k) is the observation noise covariance matrix, M(k) is the innovation covariance matrix, K d (k) is the unknown input gain, is the slowly varying unknown input estimate, is the rapidly varying unknown input estimate, ρ(k) is the intermittency coefficient matrix, θ(k) is the rapidly varying part parameter, K(k) is the Kalman gain, is the posterior state estimate at time k, P(k|k) is the posterior error covariance matrix at time k, F(k) is the unknown input matrix, I is the identity matrix, ε(k) is the innovation sequence, S(k) is the innovation covariance matrix, λ is the forgetting factor;
[0050] Expand the observation equation for data fusion:
[0051] y(k) = [C1(k); C2(k);...; C a (k)]x(k) + v(k)
[0052] Wherein, C1(k); C2(k);...; C a (k) respectively represent the observation matrices of different data sources.
[0053] Based on the above technical solution, preferably, the initial state space model parameters are estimated based on the maximum likelihood estimation method, and the optimization method is a hybrid optimization method based on particle swarm optimization and simulated annealing.
[0054] Based on the above technical solution, preferably, the feature engineering module includes:
[0055] A feature extraction unit, configured to perform signal decomposition on the fused data according to the VMD decomposition algorithm, separate K IMF components from each data, and use the trained autoencoder to perform feature extraction on the K IMF components respectively to obtain multi-dimensional features of the data;
[0056] A feature selection unit, configured to calculate the relative chaos degree of the multi-dimensional features, sort the multi-dimensional features in ascending order of the relative chaos degree, and select the first M-dimensional features as the features to be fused;
[0057] A feature fusion unit, configured to fuse the features to be fused of each data according to the fusion formula to obtain fused features.
[0058] Based on the above technical solution, preferably, the loss function during the training of the autoencoder is:
[0059]
[0060] Wherein, m is the number of samples, s represents the s-th sample, O s represents the output data, I s represents the input data, when z = 1, it represents is the weight matrix of the encoder, when z = 2, it represents is the weight matrix of the decoder, r h is the number of hidden layers of the encoder, r g is the number of hidden layers of the decoder, γ is the weight decay parameter, SG[·] is to prohibit the backpropagation of gradients, h e (s) is the vector representation of mapping the sample s to the latent space by the encoder, e is the embedding vector, represents the squared L2 norm.
[0061] Based on the above technical solution, preferably, the relative chaos degree calculation formula is:
[0062]
[0063] wherein, RD(X i ) is the relative chaos degree of the i-th dimensional feature X i , AC(X i ) is the first-order autocorrelation coefficient of X i , is the average mutual information of X i , N is the total dimension of the multi-dimensional feature, MI(X i , X j ) is the mutual information between X i and X j , H(X i ) is the information entropy of X i , and α and β are adjustable parameters used to adjust the influence degrees of the autocorrelation coefficient and the mutual information.
[0064] Based on the above technical solution, preferably, the fusion formula is:
[0065]
[0066] wherein, F is the fusion feature, and w i is the weight.
[0067] Based on the above technical solution, preferably, the fault prediction model is SVM. During pre-training, an optimization algorithm is used to optimize the penalty parameter and the kernel function parameter of SVM, and after optimization, SVM is trained based on the sample set.
[0068] The present invention has the following beneficial effects compared with the prior art:
[0069] (1) By integrating multiple modules, the whole process optimization from data acquisition, preprocessing, fusion to feature extraction and fault prediction is realized. This systematic method significantly improves the accuracy and reliability of ship fault prediction, and effectively reduces the ship operation risk and maintenance cost;
[0070] (2) The improved Kalman filtering algorithm can better handle non-linearity and uncertainty by introducing unknown input estimation and dynamically adjusting the noise covariance matrix, and improves the accuracy of data fusion. At the same time, the design of the extended observation equation makes the fusion of multi-source data more flexible and effective, and enhances the adaptability of the system to the complex ship operation environment;
[0071] (3) The feature extraction method combining the VMD decomposition algorithm and the autoencoder, as well as the feature selection strategy based on relative chaos degree, can extract the most representative and informative features from complex ship operation data. This not only reduces the data dimension, but also improves the efficiency and accuracy of subsequent fault prediction;
[0072] (4) By optimizing and pre-training the SVM model parameters, this method enhances the model's learning ability for ship fault modes. This optimized SVM model can capture various complex fault modes more accurately, improving the accuracy and reliability of fault prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0074] Figure 1 It is the system framework diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0076] As Figure 1 shown, the present invention provides a ship fault prediction system for multi-source data fusion, including:
[0077] A data acquisition module for acquiring the original signal data of the engine room equipment related to the ship operation;
[0078] A data preprocessing module for aligning the original signal data from different sources. After data alignment, the original signal data after data alignment is preprocessed according to the preprocessing steps to obtain preprocessed data;
[0079] A data fusion module for performing multi-source fusion on the preprocessed data by using an improved Kalman filtering algorithm to obtain fused data;
[0080] A feature engineering module for performing multi-dimensional feature extraction, feature selection on the fused data, and fusing the selected features to obtain fused features;
[0081] A fault prediction module for receiving the fused features and performing ship fault prediction based on the pre-trained fault prediction model according to the fused features.
[0082] Specifically, in one embodiment of the present invention, the data acquisition module targets the following engine room equipment in a ship: main engine system, auxiliary power generation system, fuel system, cooling system, lubrication system, and exhaust system. The original signal data is collected from the operating states of these engine room equipment using sensors, data acquisition cards, etc. The specific data types include physical quantities such as temperature, pressure, and vibration measured by sensors, operating parameters such as engine speed, fuel consumption rate, and exhaust temperature. Additionally, some environmental data such as sea conditions, air temperature, and humidity are collected as backup data. When collecting data, the sampling frequency is set according to the characteristics of different equipment and parameters. For key parameters such as engine speed, engine temperature, fuel consumption rate, exhaust temperature, cooling water temperature, and lubricating oil pressure, high-frequency sampling can be set, and for other parameters, low-frequency sampling can be set.
[0083] Specifically, in one embodiment of the present invention, the data preprocessing module includes:
[0084] A data alignment unit, which is used to perform a preliminary check on the original signal data from different sources, standardize the timestamps, create a unified time index, map the original signal data from different sources to the unified time index, and make the original signal data in a time series format to complete data alignment;
[0085] A data preprocessing unit, which is used to preprocess the original signal data after data alignment, including data interpolation and outlier filtering, to obtain the preprocessed data.
[0086] In this embodiment, the preliminary data check includes integrity check, format check, and range check. For example, check whether there is missing data, ensure that the formats of all data sources are consistent, and verify whether the data is within a reasonable range. Timestamp standardization means converting the timestamps of all data sources into the same format, such as the ISO 8601 standard (YYYY-MM-DD HH:MM:SS). Ensure that all timestamps have the same precision, for example, accurate to seconds or milliseconds.
[0087] In this embodiment, the process of creating a unified time index is as follows: Based on the timestamps of all data sources, determine the overall start and end times; according to the nature and sampling frequency of the data, determine an appropriate time interval; use the determined time range and interval to create a continuous time series as the unified index. According to the sampling frequency set when collecting data, for data with a sampling frequency lower than the unified index, linear interpolation is used to fill in the missing time points, and for data with a sampling frequency higher than the unified index, downsampling is performed by averaging or other appropriate methods to map the data to the unified index and complete data alignment.
[0088] Specifically, in one embodiment of the present invention, the data interpolation process is:
[0089] Obtain the missing data detected during the preliminary inspection. For each piece of missing data, use it as the data T to be interpolated, set a time period, and centered on T, obtain the data within the previous time period and the next time period as candidate data;
[0090] Calculate the Euclidean distance d between each data in the candidate data and the data T to be interpolated;
[0091] Arrange the candidate data in ascending order according to the Euclidean distance d;
[0092] Select the first k candidate data as the nearest neighbor dataset;
[0093] Use the interpolation formula to calculate the interpolation result of the data T to be interpolated based on the nearest neighbor dataset.
[0094] Among them, the interpolation formula is as follows:
[0095]
[0096] In the formula, y(T) is the interpolation result of T, w i is the weight calculated based on the Euclidean distance, x i represents the i-th nearest neighbor data, i = 1, 2,..., k, d(T, x i ) represents the Euclidean distance between the nearest neighbor data x i and T, and p is the weight exponent, taking a value of 1 or 2.
[0097] Specifically, in an embodiment of the present invention, outlier filtering uses the k-means algorithm to identify outliers. Use the k-means algorithm to iteratively divide the data into clustering clusters. After the iteration ends, calculate the distance between each data point and its clustering center within each cluster, and regard the data points with a distance greater than the threshold as outliers. According to the outlier situation of the outliers and the importance of the outliers, select to repair the outliers or directly remove them. The repair method can regard the outliers as missing points and use the above data interpolation method to interpolate them.
[0098] Specifically, in an embodiment of the present invention, the improved Kalman filter algorithm includes:
[0099] Based on the characteristics of the engine room equipment related to ship operation, establish a state space model including key state variables. This state space model consists of a state equation and an observation equation, which respectively describe the evolution of the system state and the observation process. Among them, the key state variables include engine speed, engine temperature, fuel consumption rate, exhaust temperature, cooling water temperature, and lubricating oil pressure;
[0100] Among them, the state equation:
[0101] x(k + 1) = A(k)x(k) + B(k)u(k) + F(k)d(k) + w(k)
[0102] Observation equation:
[0103] y(k) = C(k)x(k) + v(k)
[0104] Where x(k + 1) is the state vector at time k + 1, x(k) is the state vector at time k, u(k) is the control input vector, d(k) is the unknown input vector, w(k) is the process noise, A(k) is the state transition matrix, B(k) is the control input matrix, F(k) is the unknown input matrix, y(k) is the observation vector, C(k) is the observation matrix, and v(k) is the observation noise. This model describes the evolution of the system state and the observation process.
[0105] Obtain the initial state - space model parameters and optimize the parameters through an optimization method to obtain the optimal parameter set;
[0106] Specifically, the initial state - space model parameters are estimated based on the maximum - likelihood estimation method. The process includes: constructing the log - likelihood function, maximizing the likelihood function using the Newton - Raphson method to obtain the initial parameter estimates.
[0107] Specifically, the optimization method is a hybrid optimization method based on particle swarm optimization and simulated annealing. The optimization process is as follows:
[0108] a) Initialize the particle swarm. Each particle represents a set of possible parameter solutions. Set the number of particles as N and the parameter dimension as D. For each particle, its position is represented as X i , X i = [x i1 , x i2 ,..., x iD and its velocity is represented as V i , V i = [v i1 , v i2 ,..., v iD . Let the individual best position be P i , and the global best position be G = argmin{f(P i )}, i = 1, 2,..., N. f represents the fitness function.
[0109] b) Calculate the fitness value for each particle. The fitness function selects the root - mean - square error (RMSE).
[0110] c) Update the individual best position and the global best position of each particle;
[0111] For each particle i:
[0112] If f(X i ) < f(P i ), then P i = X i ;
[0113] If f(P i ) < G, then G = P i .
[0114] d) Update the velocity and position of the particle according to the velocity update formula and the position update formula.
[0115] The update formulas are as follows:
[0116] V i (t + 1) = χ * (ω * V i (t) + c1 * r1 * (P i - X i (t)) + c2 * r2 * (G - X i (t)))
[0117] X i (t + 1) = X i (t) + V i (t + 1)
[0118] Where: χ is the contraction factor, ω is the inertia weight, ω = ω max - (ω max - ω min ) * t / T max ; c1, c2 are acceleration constants; r1, r2 are random numbers uniformly distributed in [0, 1]; t is the current iteration number; T max is the maximum number of iterations.
[0119] Specifically, when adjusting c1 and c2, the following formulas are used for adjustment:
[0120] c1 = c 1min + (c 1max - c 1min ) * (T max - t) / T max
[0121] c2 = c 2max - (c 2max - c 2min ) * (T max - t) / T max .
[0122] e) Apply the simulated annealing operation to the global optimal solution G of each iteration to increase the probability of jumping out of the local optimum:
[0123] Generate a new solution θ after perturbation new = G + σ * N(0, 1), where σ is the perturbation intensity and N(0, 1) is the standard normal distribution.
[0124] Calculate the energy difference ΔE = f(θ new ). - f(G).
[0125] If ΔE < 0, G = θ new ; Otherwise, accept the new solution θ with probability p = exp(-ΔE / T) new , where T is the current temperature.
[0126] Reduce the temperature according to the exponential decay rule: T = αT, where α is the cooling coefficient and its value ranges from 0.9 to 1.
[0127] f) Repeat steps b - e until the maximum number of iterations is reached or the convergence condition is satisfied. Output the optimal parameter solution G and the corresponding fitness value f(G).
[0128] This hybrid optimization method improves the convergence speed and global search ability of the algorithm by introducing a contraction factor, adaptive parameter adjustment, and local search strategy. At the same time, the simulated annealing operation enhances the algorithm's ability to jump out of local optima, making it more suitable for dealing with complex parameter optimization problems.
[0129] Iterate the Kalman filter algorithm on the pre - processed data:
[0130] Prediction step: Predict the current state based on the previous state estimate and control input;
[0131]
[0132] P(k|k - 1) = A(k - 1)P(k - 1|k - 1)A(k - 1) T + Q(k - 1)
[0133] Unknown input estimation: Estimate the slow - changing and fast - changing unknown input terms respectively;
[0134] M(k) = C(k)P(k|k - 1)C(k) T + R(k)
[0135] K d (k) = P(k|k - 1)C(k) T M(k) -1
[0136]
[0137]
[0138] Update step: Update the state estimate and the error covariance matrix by combining the observation value and the position input estimate;
[0139] K(k) = P(k|k - 1)C(k) T (C(k)(k|k - 1)C(k) T + R(k)) -1
[0140]
[0141] P(k|k) = (I - K(k)C(k))P(k|k - 1)
[0142] Adjustment step: Dynamically adjust the process noise covariance matrix and the observation noise covariance matrix according to the innovation sequence;
[0143]
[0144] S(k) = C(k)P(k|k - 1)C(k) T + R(k)
[0145] Q(k) = λQ(k - 1)+(1 - λ)K(k)ε(k)ε(k) T K(k) T
[0146] R(k) = λR(k - 1)+(1 - λ)(ε(k)ε(k) T - C(k)P(k|k - 1)C(k) T )
[0147] Wherein, is the prior state estimate at time k, is the posterior state estimate at time k - 1, P(k|k - 1) is the prior error covariance matrix at time k, P(k - 1|k - 1) is the posterior error covariance matrix at time k - 1, Q(k - 1) is the process noise covariance matrix, R(k) is the observation noise covariance matrix, M(k) is the innovation covariance matrix, K d (k) is the unknown input gain, is the slowly varying unknown input estimate, is the fast varying unknown input estimate, ρ(k) is the intermittency coefficient matrix, θ(k) is the fast varying part parameter, K(k) is the Kalman gain, is the posterior state estimate at time k, P(k|k) is the posterior error covariance matrix at time k, F(k) is the unknown input matrix, I is the identity matrix, ε(k) is the innovation sequence, S(k) is the innovation covariance matrix, λ is the forgetting factor.
[0148] By expanding the observation equation, data from different sources are fused, and the observation noise covariance matrix is adjusted according to the reliability of each data source;
[0149] y(k) = [C1(k); C2(k);...; C a (k)]x(k) + v(k)
[0150] In the formula, C1(k); C2(k);...; C a (k) respectively represent the observation matrices of different data sources.
[0151] Output the state estimation value after denoising and fusion as the fused data.
[0152] Specifically, this algorithm takes into account the influence of unknown inputs, and separately estimates the slowly varying and rapidly varying unknown inputs. Dynamically adjusts the process noise and observation noise covariance matrices, improving the self - adaptability of the algorithm. Achieves multi - source data fusion by expanding the observation equation, improving the accuracy of estimation.
[0153] Specifically, in an embodiment of the present invention, the feature engineering module includes:
[0154] A feature extraction unit, which is used to decompose the fused data according to the VMD decomposition algorithm. Each data is separated into K IMF components, and the trained auto - encoder is used to extract features from the K IMF components respectively to obtain the multi - dimensional features of the data;
[0155] A feature selection unit, which is used to calculate the relative chaos degree of the multi - dimensional features, sort the multi - dimensional features from small to large according to the relative chaos degree, and select the first M - dimensional features as the features to be fused;
[0156] A feature fusion unit, which is used to fuse the features to be fused of each data according to the fusion formula to obtain the fused features.
[0157] In this embodiment, the feature extraction unit first uses the variational mode decomposition (VMD) algorithm to decompose the fused data. Each data is separated into K intrinsic mode function (IMF) components. The purpose of this step is to decompose the complex signal into multiple simple components. Then, an auto - encoder model is designed and trained. The auto - encoder is an unsupervised learning algorithm used to learn the effective representation (encoding) of data. The following loss function is used in the training process:
[0158]
[0159] In the formula, m is the number of samples, s represents the s - th sample, O s represents the output data, I s represents the input data, when z = 1, it means is the weight matrix of the encoder. When z = 2, it means is the weight matrix of the decoder, r h is the number of hidden layers of the encoder, r g is the number of hidden layers of the decoder, γ is the weight decay parameter, SG[·] is to prohibit the gradient from backpropagating, h e (s) is the vector representation that the encoder maps the sample s to the latent space, e is the embedding vector, represents the squared L2 norm.
[0160] Use the trained autoencoder to extract features from each IMF component. The specific steps are as follows:
[0161] a) Take each IMF component as the input of the autoencoder.
[0162] b) Use the encoder part to map the IMF component to the latent space to obtain its compressed representation.
[0163] c) This compressed representation is the extracted feature.
[0164] For each original data, there are K IMF components now, and each component has its corresponding feature. Combine these features to form the multi-dimensional feature representation of this data.
[0165] Repeat the above process for all the fused data, and finally obtain the multi-dimensional feature representation of each data.
[0166] In this embodiment, the loss function of the autoencoder includes three terms. The first term is the reconstruction error term, which prompts the encoder to learn an effective representation of the input data, and the decoder can recover the original data from these representations; the second term is the regularization term, which prevents overfitting and simplifies the model by penalizing large weight values; the third term is the contrastive learning term, which guides the model to learn more meaningful and structured feature representations. Designing this loss function is to enable the autoencoder to learn a feature representation that can both reconstruct the original data and has a specific structure through the combined action of the reconstruction error term and the contrastive learning term.
[0167] Specifically, in this embodiment, the feature selection unit selects the most valuable feature subset from the multi-dimensional features. For this purpose, a calculation method of relative chaos is introduced. This relative chaos comprehensively considers the self-correlation of the features, the mutual information between the features, and the uncertainty of the features to represent the value of the features. The specific calculation formula is as follows:
[0168]
[0169] In the formula, RD(X i ) is the relative chaos of the i-th dimensional feature X i , AC(X i) is X i is the first-order autocorrelation coefficient of is X i is the average mutual information of, N is the total dimension of the multi-dimensional feature, MI(X i , X j ) is X i and X j is the mutual information between, H(X i ) is the information entropy of X i , α and β are adjustable parameters used to adjust the influence degree of the autocorrelation coefficient and the mutual information.
[0170] In the calculation formula of the relative chaos degree, the numerator part considers the autocorrelation of the feature. When AC(X i ) is relatively large, tends to 0, and the numerator tends to 1. High autocorrelation means that the feature is more stable in the time series and is more likely to contain useful information. The denominator part comprehensively considers the product of the average mutual information, autocorrelation, and mutual information. Among them, the average mutual information reflects the overall correlation between the feature and other features, and log(1 + AC(X i ) × MI(X i )) reflects the joint effect of autocorrelation and mutual information. In the calculation of the average mutual information, normalization is performed according to the information entropy to make the mutual information of different features comparable.
[0171] After calculating the RD value of each feature, all features are sorted in ascending order of the RD value, and the first M features with the smallest RD value are selected as the features to be fused. The smaller the RD value, the more important the feature. Important features usually have high autocorrelation, high mutual information, and low entropy.
[0172] Through the calculation of the relative chaos degree and feature selection, the most valuable information can be effectively extracted from complex multi-source ship data, laying a foundation for subsequent feature fusion and fault prediction.
[0173] Specifically, in this embodiment, the feature fusion unit constructs a fusion formula according to the calculation result of the relative chaos degree and fuses the selected features to be fused. The fusion formula is as follows:
[0174]
[0175] In the formula, F is the fused feature, w i is the weight.
[0176] Specifically, this fusion formula uses the relative chaos degree as the reciprocal of the weight, realizing an adaptive feature fusion method. It can not only effectively integrate the information of multiple features, but also automatically emphasize important features and weaken unimportant features, thus providing a high-quality and low-dimensional input feature for the ship fault prediction system.
[0177] Specifically, in one embodiment of the present invention, the fault prediction model adopts SVM. During pre-training, an optimization algorithm is used to optimize the penalty parameter and kernel function parameter of SVM, and after optimization, SVM is trained based on the sample set. The optimization algorithm can adopt the sparrow optimization algorithm. Inputting the fusion features into SVM can obtain the fault diagnosis result.
[0178] It should be noted that in the process of each module of the present invention executing various algorithms to achieve the purpose of fault prediction, the specific process is as follows: The data acquisition module collects data of multiple engine room devices from different sources. The engine room devices include the main engine system, auxiliary power generation system, fuel system, cooling system, lubrication system, and exhaust system. The data preprocessing module preprocesses the original data. The data fusion module uses an improved Kalman filtering algorithm to fuse the data from different sources to obtain the fused data, that is, each engine room device will have a set of fused time series data. The feature engineering module extracts features from the fused data, decomposes each data into k IMF components, and then extracts multiple features from each IMF component. For example, if each IMF component extracts B features, then each data finally extracts K + B-dimensional features, and then feature selection is performed, that is, from these K + B dimensions, the most important and valuable M dimensions are selected, and then each data will obtain M-dimensional features to be fused. During feature fusion, the features to be fused of each data are fused to finally obtain the fused features of each data. Finally, the fused features of all data are input into SVM to diagnose the faults of the ship. Since each data represents an engine room device, the diagnosis result of the corresponding fused features is the diagnosis result of the engine room device.
[0179] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. The ship fault prediction system based on multi-source data fusion is characterized by: include: Data acquisition module, used to obtain the original signal data of the engine room equipment related to the ship operation; A data preprocessing module is used to align the original signal data from different sources. After the data is aligned, the original signal data after the data alignment is preprocessed according to the preprocessing steps to obtain the preprocessed data; The data fusion module is used to perform multi-source fusion on the pre-processed data using an improved Kalman filter algorithm to obtain fused data; The feature engineering module is used to extract and select multi-dimensional features from the fused data, and fuse the selected features to obtain fused features; A fault prediction module, used for receiving fused features and performing ship fault prediction according to the fused features based on a pre-trained fault prediction model; The improved Kalman filter algorithm includes: Based on the characteristics of the engine room equipment related to ship operation, a state space model containing key state variables is established. The state space model consists of state equations and observation equations, which respectively describe the evolution and observation process of the system state. Among them, the key state variables include engine speed, engine temperature, fuel consumption rate, exhaust temperature, cooling water temperature and lubricating oil pressure. Obtain the initial state space model parameters, and optimize the parameters through the optimization method to obtain the optimal parameter set; Iterate the Kalman filter algorithm on the preprocessed data: Prediction step: predict the current state based on previous state estimates and control inputs; Unknown input estimation: Estimate slowly changing and rapidly changing unknown inputs separately; Update step: Combine the observations and position input estimates to update the state estimate and error covariance matrix; Adjustment step: dynamically adjust the process noise covariance matrix and the observation noise covariance matrix according to the innovation sequence; By extending the observation equation, data from different sources are fused, and the observation noise covariance matrix is adjusted according to the reliability of each data source; Output the denoised and fused state estimate as the fused data; In the state space model: Equation of state: x(k+1)=A(k)x(k)+B(k)u(k)+F(k)d(k)+w(k) Observation equation: y(k)=C(k)x(k)+v(k) Where x(k+1) is the state vector at time k+1, x(k) is the state vector at time k, u(k) is the control input vector, d(k) is the unknown input vector, w(k) is the process noise, A(k) is the state transfer matrix, B(k) is the control input matrix, F(k) is the unknown input matrix, y(k) is the observation vector, C(k) is the observation matrix, and v(k) is the observation noise; Kalman filter algorithm iteration: Prediction steps: P(k|k-1)=A(k-1)P(k-1|k-1)A(k-1) T +Q(k-1) Unknown input estimation: M(k)=C(k)P(k|k-1)C(k) T +R(k) K d (k)=P(k|k-1)C(k) T M(k) -1 Update steps: K(k)=P(k|k-1)C(k) T (C(k)(k|k-1)C(k) T +R(k)) -1 P(k|k)=(IK(k)C(k))P(k|k-1) Adjustment steps: S(k)=C(k)P(k|k-1)C(k) T +R(k) Q(k)=λQ(k-1)+(1-λ)K(k)ε(k)ε(k) T K(k) T R(k)=λR(k-1)+(1-λ)(ε(k)ε(k) T -C(k)P(k|k-1)C(k) T ) In the formula, is the prior state estimate at time k, is the a priori state estimate at time k-1, P(k|k-1) is the a priori error covariance matrix at time k, P(k-1|k-1) is the a priori error covariance matrix at time k-1, Q(k-1) is the process noise covariance matrix, R(k) is the observation noise covariance matrix, M(k) is the innovation covariance matrix, K d (k) is the unknown input gain, For slowly varying unknown input estimates, is the estimation of the rapidly changing unknown input, ρ(k) is the intermittent coefficient matrix, θ(k) is the rapidly changing part parameter, K(k) is the Kalman gain, is the posterior state estimate at time k, P(k|k) is the posterior error covariance matrix at time k, F(k) is the unknown input matrix, I is the identity matrix, ε(k) is the innovation sequence, S(k) is the innovation covariance matrix, and λ is the forgetting factor; Extend the observation equation for data fusion: y(k)=[C1(k);C2(k);...;C a (k)]x(k)+v(k) Where, C1(k); C2(k); ...; C a (k) represent the observation matrices of different data sources.
2. The ship fault prediction system based on multi-source data fusion according to claim 1, characterized in that: The data preprocessing module includes: The data alignment unit is used to perform preliminary checks on the original signal data from different sources, standardize the timestamps, create a unified time index, map the original signal data from different sources to the unified time index, make the original signal data in time series format, and complete data alignment; The data preprocessing unit is used to preprocess the original signal data after data alignment, including data interpolation and outlier filtering, to obtain preprocessed data.
3. The ship fault prediction system based on multi-source data fusion according to claim 1, characterized in that: The initial state space model parameters are estimated based on the maximum likelihood estimation method, and the optimization method is a hybrid optimization method based on particle swarm optimization and simulated annealing.
4. The ship fault prediction system based on multi-source data fusion according to claim 1, characterized in that: Feature engineering modules include: The feature extraction unit is used to perform signal decomposition on the fused data according to the VMD decomposition algorithm, separate K IMF components from each data, and use the trained autoencoder to perform feature extraction on the K IMF components respectively to obtain the multi-dimensional features of the data; The feature selection unit is used to calculate the relative disorder of the multi-dimensional features, sort the multi-dimensional features from small to large according to the relative disorder, and select the first M dimensional features as the features to be fused; The feature fusion unit is used to fuse the features to be fused of each data according to the fusion formula to obtain the fused features.
5. The ship fault prediction system of multi-source data fusion according to claim 4, characterized in that: The loss function of the autoencoder during training is: In the formula, m is the number of samples, s represents the sth sample, O s Indicates output data, I s Represents input data. When z=1, it means is the weight matrix of the encoder, when z=2, it means is the weight matrix of the decoder, r h is the number of hidden layers of the encoder, r g is the number of hidden layers in the decoder, γ is the weight decay parameter, SG[·] is the prohibition of gradient backpropagation, and h e (s) is the vector representation of the encoder mapping sample s to the latent space, e is the embedding vector, represents the squared L2 norm.
6. The ship fault prediction system of multi-source data fusion according to claim 4, characterized in that: The relative disorder calculation formula is: In the formula, RD(X i ) is the i-th dimension feature X i The relative disorder degree, AC(X i ) is X i The first-order autocorrelation coefficient of For X i The average mutual information of N is the total dimension of multidimensional features, MI(X i ,X j ) is X i and X j The mutual information between i ) is X i The information entropy of , α and β are adjustable parameters used to adjust the influence of autocorrelation coefficient and mutual information.
7. The ship fault prediction system of multi-source data fusion according to claim 6, characterized in that: The fusion formula is: In the formula, F is the fusion feature, w i is the weight.
8. The ship fault prediction system of multi-source data fusion according to claim 1, characterized in that: The fault prediction model is SVM. During pre-training, the optimization algorithm is used to optimize the penalty parameters and kernel function parameters of SVM. After optimization, the SVM is trained based on the sample set.
Citation Information
Patent Citations
Offshore wind turbine generator operation state evaluation method and related device
CN118242232A