Bearing fault diagnosis method and system based on image feature fusion
By using image feature fusion and optimizing SVMD and ITSMIxer models with RSNPD, the problem of insufficient multi-dimensional feature extraction in existing bearing fault diagnosis is solved, achieving higher diagnostic accuracy and reliability.
Patent Information
- Application Number
- CN202510816391.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-17
AI Technical Summary
Most existing bearing fault diagnosis methods are based on the analysis of one-dimensional vibration signals and lack the extraction and fusion of multi-dimensional features of the signals. This makes it difficult to fully capture signal feature information when processing non-stationary signals under complex working conditions, resulting in low accuracy and reliability of the diagnostic results.
A bearing fault diagnosis method based on image feature fusion is adopted. The successive variational mode decomposition (SVMD) is optimized by the reverse spiral neural population dynamics algorithm (RSNPD) to decompose the vibration signal into modal component signals and convert them into RGB images and two-dimensional image signals. The improved time series mixer (ITSMixer) model is then used for feature extraction and diagnosis.
It significantly improves the accuracy of bearing fault diagnosis. Through multi-scale temporal feature capture and sparse attention mechanism, it achieves accurate modal separation and fault identification of complex temporal signals.
Smart Images

Figure CN120804866A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a bearing fault diagnosis method and system, in particular to a bearing fault diagnosis method and system based on image feature fusion, and belongs to the technical field of bearing fault diagnosis. BACKGROUND
[0002] As one of the most critical and easily faulted parts in industrial equipment mechanical systems, bearings are long-term operated under high load. When a bearing fails or irreversibly wears, it may cause accidents and even huge economic losses. Therefore, effective health monitoring and fault diagnosis of bearings are of great significance to ensure the safe and stable operation of industrial equipment.
[0003] Traditional bearing fault diagnosis methods mainly include signal analysis-based methods and data-driven methods. Signal analysis-based methods, such as wavelet transform and empirical mode decomposition (EMD), extract fault features by decomposing and analyzing vibration signals. For example, wavelet transform can decompose vibration signals into components of different frequencies and time scales, achieving time-frequency positioning, but its parameter selection and adaptability have a great impact on the final results. Empirical mode decomposition (EMD) can decompose the original signal into multiple intrinsic mode functions (IMF), each of which represents a different frequency and amplitude in the signal, but it may have problems such as mode mixing and end effect when processing non-stationary signals.
[0004] With the development of artificial intelligence technology, data-driven fault diagnosis methods have gradually attracted attention. These methods mainly use machine learning and deep learning models, such as convolutional neural networks (CNN), recurrent neural networks (RNN), and support vector machines (SVM), to extract features and classify vibration signals. For example, CNN can automatically extract time-frequency features of vibration signals and is suitable for processing high-dimensional data, but it relies on a large amount of data and has high computational complexity. RNN and its variants, such as long short-term memory networks (LSTM), can handle sequence data and are suitable for time series analysis, but they may have problems such as gradient vanishing and low computational efficiency when processing long sequence data.
[0005] Currently, most existing bearing fault diagnosis methods are based on one-dimensional vibration signal analysis, lacking multi-dimensional feature extraction and fusion of signals. This makes it difficult to fully capture the feature information of signals when processing non-stationary signals under complex working conditions, resulting in low accuracy and reliability of diagnosis results. SUMMARY
[0006] The purpose of the present application is to provide a bearing fault diagnosis method and system based on image feature fusion, which can improve the accuracy of bearing fault diagnosis.
[0007] Technical scheme: The bearing fault diagnosis method based on image feature fusion provided by the application comprises:
[0008] (1) obtaining original vibration signal data of bearing fault, and classifying according to generated fault type and damage degree;
[0009] (2) optimizing SVMD parameters by RSNPD to obtain optimized SVMD;
[0010] (3) decomposing original vibration signal data by using optimized SVMD to obtain modal component signals;
[0011] (4) collecting decomposed modal component signals, randomly cutting three vibration signals from each modal component signal to generate different fault samples, converting the fault samples into RGB images to obtain time domain features of original signals;
[0012] (5) converting original vibration signal data into two-dimensional image signals by using CWT to obtain frequency domain features of original signals;
[0013] (6) mixing RGB images and two-dimensional image signals to obtain image data sets, inputting image features of the image data sets into an improved ITSMixer model for training, optimizing hyperparameters of the ITSMixer model by using RSNPD to obtain a bearing fault diagnosis model;
[0014] (7) using the trained and optimized bearing fault diagnosis model to diagnose bearing faults to obtain corresponding diagnosis results.
[0015] Further, the step (2) is specifically that: by using RSNPD, a lens imaging reverse learning strategy is used for random initialization of population positions, a spiral sine strategy is introduced to enhance disturbance of position update of optimal solutions, and diversity of the population is increased, comprising:
[0016] (21) setting an objective function of RSNPD as envelope entropy and initializing parameters, the parameters comprising: population size, iteration number, dimension size and penalty parameter a;
[0017] (22) using a lens imaging reverse learning strategy for random initialization of population positions, the formula being:
[0018]
[0019] wherein lb is a lower limit, ub is an upper limit, a i,jh is the length of the incident light, h* is the length of the refracted light;
[0020] (23) The fitness value is calculated according to the envelope entropy, and the optimal solution is obtained according to the calculated fitness:
[0021] fitness = min [E p (i)]
[0022]
[0023]
[0024] wherein E p is the envelope entropy of each IMF, is the envelope signal of each mode after Hilbert demodulation, p j is the normalized form of , and K is the number of IMFs;
[0025] (24) The RSNPD algorithm realizes the optimization process through the attractor trend strategy, coupling interference strategy and information projection strategy.
[0026] (25) The current solution is constantly updated through the RSNPD optimization process, and the global optimal solution is synchronously updated.
[0027] (26) It is judged whether the given maximum number of iterations is reached, if not, it is turned to step (23), and the optimal result is output within the specified number of iterations.
[0028] Further, the step (24) comprises:
[0029] (241) In the attractor trend strategy, the optimal neural state is taken as an attractor, and the neural states of the remaining neural groups are randomly converged to one of the attractors, the number of attractors is set as [c·nd], wherein c is the proportion of attractors in all neural states in the neural population, nd is the number of neural populations, and the formula is:
[0030]
[0031] wherein l is the proportion factor of the attractor trend, r1 is a random number in [0,1], att g is the gth attractor, g is a random number in {1,2,…,[c·nd]}, a i is the neural state of the ith neural population, W i is the zero-mean Gaussian noise vector of the ith neural state, and the covariance is U-L, U is the peak value of all neural states, and L is the minimum boundary of all neural states.
[0032] (242)Assuming that the interaction between neural populations occurs through additive or diffusive coupling, in which the input of a neural population is a function of the sum of the neural states of other neural populations, represented as the average a of the sum of all neural states, except for the optimal neural population add,i , the formula is as follows:
[0033]
[0034] where r2 is a random number in [0, 0.5], and nd is the number of neural populations;
[0035] In diffusive coupling, the input of a neural population is the sum of the difference between its neural state and the neural state of other neural populations a dif,i , the formula is as follows:
[0036]
[0037] where r3 is a random number in [0, 0.5], and a i is the neural state of the i-th neural population;
[0038] The influence of other neural populations is simulated using a combination of the two couplings, and the formula is:
[0039] a couple,i = d·(a add,i + a dif,i )
[0040] where d is the proportion factor of coupling disturbance;
[0041] (243)The spiral sine strategy is introduced to perturb the role stage, and the formula is:
[0042]
[0043] The new coupling disturbance strategy calculation formula is:
[0044]
[0045] where η b+1 is the spiral sine calculation formula, b is the iteration number, g is the spiral coefficient, which is in the range [-1, 1], and ψ1 is a random number;
[0046] (244)In the attraction trend, the information transmission between the attractor and the other neural state of the neural population is realized by a 1×D adjacency matrix C attract,i and a 1×D communication strength matrix R attract,i , and the formula is as follows:
[0047] a C_attract,i = C attract,i ·Rattract,i ·a attract,i
[0048] where C attract,i is a 1 x D matrix with binary random values, R attract,i is a 1 x D matrix with values randomly generated in [0, 1], a attract,i is the dynamic vector obtained from the attractor trend strategy;
[0049] In the coupling disturbance, the information transmission between neural populations is utilized by another 1 x D adjacency matrix C couple,i and a 1 x D communication strength matrix R couple,i , and the formula used is as follows:
[0050] a C_coupie,i = C coupie,i · R coupie,i · a coupie,i · r4· (1 - FE / FE max )
[0051] where C couple,i is a 1 x D matrix with binary random values, R couple,i is a 1 x D matrix with values randomly generated in [0, 1], a couple,i is the dynamic vector obtained from the coupling disturbance strategy, r4 is a random number in [0, 1], FE is the function evaluation number, FE max is the maximum function evaluation number, and FE gradually increases as the iteration process proceeds;
[0052] The neural state of each neural population is affected by the attractor and the neural states of other neural populations, and the neural state a new,i of the current neural population is updated as follows:
[0053] a new,i = a i + a C_attract,i + a C_coupie,i .
[0054] Further, the step (3) comprises:
[0055] (31) Assuming that the original vibration signal x(t) is decomposed into two signals, including K modes u K (t) and a residual signal x r (t), the formula is as follows:
[0056] x(t) = u K (t) + x r (t)
[0057] where t is the time step, and the residual signal x r(t) includes the unprocessed part x u (t) and the modal part is obtained as follows:
[0058]
[0059] (32) The Kth modal minimization satisfies the following criteria:
[0060]
[0061] wherein, is a quadratic penalty factor for reducing the interference of noise, δ(t) is a Dirac distribution function, ω K is the center frequency of the Kth modal, g is a rotation factor, * is a convolution operation, u K (t) is the subsequence of the Kth modal;
[0062] (33) The constraint is set to minimize the spectral overlap between u k (t) and x r (t), and a filter with the following frequency response is used:
[0063]
[0064] wherein, α is a penalty parameter, ω is a frequency, and the established constraint is:
[0065]
[0066] wherein, β K (t) is the impulse response in the filter, and ‖‖ represents the norm;
[0067] (34) The constraint is set to minimize the energy of u k (t) within ±5% of the center frequency of the modal, and a filter with the following frequency response is used:
[0068]
[0069] The established constraint is:
[0070]
[0071] wherein, β k (t) is the impulse response in the filter;
[0072] (35) The extraction task of the Kth modal is converted into a constrained minimization problem, expressed as follows:
[0073]
[0074] wherein, α is a parameter for balancing J1, J2, and J3;
[0075] (36) Constructing a quadratic penalty term and a Lagrange multiplier, establishing an augmented Lagrange function, reconstructing the original signal, and the expression is as follows:
[0076]
[0077] Where λ is the Lagrange multiplier coefficient, and after solving the optimal solution, the alternating direction multiplier method is used for continuous iteration To obtain the decomposed K components.
[0078] Further, the step (4) comprises:
[0079] (41) The modal component signal is expressed as:
[0080] B r ={b i1j1 ,b ij ∈R n×m},i1=1,2,…,n,j1=1,2,…,m
[0081] Where B r represents an n-row and m-column matrix, n represents the number of channels, and takes a value of 3, and m represents the number of sampling points;
[0082] (42) Randomly intercepting three vibration signals to generate different fault samples with a size of S t =D′×D′×3, D′ is the sample length, each sample is randomly sampled from [1, m-D′×D′] without overlapping; then according to the sampling point, the randomly generated sample length is k×k:
[0083]
[0084] Where PM i (α,β) is a pixel matrix of each channel, α=1,…,D′; β=1,…,D′; z=1,2,3; let N represent the total number of samples, represents the signal value, gt=1,2,…,N; h=1,2,…,D ′2 , after normalization, the value range of each pixel matrix is from 0 to 255;
[0085] (43) Fusing the three D′×D′ matrices into an RGB pixel matrix to obtain an RGB image; the RGB pixel matrix RGB img (α,β,c), c=1,2,3, is obtained from RGB img (α,β,c)=(PM z (α,β),c).
[0086] Further, the step (5) comprises:
[0087] (51) Set γ(t) is a mother wavelet, and a continuous wavelet function γ(t) is obtained by stretching and translating the wavelet function m′,n′
[0088]
[0089] Wherein, m' is a scaling factor, n' is a translation factor, t is time, γ is a mother wavelet function, and R is a real number;
[0090] (52) Given a square integrable signal f(t), the continuous wavelet transform expression of f(t) is:
[0091]
[0092] Wherein, γ' is the derivative of the continuous wavelet function;
[0093] (53) When C f (m', n') is inversely transformed, the following is obtained:
[0094]
[0095] Wherein, w is the frequency of wavelet transform, when the value of m' is greater than the reciprocal of the sampling frequency of the signal, the low frequency characteristics in the signal are extracted, otherwise, the high frequency characteristics in the signal are extracted.
[0096] Further, the step (6) comprises:
[0097] (61) Mix the RGB image and the two-dimensional image signal to obtain an image data set, and input the images in the data set to the ITSMixer model;
[0098] (62) Given an input matrix Time projection TP is a linear layer acting on the columns X *,i of X, and is shared between all columns to project the time series from the input length to the prediction length, and the operation formula is:
[0099]
[0100] Wherein, And The weight and bias of the linear layer are respectively C' represents the number of time series, The mapping relationship between the input dimension And the output dimension ;
[0101] (63) The convolution kernel weight of the time convolution network TCN is updated as sp , K sp , and V sp , and the trend parameter ζ of ITSMixer 趋势 , the seasonal parameter ζ 季节性 , and the residual parameter ζ 残差 are put into the RSNPD algorithm for optimization:
[0102]
[0103] μ T , μ S , and μ R = RSNPD(X TCN , X Atten )
[0104] where θ contains the convolution kernel weights of TCN , and the attention matrix Q sp , K sp , and V sp , br is the weight decay coefficient, μ T , μ S , and μ R are dynamic weights generated by RSNPD, X TCN and X Attn represent feature vectors, respectively;
[0105] (64) Apply time mixing TM to all columns of X and perform time feature conversion, replace the original time mixing layer TM with a multi-scale TCN module to capture local to global time sequence features, the operation formula is:
[0106]
[0107]
[0108] where TM-TCN(X) j,i is the defined dilated causal convolution, kc is the size of the convolution kernel, d is the dilation rate, share the convolution kernel weights across channels, TM-TCN(X) *,i is the output after applying multi-scale time convolution to the i-th feature channel, is the residual output of time mixing, Drop() represents dropout, and Norm() represents batch normalization;
[0109] (65) Apply feature mixing FM to the rows of the input matrix , share across all rows, and apply to each row X j,* of the input matrix, then introduce cross-variable sparse attention mechanism in the feature mixing layer to enhance feature mixing, the operation formula is:
[0110] Q sp = XW Q ,K sp = XW K ,V sp = XW V
[0111]
[0112] U′ j,* = Drop(σ(Q2X j,* + Ab2))
[0113]
[0114] where Attn() is an attention mechanism, Softmax() is a classification function, σ is an activation function, Q2 and Q3 ∈ R C ′×C′ , Ab2 and Ab3 ∈ R C′ ,Q sp , K sp and V sp are attention calculation matrices, W Q , W K and W V are projection matrices, is the dimension of K sp , is a feature mixing with attention mechanism added;
[0115] When it is necessary to project the features to different sizes of H (H = C), ITSMixer performs linear transformation on the residual term:
[0116]
[0117] where Q3, Q H ∈ R H×C′ and Ab2, Ab3 ∈ R C′ ;
[0118] (66) Conditional Feature Mixing (CFM) input sequence in addition to considering the associated static features convert the static features into hidden features, and the operation formula is:
[0119]
[0120] where V′ j,* represents the matrix after the static features are associated, expand the input along the time dimension, and repeat the input times, is the concatenation of X and V′ along the feature dimension, represents from space C s to space H, CFM C′→H represents a conditional feature mixing layer that converts static features into hidden features;
[0121] (67) Apply Mix and CMix blocks for time and feature conversion respectively:
[0122]
[0123] where Mix C′→H represents a mixing operation in the mixing layer for converting input features in both time and feature dimensions, CMix C′→H represents a mixing operation in the conditional mixing layer for converting input features in both time and feature dimensions, CFR C′→H represents a feature reshaping operation in the conditional mixing layer;
[0124] (68) Obtain the final classification output Y out using a softmax layer;
[0125] Y out = Softmax (Q c [X TCN , X Attn ]+Ab c )
[0126] where Q c and Ab c are the weights and biases of the linear layer respectively, X TCN and X Attn are the feature vectors.
[0127] Based on the same inventive concept, the application also provides a bearing fault diagnosis system based on image feature fusion, comprising:
[0128] a data acquisition module for acquiring original vibration signal data of bearing faults and classifying them according to the types of faults and the degrees of damage;
[0129] a data processing module for optimizing SVMD parameters by using a reverse spiral neural population dynamics algorithm (RSNPD) to obtain optimized SVMD, decomposing the original vibration signal data by using the optimized SVMD to obtain modal component signals, collecting the decomposed modal component signals, randomly cutting three vibration signals from each modal component signal to generate different fault samples, converting the fault samples into RGB images to obtain time domain features of the original signals, and converting the original vibration signal data into two-dimensional image signals by using a continuous wavelet transform (CWT) to obtain frequency domain features of the original signals;
[0130] The bearing fault diagnosis module is used for mixing the RGB image and the two-dimensional image signal to obtain an image dataset, inputting image features of the image dataset into an improved time series mixer (ITSMixer) model for training, optimizing hyperparameters of the ITSMixer model by using an RSNPD algorithm, and obtaining a bearing fault diagnosis model; and the bearing fault diagnosis model after training and optimization is used for bearing fault diagnosis, and a corresponding diagnosis result is obtained.
[0131] Based on the same inventive concept, the application further provides a computing device, comprising one or more processors, one or more memories, and one or more programs stored in the memories and configured for execution by the processors, wherein the programs, when loaded into the processors, implement the steps of the bearing fault diagnosis method based on image feature fusion according to any one of the above.
[0132] Based on the same inventive concept, the application further provides a storage medium, wherein the storage medium stores a computer program, and the computer program comprises program instructions, and the program instructions, when executed by a processor, cause the processor to perform the steps of the bearing fault diagnosis method based on image feature fusion according to any one of the above.
[0133] Beneficial effects: compared with the prior art, the present application enhances the data decomposition effect by using the reverse spiral neural population dynamics algorithm (RSNPD) to optimize the successive variation modal decomposition (SVMD), the parameters of the successive variation modal decomposition are optimized by introducing the reverse spiral neural population dynamics algorithm (RSNPD), so that the inherent mode in the data can be separated more accurately, thereby significantly improving the precision of the bearing fault diagnosis result; the present application proposes an improved time series mixer (ITSMixer) model by combining a time series mixer (TSMixer) model with a time convolution network (TCN) and a sparse attention mechanism, solves the conflict problem between the insufficient local feature focusing and the global dependence calculation redundancy in the traditional time sequence model. By replacing the original time mixing layer (TM) with a multi-scale TCN module, the local time sequence features can be better captured, by introducing a cross-variable sparse attention mechanism in the feature mixing layer (FM), the key time steps can be adaptively selected in the local window, after the double-path feature fusion, the precise modal separation of the complex time sequence signal is realized, and the accuracy of the fault diagnosis is effectively improved; the present application establishes the ITSMixer model to diagnose the bearing fault, converts the decomposed multi-channel signal into an RGB image to obtain the time domain feature information; the one-dimensional original vibration signal is converted into a two-dimensional image signal by using continuous wavelet transform (CWT) to obtain the frequency domain features of the original signal; then the image features are fused and sent into the ITSMixer model, and the ITSMixer model is optimized by using the RSNPD, to obtain the final bearing fault diagnosis result, which significantly improves the accuracy of the model bearing fault diagnosis. BRIEF DESCRIPTION OF DRAWINGS
[0134] Figure 1 The method flowchart of the embodiment of the present application is shown in the figure.
[0135] Figure 2 The RSNPD algorithm flowchart of the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0136] In order to enable the personnel in the technical field to better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application.
[0137] As shown in the figure, the bearing fault diagnosis method based on image feature fusion of the present embodiment comprises: Figure 1
[0138] (1) obtaining the original vibration signal data of the bearing fault, and classifying according to the generated fault type and damage degree;
[0139] (2) optimizing the parameters of the successive variation modal decomposition (SVMD) by using the reverse spiral neural population dynamics algorithm (RSNPD), to obtain the optimized successive variation modal decomposition (SVMD);
[0140] (3) The original vibration signal data is decomposed by using an optimized successive variation modal decomposition (SVMD) to obtain modal component signals;
[0141] (4) The decomposed modal component signals are collected, three vibration signals are randomly intercepted from each modal component signal to generate different fault samples, the fault samples are converted into RGB images to obtain time domain features of the original signals;
[0142] (5) The original vibration signal data is converted into a two-dimensional image signal by using a continuous wavelet transform (CWT) to obtain frequency domain features of the original signals;
[0143] (6) The RGB images and the two-dimensional image signals are mixed to obtain an image dataset, and image features of the image dataset are input into an improved time series mixer (ITSMixer) model for training, the hyperparameters of the ITSMixer model are optimized by using an RSNPD algorithm to obtain a bearing fault diagnosis model;
[0144] (7) The bearing fault diagnosis model after training and optimization is used for bearing fault diagnosis to obtain corresponding diagnosis results.
[0145] Specifically, in step (1), the original vibration signals of bearing faults are obtained, including inner ring fault vibration signals, outer ring fault vibration signals, rolling element fault vibration signals, and fault vibration signals caused by wear, and corresponding labels are made according to the types and degrees of damage;
[0146] In step (2), the reverse spiral neural population dynamic algorithm (RSNPD) includes random initialization of population positions by using a lens imaging reverse learning strategy, introduction of a spiral sine strategy to enhance disturbance of position update of the optimal solution, and increase of diversity of the population, which is helpful to jump out of local optimum, and specifically includes the following steps:
[0147] (21) The objective function of the RSNPD algorithm is set as envelope entropy, and related parameters are initialized, including population size, iteration number, dimension size, and penalty parameter a;
[0148] (22) As shown in Figure 2 , the random initialization of population positions is performed by using a lens imaging reverse learning strategy, and this process is represented by a mathematical formula as follows:
[0149]
[0150] Wherein, lb is the lower limit, ub is the upper limit, a i,j is the position of the i-th neural population in the j-th dimension, h is the length of the incident light, and h* is the length of the refracted light;
[0151] (23) The fitness value is calculated according to the envelope entropy, and the optimal solution is obtained according to the calculated fitness value;
[0152] fitness = min [E p (i)]
[0153]
[0154] wherein E p is the envelope entropy of each IMF, is the envelope signal of each mode after Hilbert demodulation, p j is the normalized form of , and K is the number of IMFs;
[0155] (24) The RSNPD algorithm is divided into three strategies in the optimization process, namely the attractor trend strategy, the coupling interference strategy and the information projection strategy, and the specific process is as follows:
[0156] (241) The neural states of the neural population tend to be multi-attractors in the neural state space. Therefore, the algorithm takes the optimal neural states as attractors, and the neural states of the remaining neural population randomly converge to one of these attractors. The number of attractors is set to [c·nd], wherein c is the proportion of attractors in all neural states in the neural population, and nd is the number of the neural population. This process is the attractor trend strategy of the NPDOA algorithm, and this process is mathematically expressed as:
[0157]
[0158] wherein l is the proportion factor of the attractor trend, r1 is a random number in [0,1], att g is the gth attractor, g is a random number in {1,2,…,[c·nd]}, a i is the neural state of the ith neural population, W i is the zero-mean Gaussian noise vector of the ith neural state, and the covariance is U-L, wherein U is the peak value of all neural states, and L is the minimum boundary of all neural states;
[0159] (242) The neural state of a specific neural population is also affected by the neural states of other neural populations. It is usually assumed that the interaction between neural populations occurs through additive or diffusive coupling. The algorithm considers the existence of these two couplings in the dynamics of the neural population. In additive coupling, the input of the neural population is a function of the sum of the neural states of other neural populations. The algorithm represents the neural state of other neural populations as the average value a add,i of the sum of all neural states, and the formula is as follows:
[0160]
[0161] where r2 is a random number in [0, 0.5], and nd is the number of neural populations;
[0162] In the diffusion coupling, the input of a neural population is the sum of the difference between its neural state and the neural states of other neural populations a dif,i , which is calculated as follows:
[0163]
[0164] where r3 is a random number in [0, 0.5], and a i is the neural state of the i-th neural population;
[0165] The combination of two couplings is used to simulate the influence of other neural populations, which can be expressed as:
[0166] a couple,i = d · (a add,i + a dif,i )
[0167] where d is the scale factor of the coupling disturbance;
[0168] (243) A spiral sine strategy is introduced to perturb the role stage, which helps to jump out of the local optimum. This process is represented by the mathematical formula as follows:
[0169]
[0170] The new coupling disturbance strategy calculation formula is:
[0171]
[0172] where η b+1 is the spiral sine calculation formula, b is the iteration number, g is the spiral coefficient, which is in the range of [-1, 1], and ψ1 is a random number.
[0173] (244) The exchange and interaction of information between neural populations is carried out through the communication subspace. In the algorithm, the communication subspace is represented by a matrix representing the adjacency degree and communication strength, which represents the information projection between neural populations.
[0174] In the attraction trend, the information transmission between the attractor and the other neural state of the neural population is realized by a 1 × D adjacency matrix C attract,i and a 1 × D communication strength matrix R attract,i . The formula used is as follows:
[0175] a C_attract,i = C attract,i · R attract,i · a attract,i
[0176] Among them C attract,i is a 1×D matrix with binary random values, R attract,i is a 1×D matrix with randomly generated values in [0,1], a attract,i is the dynamic vector obtained from the attractor trend strategy;
[0177] In coupled perturbations, information transfer between neural populations is also controlled by another 1×D adjacency matrix C couple,i , and the 1×D communication intensity matrix R couple,i Furthermore, disturbances in the neural state should have a greater impact early in the process, before the neural state stabilizes. Therefore, the influence of coupled perturbations gradually decreases, using the following formula:
[0178] a C_coupie,i =C coupie,i ·R coupie,i ·a coupie,i ·r4·(1-FE / FE max )
[0179] Among them C couple,i is a 1×D matrix with binary random values, R couple,i is a 1×D matrix with randomly generated values in [0,1], a couple,i is the dynamic vector obtained by the coupled perturbation strategy, r4 is a random number in [0,1], FE is the function evaluation number, FE max is the maximum number of function evaluations. As the iteration process proceeds, FE gradually increases, while the influence of the coupled perturbation gradually decreases.
[0180] Ultimately, the neural state of each neural population is affected by the attractor and the neural states of other neural populations. The neural state of the current neural population a new,i is updated as follows:
[0181]
[0182] (25) The current solution is continuously updated through the RSNPD optimization process, and the global optimal solution is updated synchronously;
[0183] (26) Determine whether the algorithm termination condition is reached by the given maximum number of iterations. If not, go to step (3) and output the optimal result within the specified number of iterations.
[0184] In step (3), the optimized successive variational mode decomposition (SVMD) is used to decompose the data into relatively stable modal components. The parameters of SVMD are optimized by the improved neural population dynamic optimization algorithm (NPDOA), which specifically includes the following steps:
[0185] (31) Suppose the original vibration signal x(t) is decomposed into two signals, including K modes u K (t) and residual signal x r (t), the formula is as follows:
[0186] x(t) = u K (t) + x r (t)
[0187] Where t is the time step, the residual signal x r (t) includes the original signal x(t) unprocessed part x u (t) and the obtained modal part, the formula is as follows:
[0188]
[0189] (32) Each mode should be compactly surrounded by its center frequency, therefore, the Kth mode minimization should meet the following criteria:
[0190]
[0191] Where, is the quadratic penalty factor, used to reduce the interference of noise, δ(t) is the Dirac distribution function, ω K is the center frequency of the Kth mode, g is the rotation factor, * is the convolution operation, u K (t) is the Kth mode subsequence.
[0192] (33) The spectral overlap between u k (t) and x r (t) should be minimized, ensuring their energy within the target modal band is minimal, to ensure stable implementation of this constraint, a suitable filter is adopted, whose frequency response is as follows:
[0193]
[0194] Where α is the penalty parameter, ω is the frequency, then the established constraint is:
[0195]
[0196] Where β K (t) is the impulse response in the filter, ‖‖ represents the norm.
[0197] (34) The energy of u k (t) within ±5% of the center frequency of the previously obtained mode should be minimized, to ensure stable implementation of this constraint, a specific filter is adopted, whose frequency response is as follows:
[0198]
[0199] The established constraint is:
[0200]
[0201] where β k (t) is the impulse response in the filter.
[0202] (35) The original signal x(t) should be completely reconstructed from K modes and x u (t). Therefore, when K-1 modes are known, the extraction task of the Kth mode can be converted into a constrained minimization problem, expressed as follows:
[0203]
[0204] where α is a parameter that balances J1, J2 and J3.
[0205] (36) In order to better converge and maximize the reconstruction of the original signal in the presence of noise, an augmented Lagrange function is established by constructing a quadratic penalty term and a Lagrange multiplier, expressed as follows:
[0206]
[0207] where λ is the Lagrange multiplier coefficient, and the optimal solution to the above problem is solved by introducing an augmented Lagrange, and then the alternating direction multiplier method is used to continuously iterate to finally obtain the decomposed K components.
[0208] In step (4), the decomposed modal component signals are collected, and each modal component signal is randomly cut into three vibration signals to generate different fault samples, which are converted into RGB images to obtain the time domain features of the original signal, including the following steps:
[0209] (41) The modal component signal signal is expressed as:
[0210] B r ={b i1j1 ,b ij ∈R n×m}, i1 = 1, 2, …, n, j1 = 1, 2, …, m
[0211] where B r represents an n-row and m-column matrix, n represents the number of channels and takes a value of 3, and m represents the number of sampling points;
[0212] (42) Randomly cut three vibration signals to generate S tDifferent fault samples of D'x D'x 3, D' is the sample length, each sample is randomly sampled from [1, m-D'x D'] without overlapping; then according to the sampling points, a sample length of kx k is randomly generated:
[0213]
[0214] Where PM i (α,β) is the pixel matrix of each channel, α=1,…,D'; β=1,…,D'; z=1,2,3; let N represent the total number of samples, represents the signal value, gt=1,2,…,N; h=1,2,…,D ′2 , after normalization, the value range of each pixel matrix is from 0 to 255.
[0215] (43) fuse the three D'x D' matrices into the RGB pixel matrix to obtain the RGB image; the RGB pixel matrix is RGB img (α,β,c), c=1,2,3, which is obtained from RGB img (α,β,c)=(PM z (α,β), c) is obtained.
[0216] In step (5), the one-dimensional original vibration signal is converted into a two-dimensional image signal by using continuous wavelet transform (CWT) to obtain the frequency domain features of the original signal, which specifically includes the following steps:
[0217] (51) Continuous wavelet transform (CWT) can display the time-frequency characteristics of the original vibration signal on the image by using a window that changes with frequency. Let γ(t) be a mother wavelet, and the continuous wavelet function γ m′,n′ (t) is obtained by stretching and translating the wavelet function, and the expression is:
[0218]
[0219] Where m' is the scaling factor, n' is the translation factor, t is the time, γ is the mother wavelet function, and R is a real number;
[0220] (52) Given a square-integrable signal f(t), the continuous wavelet transform expression of f(t) is:
[0221]
[0222] Where γ' is the derivative of the continuous wavelet function;
[0223] (53) When C f (m',n') is inversely transformed, we can get:
[0224]
[0225] in, w is the frequency of wavelet transform. When the value of m′ is greater than the inverse of the signal sampling frequency, it is suitable for extracting low-frequency features in the signal. Otherwise, it is suitable for extracting high-frequency features in the signal.
[0226] In step (6), the fused image features are input into the ITSMixer model for training, and the hyperparameters of the ITSMixer model are optimized using the RSNPD algorithm to obtain a bearing fault diagnosis model, which specifically includes the following steps:
[0227] (61) Mixing the RGB image and the two-dimensional image signal to obtain an image dataset, and inputting the images in the dataset into the ITSMixer model;
[0228] (62) Given an input matrix Time projection (TP) is a column acting on X (denoted as X *,i ) and shared across all columns to project the time series from the input length to the prediction length. The operation is defined as:
[0229]
[0230] in, and are the weights and biases of the linear layer, C′ represents the number of time series, Represents the input dimension and output dimension The mapping relationship between them;
[0231] (63) The convolution kernel weight of the temporal convolutional network TCN Attention matrix Q sp , K sp and V sp , and ITSMixer's trend parameter ζ 趋势 , seasonal parameter ζ 季节性 and the residual parameter ζ 残差 Put it into the RSNPD algorithm for optimization and then proceed to the subsequent steps;
[0232]
[0233] μ T ,μ S ,μ R =RSNPD(X TCN ,X Atten )
[0234] Among them, θ contains the core weight of TCN Attention matrix Q sp , Ksp , V sp , br is the weight decay coefficient, μ T , μ S , μ R Dynamic weights generated by RSNPD, X TCN and X Attn are feature vectors, respectively.
[0235] (4) Temporal mixing (TM) acts on all columns of X and performs temporal feature transformation, replacing the original temporal mixing layer (TM) with a multi-scale TCN module to capture local to global temporal features, and the operation is defined as:
[0236]
[0237] where TM-TCN(X) j,i is the defined dilated causal convolution, kc is the size of the convolution kernel, and d is the dilation rate, channel-shared convolution kernel weights, TM-TCN(X) *,i is the output after applying multi-scale temporal convolution to the i-th feature channel, is the residual output of temporal mixing, Drop(·) is dropout, and Norm(·) is batch normalization.
[0238] (65) Feature mixing (FM) is a two-layer residual MLP that acts on the rows of the input matrix and is shared between all rows and applied to each row X j,* of the input matrix, and then introduces a cross-variable sparse attention mechanism in the feature mixing layer to enhance feature mixing, and the operation is defined as:
[0239] Q sp = XW Q , K sp = XW K , V sp = XW V
[0240]
[0241] U′ j,* = Drop(σ(Q2X j,* + Ab2))
[0242]
[0243] where Attn() is the attention mechanism, Softmax() is the classification function, σ is the activation function, Q2 and Q3 ∈ R C ′×C′ , Ab2 and Ab3 ∈ R C′ , Qsp , K sp and V sp are attention computation matrices, W Q , W K and W V are projection matrices, is the dimension of K sp , is the feature mixing with attention mechanism added;
[0244] When the features need to be projected to different sizes of H (H = C), ITSMixer performs linear transformation on the residual term:
[0245]
[0246]
[0247] where Q3, Q H ∈ R H×C′ and Ab2, Ab3 ∈ R C′
[0248] (66) Conditional Feature Mixing (CFM) is a variant of FM block, in addition to the input sequence , the associated static features are also considered. The operation of converting static features to hidden features is defined as:
[0249]
[0250] where V′ j,* represents the matrix after associating the static features, extends the input along the time dimension, repeating the input times, is the concatenation of X and V′ along the feature dimension, represents the projection from the space C s ′ to the space H, CFM C′→H represents the operation of converting static features to hidden features in the conditional feature mixing layer;
[0251] (67) Mixer layer (Mix) is the combination of time mixing layer and feature mixing layer, while conditional mixer layer (CMix) is the combination of time mixing layer and conditional feature mixing layer, Mix and CMix blocks apply time and feature transformations respectively:
[0252]
[0253] where Mix C′→H represents the mixing operation in the mixing layer, which is used to transform the input features in both time and feature dimensions, CMix C′→Hdenotes a mixing operation in the conditional mixing layer for transforming the input features simultaneously in time and feature dimensions, CFR C′→H denotes a feature reshaping operation in the conditional mixing layer;
[0254] (68) using a softmax layer to obtain the final classification output result;
[0255] Y out = Softmax (Q c [X TCN , X Attn ]+Ab c )
[0256] wherein Q c and Ab c are weights and biases of the linear layer respectively, X TCN and X Attn are feature vectors respectively;
[0257] In step (7), the bearing fault diagnosis model optimized by training is used to diagnose the bearing fault, and the corresponding diagnosis result is obtained;
[0258] Further, it further comprises the step (8) of displaying the bearing fault diagnosis result obtained in the step (7) on the front-end interface, so as to facilitate the supervision personnel to handle;
[0259] Based on the same inventive concept, the embodiment also provides a bearing fault diagnosis system based on image feature fusion, comprising:
[0260] A data acquisition module is configured to acquire original vibration signal data of bearing faults and classify the original vibration signal data according to fault types and damage degrees;
[0261] The data processing module is configured to optimize SVMD parameters by using a reverse spiral neural population dynamics algorithm (RSNPD) to obtain optimized SVMD, decompose the original vibration signal data by using the optimized SVMD to obtain modal component signals, collect the decomposed modal component signals, randomly cut three vibration signals from each modal component signal to generate different fault samples, convert the fault samples into RGB images to obtain time domain features of the original signal, and convert the original vibration signal data into a two-dimensional image signal by using a continuous wavelet transform (CWT) to obtain frequency domain features of the original signal.
[0262] The bearing fault diagnosis module is used for mixing the RGB image and the two-dimensional image signal to obtain an image dataset, inputting image features of the image dataset into an improved time series mixer (ITSMixer) model for training, optimizing hyperparameters of the ITSMixer model by using an RSNPD algorithm, and obtaining a bearing fault diagnosis model; and the bearing fault diagnosis model after training and optimization is used for bearing fault diagnosis, and a corresponding diagnosis result is obtained.
[0263] Based on the same inventive concept, the embodiment further provides a computing device, comprising one or more processors, one or more memories, and one or more programs stored in the memories and configured to be executed by the processors, and the programs, when loaded into the processors, implement the steps of the bearing fault diagnosis method based on image feature fusion according to any one of the above.
[0264] Based on the same inventive concept, the embodiment further provides a storage medium, characterized in that the storage medium stores a computer program, and the computer program comprises program instructions, and the program instructions, when executed by a processor, cause the processor to execute the steps of the bearing fault diagnosis method based on image feature fusion according to any one of the above.
[0265] The above specific implementation cases are used to explain and illustrate the application, and are only preferred embodiments of the application, but not limit the application, and any modification, equivalent replacement, improvement, etc. of the application within the spirit and protection scope of the claims of the application, all fall into the protection scope of the application.
Claims
1. A bearing fault diagnosis method based on image feature fusion, characterized in that: include: (1) Obtain the original vibration signal data of the bearing fault and classify it according to the fault type and damage degree; (2) The parameters of successive variational mode decomposition (SVMD) are optimized by using the reverse spiral neural population dynamics algorithm (RSNPD) to obtain the optimized successive variational mode decomposition (SVMD); (3) Using optimized successive variational modal decomposition (SVMD) to decompose the original vibration signal data to obtain modal component signals; (4) Collecting the decomposed modal component signals, randomly intercepting three segments of vibration signals from each modal component signal to generate different fault samples, converting the fault samples into RGB images, and obtaining the time domain features of the original signal; (5) Using continuous wavelet transform (CWT) to transform the original vibration signal data into a two-dimensional image signal to obtain the frequency domain characteristics of the original signal; (6) The RGB image and the two-dimensional image signal are mixed to obtain an image dataset, and the image features of the image dataset are input into the improved time series mixer ITSMixer model for training. The hyperparameters of the ITSMixer model are optimized using the RSNPD algorithm to obtain a bearing fault diagnosis model. (7) Use the trained and optimized bearing fault diagnosis model to diagnose bearing faults and obtain the corresponding diagnosis results.
2. The bearing fault diagnosis method based on image feature fusion according to claim 1 is characterized in that: The step (2) specifically includes randomly initializing the population position by using the reverse spiral neural population dynamics algorithm RSNPD, adopting the lens imaging reverse learning strategy, introducing the spiral sine strategy to enhance the disturbance of the optimal solution position update, and increasing the diversity of the population, including: (21) Setting the objective function of the RSNPD algorithm to envelope entropy and initializing parameters, the parameters include: population size, number of iterations, dimension size, and penalty parameter α; (22) The population position is randomly initialized using the lens imaging reverse learning strategy, and the formula is: Among them, lb is the lower limit, ub is the upper limit, a i,j is the position of the i-th neural population in the j-th dimension, h is the length of the incident light, and h* is the length of the refracted light; (23) Calculate the fitness value based on the envelope entropy, and obtain the optimal solution based on the calculated fitness: fitness=my[E p (in)] Among them, E p is the envelope entropy of each IMF, á(j) is the envelope signal of each mode after Hilbert demodulation, p j is the normalized form of á(j), K is the number of IMFs; (24) RSNPD algorithm realizes the optimization process through attractor trend strategy, coupled interference strategy and information projection strategy; (25) The current solution is continuously updated through the RSNPD optimization process, and the global optimal solution is updated synchronously; (26) Determine whether the given maximum number of iterations has been reached. If not, go to step (23) and output the optimal result within the specified number of iterations.
3. The bearing fault diagnosis method based on image feature fusion according to claim 2 is characterized in that: The step (24) comprises: (241) In the attractor trend strategy, the optimal neural state is used as the attractor, and the neural states of the remaining neural populations randomly converge to one of the attractors. The number of attractors is set to [c·nd], where c is the proportion of attractors in all neural states in the neural population, and nd is the number of neural populations. The formula is: Where l is the scaling factor of the attractor trend, r1 is a random number in [0,1], and att g is the g-th attractor, g is a random number in {1,2,…,[c·nd]}, a i is the neural state of the ith neural population, W i is the zero-mean Gaussian noise vector of the i-th neural state with covariance UL, U is the peak value of all neural states, and L is the minimum boundary of all neural states; (242) Assume that the interaction between neural populations occurs through additive or diffusive coupling, in which the input of a neural population is a function of the sum of the neural states of all neural populations except the optimal neural population, and the neural states of other neural populations are represented as the average value a of the sum of all neural states. add,i , the formula is as follows: Where r2 is a random number in [0,0.5], and nd is the number of neural populations; In diffuse coupling, the input to a neural population is the sum of the differences between its neural state and the neural states of other neural populations. dif,i , the formula is as follows: Among them, r3 is a random number in [0,0.5], a i is the neural state of the i-th neural population; The influence of other neural groups is simulated using a combination of two couplings, as follows: a couple,i =d·(a add,i +a dif,i ) Where d is the proportional factor of the coupled disturbance; (243) The spiral sine strategy is introduced to perturb the role stage, and the formula is: The new coupling interference strategy calculation formula is: Among them, η b+1 is the spiral sine calculation formula, b is the number of iterations, g is the spiral coefficient, the value range is [-1,1], ψ1 is a random number; (244) In the attractor trend, the information transmission between the attractor and another neural state of the neural population is represented by the 1×D adjacency matrix C attract,i and the 1×D communication intensity matrix R attract,i The implementation formula is as follows: a C_attract,i =C attract,i ·R attract,i ·a attract,i Among them C attract,i is a 1×D matrix with binary random values, R attract,i is a 1×D matrix with randomly generated values in [0,1], a attract,i is the dynamic vector obtained from the attractor trend strategy; In coupled perturbations, information transfer between neural populations is replaced by another 1×D adjacency matrix C couple,i , and the 1×D communication intensity matrix R couple,i The formula used is as follows: a C_coupie,i =C coupie,i ·R coupie,i ·a coupie,i ·r4·(1-FE / FE max ()) Among them C couple,i is a 1×D matrix with binary random values, R couple,i is a 1×D matrix with randomly generated values in [0,1], a couple,i is the dynamic vector obtained by the coupled perturbation strategy, r4 is a random number in [0,1], FE is the function evaluation number, FE max is the maximum number of function evaluations. As the iteration process proceeds, FE gradually increases; The neural state of each neural population is affected by the attractor and the neural state of other neural populations. The neural state of the current neural population is a new,i Updates as follows: a new,i =a i +a C_attract,i +a C_coupie,i 。 4. The bearing fault diagnosis method based on image feature fusion according to claim 1 is characterized in that: The step (3) comprises: (31) Assume that the original vibration signal x(t) is decomposed into two signals, including K modes u K (t) and the residual signal x r (t), the formula is as follows: x(t)=u K (t)+x r (t) Where t is the time step, the residual signal x r (t) includes the unprocessed part x of the original signal x(t) u (t) and obtain the modal part, the formula is as follows: (32) Let the K-th order mode be minimized to meet the following criteria: in, is a quadratic penalty factor used to reduce noise interference, δ(t) is the Dirac distribution function, ω K is the center frequency of the K-order mode, g is the rotation factor, * is the convolution operation, u K (t) is the subsequence of the Kth mode; (33) Set the constraint so that u k (t) and x r The spectral overlap between (t) is minimized by using a filter with the following frequency response: Among them, α is the penalty parameter and ω is the frequency, then the established constraint is: Among them, β K (t) is the impulse response in the filter, ‖‖ indicates the norm; (34) Set the constraint so that u k (t) Minimize the energy within ±5% of the modal center frequency using a filter with the following frequency response: The constraints established are: Among them, β k (t) is the impulse response in the filter; (35) The task of extracting the Kth modality is transformed into a constrained minimization problem, which is expressed as follows: Among them, α is the parameter that balances J1, J2 and J3; (36) Construct the quadratic penalty term and Lagrangian multiplier, establish the augmented Lagrangian function, and reconstruct the original signal. The expression is as follows: Among them, λ is the Lagrange multiplier coefficient. After solving the optimal solution, the alternating direction multiplier method is used to iterate continuously. Get the decomposed K components.
5. The bearing fault diagnosis method based on image feature fusion according to claim 1, characterized in that: The step (4) comprises: (41) The modal component signal is expressed as: B r ={b i1j1 ,b ij ∈R n×m },i1=1,2,…,n,j1=1,2,…,m Among them, B r Represents a matrix of n rows and m columns, n represents the number of channels, which is 3, and m represents the number of sampling points; (42) Randomly intercept three vibration signals and generate a signal of size S t = D′×D′×3 different fault samples, where D′ is the sample length. Each sample is randomly sampled from [1,mD′×D′] without overlapping. Then, based on the sampling point, the randomly generated sample length is k×k: Among them, PM i (α, β) is the pixel matrix of each channel, α=1,…,D′; β=1,…,D′; z=1,2,3; let N represent the total number of samples, Indicates the signal value, gt=1,2,…,N; h=1,2,…,D ′2 ,After normalization, each pixel matrix value ranges from 0 to 255; (43) The three segments of D′×D′ matrix are fused into the RGB pixel matrix to obtain an RGB image; the RGB pixel matrix RGB img (α,β,c),c=1,2,3,by RGB img (α,β,c)=(PM z (α,β),c) is obtained.
6. The bearing fault diagnosis method based on image feature fusion according to claim 1 is characterized in that: The step (5) comprises: (51) Let γ(t) be a mother wavelet, and the continuous wavelet function γ is obtained by stretching and translating the wavelet function. m′,n′ (t): Among them, m′ is the scaling factor, n′ is the translation factor, t is the time, γ is the mother wavelet function, and R is a real number; (52) Given a square integrable signal f(t), the continuous wavelet transform expression of f(t) is: Where γ′ is the derivative of the continuous wavelet function; (53) When C f When (m′,n′) is inversely transformed, we get: in, w is the frequency of wavelet transform. When the value of m′ is greater than the inverse of the signal sampling frequency, the low-frequency features in the signal are extracted. Otherwise, the high-frequency features in the signal are extracted.
7. The bearing fault diagnosis method based on image feature fusion according to claim 1 is characterized in that: The step (6) comprises: (61) Mixing the RGB image and the two-dimensional image signal to obtain an image dataset, and inputting the images in the dataset into the ITSMixer model; (62) Given an input matrix The time projection TP is the column X acting on X *,i A linear layer is created and shared among all columns to project the time series from the input length to the prediction length. The operation formula is: in, and are the weights and biases of the linear layer, C′ represents the number of time series, Represents the input dimension and output dimension The mapping relationship between them; (63) The convolution kernel weight of the temporal convolutional network TCN Attention matrix Q sp , K sp and V sp , and ITSMixer's trend parameter ζ 趋势 , seasonal parameter ζ 季节性 and the residual parameter ζ 残差 Put it into RSNPD algorithm for optimization: m T ,m S ,m R =RSNPD(X TCN ,X Atten ) Among them, θ contains the convolution kernel weights of TCN And the attention matrix Q sp , K sp and V sp , br is the weight attenuation coefficient, μ T 、μ S and μ R is the dynamic weight generated by RSNPD, X TCN and X Attn They represent eigenvectors respectively; (64) Apply the time mixing TM to all columns of X and perform time feature conversion. Replace the original time mixing layer TM with a multi-scale TCN module to capture local to global temporal features. The operation formula is: Among them, TM-TCN(X) j,i is the defined dilated causal convolution, kc is the size of the convolution kernel, d is the dilation rate, Channel-shared convolution kernel weights, TM-TCN(X) *,i is the output after applying multi-scale temporal convolution to the i-th feature channel, It is the residual output of time mixing, Drop() represents dropout, and Norm() represents batch normalization; (65) Apply the feature mixture FM to the input matrix The rows of , shared between all rows, and applied to each row of the input matrix X j,* , then introduce the cross-variable sparse attention mechanism in the feature mixing layer to enhance feature mixing. The operation formula is: Q sp =XW Q ,K sp =XW K ,V sp =XW V U′ j,* =Drop(σ(Q2X j,* +Ab2)) Among them, Attn() is the attention mechanism, Softmax() is the classification function, σ is the activation function, Q2 and Q3∈R C′×C′ , Ab2 and Ab3∈R C′ ,Q sp , K sp and V sp is the attention calculation matrix, W Q , W K and W V is the projection matrix, K sp Dimensions, Feature blending to add attention mechanism; When it is necessary to project features to H of different sizes (H = C), ITSMixer performs a linear transformation on the residual term: Among them, Q3,Q H ∈R H×C′ and Ab2,Ab3∈R C′ ; (66) Input the conditional feature mixture CFM into the sequence In addition, consider the static characteristics of the association Convert static features into hidden features. The operation formula is: Among them, V′ j,* Represents the matrix after associating static features, Expand the input along the time dimension and repeat the input Second-rate, is the concatenation of X and V′ along the characteristic dimension, Indicates that from space C s ′ to space H, CFM C′→H Indicates that static features are converted into hidden features in the conditional feature mixing layer; (67) Apply the Mix and CMix blocks to the time and feature transformations, respectively: Among them, Mix C′→H Represents the mixing operation in the mixing layer, which is used to transform the input features in both time and feature dimensions. CMix C′→H Represents the mixing operation in the conditional mixing layer, which is used to transform the input features in both time and feature dimensions. C′→H Represents the feature reshaping operation in the conditional mixture layer; (68) Use the softmax layer to get the final classification output result Y out ; Y out =Softmax(Q c [X TCN ,X Attn ]+Ab c ) Among them, Q c and Ab c are the weights and biases of the linear layer, X TCN and X Attn are the eigenvectors respectively.
8. A bearing fault diagnosis system based on image feature fusion, characterized in that: include: The data acquisition module is used to obtain the original vibration signal data of the bearing fault and classify it according to the fault type and damage degree; A data processing module is configured to optimize the parameters of a successive variational modal decomposition (SVMD) using a reverse spiral neural population dynamics (RSNPD) algorithm to obtain an optimized successive variational modal decomposition (SVMD); decompose the original vibration signal data using the optimized successive variational modal decomposition (SVMD) to obtain modal component signals; collect the decomposed modal component signals, randomly intercept three segments of vibration signals from each modal component signal to generate different fault samples, convert the fault samples into RGB images, and obtain time domain features of the original signal; Continuous wavelet transform (CWT) is used to transform the original vibration signal data into a two-dimensional image signal to obtain the frequency domain characteristics of the original signal; The bearing fault diagnosis module is used to mix RGB images and two-dimensional image signals to obtain an image dataset, and input the image features of the image dataset into the improved time series mixer ITSMixer model for training. The hyperparameters of the ITSMixer model are optimized using the RSNPD algorithm to obtain a bearing fault diagnosis model. The bearing fault diagnosis model after training and optimization is used to diagnose bearing faults and obtain the corresponding diagnosis results.
9. A computing device, characterized in that include: One or more processors, one or more memories, and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the bearing fault diagnosis method based on image feature fusion according to any one of claims 1 to 7 are implemented.
10. A storage medium, characterized in that: The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor executes the steps of the bearing fault diagnosis method based on image feature fusion according to any one of claims 1 to 7.