Ship bearing fault diagnosis method under unbalanced data
By using signal processing and Gram image encoding technology in ship bearing fault diagnosis, forming a fault image sample set and establishing an optimized training bearing fault diagnosis model, the fault diagnosis challenge under unbalanced data is solved, and fault identification with high accuracy is achieved.
Patent Information
- Application Number
- CN202510102604.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately, efficiently and targetedly diagnose ship bearing faults under the condition of unbalanced sample data, especially for faults with low frequency and few samples.
By obtaining the vibration signal of normal generator bearings, the fault signal is converted into a Gram matrix using signal processing technology and Gram image encoding technology to form a bearing fault image sample set. Then, a bearing fault diagnosis model including a convolutional layer, a deep bearing fault extraction module and an output layer is established, and optimized training is performed through a gradient coordination class balance loss function to handle unbalanced data.
It realizes the high accuracy of the ship bearing fault diagnosis model under unbalanced sample conditions, and can deeply dig and extract the discrimination information of a small number of sample-type faults, solving the problem that it is difficult to diagnose small sample-type faults.
Smart Images

Figure CN119942218A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of bearing fault diagnosis, and in particular to a ship bearing fault diagnosis method under unbalanced data. Background Art
[0002] Bearings are the most common components in marine machinery and equipment. They work in harsh and complex environments for a long time, and the probability of failure is high. Failure to detect bearing failures in a timely and accurate manner may lead to significant economic losses or even casualties. Therefore, in order to ensure the normal operation of marine machinery and equipment, it is necessary to conduct fault diagnosis research on bearings, a common component.
[0003] However, the conditions for safe operation at sea limit the collection and marking of fault data. Compared with the amount of data under healthy operating conditions, bearing faults rarely occur and are difficult to mark in a timely manner, resulting in limited fault-related data and an imbalance in the number of fault samples and healthy samples. This causes model bias in data-driven fault diagnosis methods, which can easily lead to the model being able to identify the healthy state of the bearing well, but having difficulty in accurately identifying various bearing fault states, especially faults with low occurrence frequency and small number of samples, which poses a challenge to ship bearing fault diagnosis methods. Summary of the invention
[0004] The present invention provides a ship bearing fault diagnosis method under unbalanced data, so as to overcome the technical problem that the existing bearing fault diagnosis method is difficult to accurately, efficiently and targetedly perform fault diagnosis under the condition of unbalanced sample data.
[0005] In order to achieve the above object, the technical solution of the present invention is:
[0006] A ship bearing fault diagnosis method under unbalanced data, the specific steps include:
[0007] S1: obtaining a vibration signal of a normal generator bearing, and processing the vibration signal of the normal generator bearing by using a signal processing technology to obtain a plurality of fault signals;
[0008] S2: using Gram image coding technology to convert the multiple fault signals into corresponding Gram matrices, thereby forming a bearing fault image sample set;
[0009] S3: Establish bearing fault diagnosis model;
[0010] S4: dividing the bearing fault image sample set into a training set and a test set, wherein the training set is set in a plurality of unbalanced ways, and inputting each training set into the bearing fault diagnosis model for optimization training, and when the loss function of the bearing fault diagnosis model converges, obtaining the optimized and trained bearing fault diagnosis model;
[0011] S5: using the test set to verify the ship bearing fault classification effect of the optimized and trained bearing fault diagnosis model under the unbalanced sample condition, and obtain the verified bearing fault diagnosis model;
[0012] S6: Classify ship bearing faults under unbalanced data based on the verified bearing fault diagnosis model.
[0013] Further, the bearing fault diagnosis model comprises a first convolutional layer, a deep bearing fault extraction module and an output layer connected in sequence;
[0014] The first convolutional layer is used to preliminarily extract bearing fault features in the bearing fault image, and transmit the obtained primary bearing fault features to the deep bearing fault extraction module;
[0015] The deep bearing fault extraction module is used to extract high-dimensional mapping features from the primary bearing fault features and transmit them to the output layer;
[0016] The output layer is used to convert the high-dimensional mapping features into a fault probability distribution, and output the bearing fault type based on the fault probability distribution.
[0017] Further, the deep bearing fault extraction module includes a first convolution block, a second convolution block, a third convolution block and a fourth convolution block connected in sequence;
[0018] The first convolution block is used to extract the first high-dimensional mapping feature from the primary bearing fault feature, obtain the first fault feature, and transmit it to the second convolution block;
[0019] The second convolution block is used to extract a second high-dimensional mapping feature of the first fault feature, obtain a second fault feature, and transmit it to the third convolution block;
[0020] The third convolution block is used to extract a third high-dimensional mapping feature of the second fault feature, obtain a third fault feature, and transmit it to the fourth convolution block;
[0021] The fourth convolution block is used to extract a fourth high-dimensional mapping feature of the third fault feature, obtain a fourth fault feature, and transmit it to the output layer.
[0022] Furthermore, in the first convolution block, the second convolution block, the third convolution block and the fourth convolution block, each convolution block includes a first channel attention branch module, a second channel attention branch module, a scale fusion layer, a first point-by-point convolution layer and a first ReLU activation function layer;
[0023] The first channel attention branch module and the second channel attention branch module are used to extract the channel attention features at a scale of 3×3 and the channel attention features at a scale of 5×5 of the input data, respectively, to obtain the first channel attention features and the second channel attention features, and transmit them to the scale fusion layer;
[0024] The scale fusion layer is used to add the first channel attention feature and the second channel attention feature to obtain a scale fusion feature, and transmit it to the first point-by-point convolution layer;
[0025] The first point-by-point convolution layer is used to extract the fault feature of the scale fusion feature, obtain the first point-by-point convolution feature, and input it to the first ReLU activation function layer;
[0026] The first ReLU activation function layer is used to perform nonlinear transformation on the first point-by-point convolution feature to obtain a first nonlinear feature.
[0027] Further, the first channel attention branch module includes a first depth-wise convolution layer, a second depth-wise convolution layer, a third depth-wise convolution layer, a second ReLU activation function layer, a first Hadamard product unit, a first global average pooling layer, a second point-wise convolution layer and a third ReLU activation function layer;
[0028] The first depth-by-depth convolution layer is used to extract fault features of input data, obtain first depth-by-depth convolution features, and transmit them to the second depth-by-depth convolution layer and the third depth-by-depth convolution layer;
[0029] The second depth-by-depth convolution layer is used to extract the fault feature of the first depth-by-depth convolution feature, obtain the second depth-by-depth convolution feature, and transmit it to the Hadamard product unit;
[0030] The third depth-by-depth convolution layer is used to extract the fault feature of the first depth-by-depth convolution feature, obtain the third depth-by-depth convolution feature, and transmit it to the second ReLU activation function layer;
[0031] The second ReLU activation function layer is used to perform nonlinear transformation on the third depth-wise convolution feature to obtain a second nonlinear feature, and transmit it to the first Hadamard product unit;
[0032] The first Hadamard product unit is used to perform Hadamard multiplication on the second depth-wise convolution feature and the second nonlinear feature to obtain a first Hadamard feature, and transmit it to the first global average pooling layer;
[0033] The first global average pooling layer is used to reduce the size of the first Hadamard feature, obtain a first average pooling feature, and transmit it to the second point-by-point convolutional layer;
[0034] The second point-by-point convolution layer is used to reduce the dimension of the first average pooling feature to be consistent with the channel dimension of the first depth-by-depth convolution feature, obtain a second point-by-point convolution feature, and transmit it to the third ReLU activation function layer;
[0035] The third ReLU activation function layer is used to perform nonlinear transformation on the second point-by-point convolution feature to obtain the first channel attention feature and transmit it to the scale fusion layer;
[0036] The second channel attention branch module includes a fourth depth-wise convolution layer, a fifth depth-wise convolution layer, a sixth depth-wise convolution layer, a fourth ReLU activation function layer, a second Hadamard product unit, a second global average pooling layer, a third point-wise convolution layer and a fifth ReLU activation function layer;
[0037] The fourth depth-by-depth convolution layer is used to extract fault features of input data, obtain fourth depth-by-depth convolution features, and transmit them to the fifth depth-by-depth convolution layer and the sixth depth-by-depth convolution layer;
[0038] The fifth depth-by-depth convolution layer is used to extract the fault feature of the fourth depth-by-depth convolution feature, obtain the fifth depth-by-depth convolution feature, and transmit it to the second Hadamard product unit;
[0039] The sixth depth-by-depth convolution layer is used to extract the fault feature of the fourth depth-by-depth convolution feature, obtain the sixth depth-by-depth convolution feature, and transmit it to the fourth ReLU activation function layer;
[0040] The fourth ReLU activation function layer is used to perform nonlinear transformation on the sixth depth-wise convolution feature to obtain a fourth nonlinear feature, and transmit it to the second Hadamard product unit;
[0041] The second Hadamard product unit is used to perform Hadamard multiplication on the fifth depth-wise convolution feature and the fourth nonlinear feature to obtain a second Hadamard feature, and transmit it to the second global average pooling layer;
[0042] The second global average pooling layer is used to reduce the size of the second Hadamard feature, obtain a second average pooling feature, and transmit it to the third point-by-point convolutional layer;
[0043] The third point-by-point convolution layer is used to reduce the dimension of the second average pooling feature to be consistent with the channel dimension of the first depth-by-depth convolution feature, obtain the third point-by-point convolution feature, and transmit it to the fifth ReLU activation function layer;
[0044] The fifth ReLU activation function layer is used to perform nonlinear transformation on the third point-by-point convolution feature to obtain the second channel attention feature and transmit it to the scale fusion layer.
[0045] Furthermore, the loss function of the bearing fault diagnosis model is a gradient coordination class balance loss function, which is based on the class balance loss function and is obtained by introducing a gradient density factor.
[0046] Furthermore, the gradient coordination class balance loss function is:
[0047]
[0048] Where N is the effective sample size of each fault type, GD(g) is the gradient density, β is a hyperparameter related to the total sample volume, z is the model's predicted output for all categories, and z c is the predicted output for category c; C is the total number of categories, g i is the gradient modulus of the i-th sample; γ i is the gradient density coordination parameter; L CB (p, y) is the class-balanced loss function;
[0049] The calculation formula of gradient density is:
[0050]
[0051] Gradient density coordination parameter γ i The calculation formula is:
[0052]
[0053] Furthermore, in S2, the process of converting the multiple fault signals into corresponding Gramian matrices using Gramian image coding technology to form a bearing fault image sample set is as follows:
[0054] Assume that the original time series of each fault in the multiple fault signals is: X = {x1, x2, x3, ..., x n}, normalize the original time series to between [-1,1] or [0,1], and the normalization formulas are:
[0055]
[0056] Mapping the normalized original time series into polar coordinates yields:
[0057]
[0058] Among them, φ i Reason The angle r is the mapping of the Cartesian coordinate system to polar coordinates. i Because iTimestamps are mapped from Cartesian coordinates to radii in polar coordinates;
[0059] The correlation of each point in the normalized original time series is quantified, that is, the sum angle formula in the trigonometric function is used to calculate and form a Gram matrix, which is expressed as:
[0060]
[0061] In the formula, φ i (i=1,2,…,n) represents the angle of the i-th point in the normalized original time series transformed to the polar coordinates, and I is the unit row vector.
[0062] Beneficial effect: The present invention constructs a bearing fault diagnosis model, sets training sets according to a variety of imbalance modes, and inputs each training set into the bearing fault diagnosis model for optimization training. The trained bearing fault diagnosis model can realize deep mining and extraction of fault discrimination information for a small number of samples, so that the bearing fault diagnosis model has a higher diagnostic accuracy when facing unbalanced fault data, solving the problem that small sample faults are difficult to distinguish and diagnose. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0064] Figure 1 It is a flow chart of a method for diagnosing ship bearing faults under unbalanced data in the present invention;
[0065] Figure 2 A schematic diagram of the structure of a bearing fault diagnosis model in an embodiment of the present invention;
[0066] Figure 3 is a schematic diagram of the structure of the first channel attention branch module and the second channel attention branch module in an embodiment of the present invention;
[0067] Figure 4 is a schematic diagram of fault clustering of a bearing fault diagnosis model in an embodiment of the present invention under a bearing fault image sample set;
[0068] Figure 5 3 is a comparison chart of the diagnostic effects of the method proposed in the embodiment of the present invention and other methods. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] This embodiment provides a method for diagnosing ship bearing faults under unbalanced data. Figure 1 As shown, the specific steps include:
[0071] S1: obtaining a vibration signal of a normal generator bearing, and processing the vibration signal of the normal generator bearing by using a signal processing technology to obtain a plurality of fault signals;
[0072] Specifically, in this embodiment, the vibration signal of the generator bearing of the ship during normal navigation is collected, and the abnormal time domain characteristic value is injected into the normal bearing signal by using the signal processing technology, so as to obtain simulated signals of various fault forms;
[0073] Specifically, the vibration sensor installed on the engine bearing cover is used to collect the normal operation vibration data of the bearing of the Wärtsilä W8L20 engine on the actual ship. The engine speed is 1000rpm and the maximum load is 1542kW. Based on the analysis of the vibration signal, according to the two types of abnormal manifestations of vibration amplitude point abnormality and collective abnormality, this embodiment simulates a total of four data modes of time domain characteristic abnormalities, among which the first three types of fault data are formed by point abnormalities, the abnormal range of point abnormal vibration amplitude is determined to be [2,6], and the position of the point abnormality in the sequence is randomly determined; the fourth type of fault is formed by collective abnormalities, the mean of the collective abnormality is set to 0, and the variance range is determined to be [1,6], and finally four fault time series are formed.
[0074] S2: using Gram image coding technology to convert the multiple fault signals into corresponding Gram matrices, thereby forming a bearing fault image sample set of RGB three channels;
[0075] S3: Establish a bearing fault diagnosis model, select the initial value of the learning rate, the neuron discard rate, and set the optimization algorithm, model loss function, and learning rate update function;
[0076] Specifically, this embodiment uses the StarNet network as the basic model of the bearing fault diagnosis model. The StarNet network uses star operations (i.e., element multiplication) to significantly amplify the advantages of the implicit feature space dimension, build a compact, lightweight, and efficient network structure, and map the input to a high-dimensional, nonlinear feature space without expanding the deep network, thereby meeting the convenience requirements of the fault diagnosis model in actual engineering.
[0077] Specifically, in this embodiment, the Adam optimization algorithm is used to update the network training parameters, the initial value of the learning rate is set to 0.001, and ReduceLROnPlateau is used to update the learning rate to realize the self-attenuation process of the learning rate. The test set accuracy is used as the adjustment index, the patience in ReduceLROnPlateau is selected as 4, and the gradient coordination class balance loss function is used to calculate the model loss. The neuron discard rate in the model is set to 0.2, the classifier is softmax, the batch size is 64, and the activation function selects the ReLU activation function.
[0078] S4: dividing the bearing fault image sample set into a training set and a test set, wherein the training set is set in a plurality of unbalanced ways, and inputting each training set into the bearing fault diagnosis model for optimization training, and when the loss function of the bearing fault diagnosis model converges, obtaining the optimized and trained bearing fault diagnosis model;
[0079] Specifically, in this embodiment, every 512 data points in the bearing fault image sample set are truncated once to generate a fault sample image with a size of 224×224RGB three channels. The sample set contains a total of 600 sample images, which are divided into a training set and a test set at a ratio of 4:1. The training set is divided according to five unbalanced methods, four of which are to divide the healthy data and the fault data according to four imbalance rates, and obtain training sets with four sample numbers including method 1, method 2, method 3, and method 4. The fifth method is to divide the training set into multiple fault and rare fault training sets according to different fault types to simulate the imbalance phenomenon of multiple faults and rare faults. Specifically, the training sets with four sample numbers including method 1, method 2, method 3, and method 4 are divided according to the imbalance rates of 2:1, 4:1, 8:1 and 20:1 between the healthy data and the fault data, respectively. The data interpretation and division are shown in Table 1.
[0080] Table 1:
[0081]
[0082] S5: using the test set to verify the ship bearing fault classification effect of the optimized and trained bearing fault diagnosis model under the unbalanced sample condition, and obtain the verified bearing fault diagnosis model;
[0083] S6: Classify ship bearing faults under unbalanced data based on the verified bearing fault diagnosis model.
[0084] In a specific embodiment, the bearing fault diagnosis model includes a first convolutional layer, a deep bearing fault extraction module and an output layer connected in sequence;
[0085] The first convolutional layer is used to preliminarily extract bearing fault features in the bearing fault image, and transmit the obtained primary bearing fault features to the deep bearing fault extraction module;
[0086] The deep bearing fault extraction module is used to extract high-dimensional mapping features from the primary bearing fault features and transmit them to the output layer;
[0087] The output layer is used to convert the high-dimensional mapping features into a fault probability distribution, and output the bearing fault type based on the fault probability distribution.
[0088] In a specific embodiment, the deep bearing fault extraction module includes a first convolution block, a second convolution block, a third convolution block and a fourth convolution block connected in sequence for deep mining of ship bearing fault features;
[0089] The first convolution block is used to extract the first high-dimensional mapping feature from the primary bearing fault feature, obtain the first fault feature, and transmit it to the second convolution block;
[0090] The second convolution block is used to extract a second high-dimensional mapping feature of the first fault feature, obtain a second fault feature, and transmit it to the third convolution block;
[0091] The third convolution block is used to extract a third high-dimensional mapping feature of the second fault feature, obtain a third fault feature, and transmit it to the fourth convolution block;
[0092] The fourth convolution block is used to extract a fourth high-dimensional mapping feature of the third fault feature, obtain a fourth fault feature, and transmit it to the output layer.
[0093] Specifically, the size of the feature map output after each convolution block is halved in length and width. For example, if the size of the first fault feature output by the first convolution block is h×w, then the size of the second fault feature output by the second convolution block is The size of the third fault feature output by the third convolutional block is The size of the fourth fault feature output by the fourth convolutional block is
[0094] In a specific embodiment, Figure 2As shown, in the first convolution block, the second convolution block, the third convolution block and the fourth convolution block, each convolution block includes a first channel attention branch module, a second channel attention branch module, a scale fusion layer, a first point-by-point convolution layer with a convolution kernel size of 7×7, and a first ReLU activation function layer;
[0095] Specifically, Figure 2 As shown, in order to distinguish each ReLU activation function layer, in the first convolution block, it is the first ReLU activation function layer, in the second convolution block, it is the sixth ReLU activation function layer; in the third convolution block, it is the eleventh ReLU activation function layer; in the fourth convolution block, it is the sixteenth ReLU activation function layer.
[0096] The first channel attention branch module and the second channel attention branch module are used to extract the channel attention features at a 3×3 scale and the channel attention features at a 5×5 scale of the input data, respectively, to obtain the first channel attention features and the second channel attention features, and transmit them to the scale fusion layer; specifically, the first channel attention features and the second channel attention features represent different degrees of importance.
[0097] The scale fusion layer is used to add the first channel attention feature and the second channel attention feature to obtain a scale fusion feature, and transmit it to the first point-by-point convolution layer;
[0098] Specifically, Figure 2 As shown, the scale fusion feature is F'', which is expressed as:
[0099]
[0100] Where F3” is the output of the attention channel with a convolution kernel size of 3, and F5” is the output of the attention channel with a convolution kernel size of 5; DW is a depth-wise convolution operation, PW is a point-wise convolution operation, σ is the ReLU activation function, is the Hadamard multiplication.
[0101] The first point-by-point convolution layer is used to extract the fault feature of the scale fusion feature, obtain the first point-by-point convolution feature, and input it to the first ReLU activation function layer;
[0102] The first ReLU activation function layer is used to perform nonlinear transformation on the first point-by-point convolution feature to obtain a first nonlinear feature.
[0103] In a specific embodiment, Figure 3As shown, the first channel attention branch module includes a first depth-wise convolution layer with a convolution kernel size of 3×3, a second depth-wise convolution layer with a convolution kernel size of 1×1, a third depth-wise convolution layer with a convolution kernel size of 1×1, a second ReLU activation function layer, a first Hadamard product unit, a first global average pooling layer, a second point-wise convolution layer with a convolution kernel size of 1×1, and a third ReLU activation function layer;
[0104] The first depth-by-depth convolution layer is used to extract fault features of input data, obtain first depth-by-depth convolution features, and transmit them to the second depth-by-depth convolution layer and the third depth-by-depth convolution layer;
[0105] The second depth-by-depth convolution layer is used to extract the fault feature of the first depth-by-depth convolution feature, obtain the second depth-by-depth convolution feature, and transmit it to the Hadamard product unit;
[0106] The third depth-by-depth convolution layer is used to extract the fault feature of the first depth-by-depth convolution feature, obtain the third depth-by-depth convolution feature, and transmit it to the second ReLU activation function layer;
[0107] The second ReLU activation function layer is used to perform nonlinear transformation on the third depth-wise convolution feature to obtain a second nonlinear feature, and transmit it to the first Hadamard product unit;
[0108] The first Hadamard product unit is used to perform Hadamard multiplication on the second depth-wise convolution feature and the second nonlinear feature to obtain a first Hadamard feature, and transmit it to the first global average pooling layer;
[0109] The first global average pooling layer is used to reduce the size of the first Hadamard feature, obtain a first average pooling feature, and transmit it to the second point-by-point convolutional layer;
[0110] The second point-by-point convolution layer is used to reduce the dimension of the first average pooling feature to be consistent with the channel dimension of the first depth-by-depth convolution feature, obtain a second point-by-point convolution feature, and transmit it to the third ReLU activation function layer;
[0111] The third ReLU activation function layer is used to perform nonlinear transformation on the second point-by-point convolution feature to obtain the first channel attention feature and transmit it to the scale fusion layer;
[0112] The second channel attention branch module includes a fourth depth-wise convolution layer with a convolution kernel size of 5×5, a fifth depth-wise convolution layer with a convolution kernel size of 1×1, a sixth depth-wise convolution layer with a convolution kernel size of 1×1, a fourth ReLU activation function layer, a second Hadamard product unit, a second global average pooling layer, a third point-wise convolution layer with a convolution kernel size of 1×1, and a fifth ReLU activation function layer;
[0113] The fourth depth-by-depth convolution layer is used to extract fault features of input data, obtain fourth depth-by-depth convolution features, and transmit them to the fifth depth-by-depth convolution layer and the sixth depth-by-depth convolution layer;
[0114] The fifth depth-by-depth convolution layer is used to extract the fault feature of the fourth depth-by-depth convolution feature, obtain the fifth depth-by-depth convolution feature, and transmit it to the second Hadamard product unit;
[0115] The sixth depth-by-depth convolution layer is used to extract the fault feature of the fourth depth-by-depth convolution feature, obtain the sixth depth-by-depth convolution feature, and transmit it to the fourth ReLU activation function layer;
[0116] The fourth ReLU activation function layer is used to perform nonlinear transformation on the sixth depth-wise convolution feature to obtain a fourth nonlinear feature, and transmit it to the second Hadamard product unit;
[0117] The second Hadamard product unit is used to perform Hadamard multiplication on the fifth depth-wise convolution feature and the fourth nonlinear feature to obtain a second Hadamard feature, and transmit it to the second global average pooling layer;
[0118] Specifically, Figure 3 As shown, F' is the second Hadamard characteristic, expressed as:
[0119]
[0120] The second global average pooling layer is used to reduce the size of the second Hadamard feature to obtain a second average pooling feature, thereby reducing the amount of parameters and calculations, and transmit it to the third point-by-point convolutional layer;
[0121] The third point-by-point convolution layer is used to reduce the dimension of the second average pooling feature to be consistent with the channel dimension of the first depth-by-depth convolution feature, so as to obtain the attention weight of each channel, obtain the third point-by-point convolution feature, and transmit it to the fifth ReLU activation function layer;
[0122] Specifically, Figure 3 As shown, the third point-by-point convolution feature is F', which is expressed as:
[0123]
[0124] The fifth ReLU activation function layer is used to perform nonlinear transformation on the third point-by-point convolution feature to obtain the second channel attention feature and transmit it to the scale fusion layer.
[0125] Specifically, in this embodiment, when designing a multi-scale parallel channel attention structure, namely, a first channel attention branch module and a second channel attention branch module, it is divided into two levels: multi-scale branch and channel attention. The multi-scale branch is used to aggregate channel features of different scales, and the channel attention is used to weight channels with higher importance to capture the information of the fault vibration signal. By aggregating channel features of different scales, focusing on key information in features of different scales and weighting channels, the focus is on learning channels with more effective information, thereby improving the feature mining and learning capabilities of the bearing fault diagnosis model.
[0126] Specifically, in this embodiment, in the first channel attention branch module, a second ReLU activation function layer is added after the third depth-wise convolution layer, so that it has nonlinear characteristics, so that the model can learn complex function mapping relationships; no activation function is added after the second depth-wise convolution layer, so that it has linear transformation characteristics, so as to retain the spatial information of the input data. At the same time, through the double-layer depth-wise convolution channel, the features are projected into a high-dimensional space, so that deep mining and extraction of information can be achieved.
[0127] In a specific embodiment, the loss function of the bearing fault diagnosis model is a gradient coordination class balance loss function, which is based on the class balance loss function and is obtained by introducing a gradient density factor. The gradient coordination class balance loss function can allocate higher attention to the bearing fault category of a minority sample, thereby accurately identifying the bearing fault of a minority sample class.
[0128] Specifically, in this embodiment, the learning cost sensitivity of the model is changed by designing a gradient coordination class balance loss function, so that the bearing fault diagnosis model is prompted to avoid outlier interference as much as possible while paying attention to categories with fewer valid samples, alleviate the negative contribution of outliers to the effective sample volume, and guide the diagnosis model to give lower attention to samples with a larger number in the gradient distribution (easy-to-separate samples and outlier samples), and focus on learning the effective features of a small number of sample classes in a targeted manner, so that the model can achieve different degrees of feature mining according to the correct proportion of the effective sample space of each type of fault, thereby improving the recognition accuracy of bearing fault categories for a small number of samples.
[0129] In a specific embodiment, the process of designing the gradient coordination class balance loss function is:
[0130] The softmax loss function is differentiated with respect to x to define the gradient modulus g, which represents the distance scalar between the true value of the sample and the predicted value, expressed as:
[0131]
[0132] Based on the gradient modulus g, the range length l within a certain gradient range is obtained ∈ (g), expressed as:
[0133]
[0134] Among them, l ∈ (g) indicates The effective length of the interval;
[0135] At the same time, the number of samples δ within a certain gradient range is calculated ∈ (x,y), expressed as:
[0136]
[0137] Among them, δ ∈ (g i ,g) means that among the N samples, the gradient modulus is distributed in The number of samples in the range;
[0138] Then, the calculation formula of gradient density is obtained as follows:
[0139]
[0140] Among them, g i is the gradient modulus of the i-th sample;
[0141] The gradient density coordination parameter γ is calculated based on the gradient density i , expressed as:
[0142]
[0143] If the samples are uniformly distributed with respect to the gradient, for any g i All have GD (g i )=N, then γ i No sample will be weighted, and samples with large gradient density will have a larger GD (g i ) and a smaller γ i , so as to increase the weight of truly valuable samples.
[0144] Based on the gradient density coordination parameter, the gradient coordination class balance loss function is expressed as:
[0145]
[0146] Where N is the effective sample size of each fault type, GD(g) is the gradient density, β is a hyperparameter related to the total sample volume, z is the model's predicted output for all categories, and z c is the predicted output for category c; C is the total number of categories, g i is the gradient modulus of the i-th sample; γi is the gradient density coordination parameter; L CB (p, y) is the class-balanced loss function;
[0147] In a specific embodiment, in S2, the process of converting the multiple fault signals into corresponding Gramian matrices using Gramian image coding technology to form a bearing fault image sample set is as follows:
[0148] Assume that the original time series of each fault in the multiple fault signals is: X = {x1, x2, x3, ..., x n}, normalize the original time series to between [-1,1] or [0,1]:
[0149] The formula for normalization to [-1,1] is:
[0150]
[0151] The formula for normalization to [0,1] is:
[0152]
[0153] Mapping the normalized original time series into polar coordinates yields:
[0154]
[0155] Among them, φ i Reason The angle r is the mapping of the Cartesian coordinate system to polar coordinates. i Because i Timestamps are mapped from Cartesian coordinates to radii in polar coordinates;
[0156] The correlation of each point in the normalized original time series is quantified, that is, the sum angle formula in the trigonometric function is used to calculate and form a Gram matrix, which is expressed as:
[0157]
[0158] In the formula, φ i (i=1,2,…,n) represents the angle of the i-th point in the normalized original time series transformed to the polar coordinates, and I is the unit row vector.
[0159] Specifically, Figure 4As shown in the figure, the t-distribution-random neighbor embedding technology is used to visualize the distribution of features in three-dimensional space, verifying the performance of the proposed method under balanced data sets. To prove the superiority of the proposed method, the proposed method is compared with other latest and most classic network models (Resnet, RepViT) and loss function (focus loss, class balance loss) to verify the fault diagnosis performance of the proposed method under data sets with various imbalance rates and imbalance modes. The results are shown in the figure. Figure 5 As shown in the figure, it can be seen that under low imbalance ratio conditions, such as mode 1 (2:1) and mode 2 (4:1), the diagnostic performance indicators of all methods can be maintained above 90%, and the diagnostic effect of the method proposed in the present invention can reach 100%. As the imbalance ratio increases, the diagnostic effect of the method proposed in the present invention decreases, but the decrease is small. Even under a high imbalance ratio, its F1 score and G-mean always exceed 97%, which is mainly due to the multi-scale parallel channel attention structure embedded in the model, which increases the information analysis of a small number of sample classes under the guidance of different penalty levels. In contrast, other comparison methods are greatly affected by the imbalance ratio, especially when the imbalance ratio reaches 20:1, the diagnostic effect of RepVit-Focal is relatively poor. One of the main reasons is that the feature learning ability and semantic expression ability of the lightweight model RepViT in this method are limited. When there are fewer samples, the feature distribution and judgment information of different fault samples cannot be learned, resulting in more serious fault confusion. This can be concluded from all the sample imbalance division methods. The fault recognition indicators of the method composed of RepVit are generally lower than those of the method composed of Resnet. The method proposed in the present invention can not only correctly identify sample types under various imbalance ratios of healthy and faulty samples, but also better handle the imbalance between fault classes.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A ship bearing fault diagnosis method under unbalanced data, characterized in that: The specific steps include: S1: obtaining a vibration signal of a normal generator bearing, and processing the vibration signal of the normal generator bearing by using a signal processing technology to obtain a plurality of fault signals; S2: using Gram image coding technology to convert the multiple fault signals into corresponding Gram matrices, thereby forming a bearing fault image sample set; S3: Establish bearing fault diagnosis model; S4: dividing the bearing fault image sample set into a training set and a test set, wherein the training set is set in a plurality of unbalanced ways, and inputting each training set into the bearing fault diagnosis model for optimization training, and when the loss function of the bearing fault diagnosis model converges, obtaining the optimized and trained bearing fault diagnosis model; S5: using the test set to verify the ship bearing fault classification effect of the optimized and trained bearing fault diagnosis model under the unbalanced sample condition, and obtain the verified bearing fault diagnosis model; S6: Classify ship bearing faults under unbalanced data based on the verified bearing fault diagnosis model.
2. The ship bearing fault diagnosis method under unbalanced data according to claim 1 is characterized in that: The bearing fault diagnosis model comprises a first convolutional layer, a deep bearing fault extraction module and an output layer connected in sequence; The first convolutional layer is used to preliminarily extract bearing fault features in the bearing fault image, and transmit the obtained primary bearing fault features to the deep bearing fault extraction module; The deep bearing fault extraction module is used to extract high-dimensional mapping features from the primary bearing fault features and transmit them to the output layer; The output layer is used to convert the high-dimensional mapping features into a fault probability distribution, and output the bearing fault type based on the fault probability distribution.
3. The ship bearing fault diagnosis method under unbalanced data according to claim 2 is characterized in that: The deep bearing fault extraction module includes a first convolution block, a second convolution block, a third convolution block and a fourth convolution block connected in sequence; The first convolution block is used to extract the first high-dimensional mapping feature from the primary bearing fault feature, obtain the first fault feature, and transmit it to the second convolution block; The second convolution block is used to extract a second high-dimensional mapping feature of the first fault feature, obtain a second fault feature, and transmit it to the third convolution block; The third convolution block is used to extract a third high-dimensional mapping feature of the second fault feature, obtain a third fault feature, and transmit it to the fourth convolution block; The fourth convolution block is used to extract a fourth high-dimensional mapping feature of the third fault feature, obtain a fourth fault feature, and transmit it to the output layer.
4. The ship bearing fault diagnosis method under unbalanced data according to claim 3 is characterized in that: In the first convolution block, the second convolution block, the third convolution block and the fourth convolution block, each convolution block includes a first channel attention branch module, a second channel attention branch module, a scale fusion layer, a first point-by-point convolution layer and a first ReLU activation function layer; The first channel attention branch module and the second channel attention branch module are used to extract the channel attention features at a scale of 3×3 and the channel attention features at a scale of 5×5 of the input data, respectively, to obtain the first channel attention features and the second channel attention features, and transmit them to the scale fusion layer; The scale fusion layer is used to add the first channel attention feature and the second channel attention feature to obtain a scale fusion feature, and transmit it to the first point-by-point convolution layer; The first point-by-point convolution layer is used to extract the fault feature of the scale fusion feature, obtain the first point-by-point convolution feature, and input it to the first ReLU activation function layer; The first ReLU activation function layer is used to perform nonlinear transformation on the first point-by-point convolution feature to obtain a first nonlinear feature.
5. The ship bearing fault diagnosis method under unbalanced data according to claim 4 is characterized in that: The first channel attention branch module includes a first depth-wise convolution layer, a second depth-wise convolution layer, a third depth-wise convolution layer, a second ReLU activation function layer, a first Hadamard product unit, a first global average pooling layer, a second point-wise convolution layer and a third ReLU activation function layer; The first depth-by-depth convolution layer is used to extract fault features of input data, obtain first depth-by-depth convolution features, and transmit them to the second depth-by-depth convolution layer and the third depth-by-depth convolution layer; The second depth-by-depth convolution layer is used to extract the fault feature of the first depth-by-depth convolution feature, obtain the second depth-by-depth convolution feature, and transmit it to the Hadamard product unit; The third depth-by-depth convolution layer is used to extract the fault feature of the first depth-by-depth convolution feature, obtain the third depth-by-depth convolution feature, and transmit it to the second ReLU activation function layer; The second ReLU activation function layer is used to perform nonlinear transformation on the third depth-wise convolution feature to obtain a second nonlinear feature, and transmit it to the first Hadamard product unit; The first Hadamard product unit is used to perform Hadamard multiplication on the second depth-wise convolution feature and the second nonlinear feature to obtain a first Hadamard feature, and transmit it to the first global average pooling layer; The first global average pooling layer is used to reduce the size of the first Hadamard feature, obtain a first average pooling feature, and transmit it to the second point-by-point convolutional layer; The second point-by-point convolution layer is used to reduce the dimension of the first average pooling feature to be consistent with the channel dimension of the first depth-by-depth convolution feature, obtain a second point-by-point convolution feature, and transmit it to the third ReLU activation function layer; The third ReLU activation function layer is used to perform nonlinear transformation on the second point-by-point convolution feature to obtain the first channel attention feature and transmit it to the scale fusion layer; The second channel attention branch module includes a fourth depth-wise convolution layer, a fifth depth-wise convolution layer, a sixth depth-wise convolution layer, a fourth ReLU activation function layer, a second Hadamard product unit, a second global average pooling layer, a third point-wise convolution layer and a fifth ReLU activation function layer; The fourth depth-by-depth convolution layer is used to extract fault features of input data, obtain fourth depth-by-depth convolution features, and transmit them to the fifth depth-by-depth convolution layer and the sixth depth-by-depth convolution layer; The fifth depth-by-depth convolution layer is used to extract the fault feature of the fourth depth-by-depth convolution feature, obtain the fifth depth-by-depth convolution feature, and transmit it to the second Hadamard product unit; The sixth depth-by-depth convolution layer is used to extract the fault feature of the fourth depth-by-depth convolution feature, obtain the sixth depth-by-depth convolution feature, and transmit it to the fourth ReLU activation function layer; The fourth ReLU activation function layer is used to perform nonlinear transformation on the sixth depth-wise convolution feature to obtain a fourth nonlinear feature, and transmit it to the second Hadamard product unit; The second Hadamard product unit is used to perform Hadamard multiplication on the fifth depth-wise convolution feature and the fourth nonlinear feature to obtain a second Hadamard feature, and transmit it to the second global average pooling layer; The second global average pooling layer is used to reduce the size of the second Hadamard feature, obtain a second average pooling feature, and transmit it to the third point-by-point convolutional layer; The third point-by-point convolution layer is used to reduce the dimension of the second average pooling feature to be consistent with the channel dimension of the first depth-by-depth convolution feature, obtain the third point-by-point convolution feature, and transmit it to the fifth ReLU activation function layer; The fifth ReLU activation function layer is used to perform nonlinear transformation on the third point-by-point convolution feature to obtain the second channel attention feature and transmit it to the scale fusion layer.
6. The ship bearing fault diagnosis method under unbalanced data according to claim 5 is characterized in that: The loss function of the bearing fault diagnosis model is a gradient coordination class balance loss function, which is based on the class balance loss function and is obtained by introducing a gradient density factor.
7. The ship bearing fault diagnosis method under unbalanced data according to claim 6 is characterized in that: The gradient coordination class balance loss function is: Where N is the effective sample size for each fault type, GD(g) is the gradient density, β is a hyperparameter related to the total sample volume, z is the model's predicted output for all categories, and z = [z1, z2, …, z C ] T , z c is the predicted output for category c; C is the total number of categories, g i is the gradient modulus of the i-th sample; γ i is the gradient density coordination parameter; L CB (p, y) is the class-balanced loss function; The calculation formula of gradient density is: Gradient density coordination parameter γ i The calculation formula is:
8. The ship bearing fault diagnosis method under unbalanced data according to claim 7 is characterized in that: In S2, the process of converting the multiple fault signals into corresponding Gram matrices using Gram image coding technology to form a bearing fault image sample set is as follows: Assume that the original time series of each fault in the multiple fault signals is: X = {x1, x2, x3, ..., x n }, normalize the original time series to between [-1,1] or [0,1], and the normalization formulas are: Mapping the normalized original time series into polar coordinates yields: Among them, φ i Reason The angle r is the mapping of the Cartesian coordinate system to polar coordinates. i Because i Timestamps are mapped from Cartesian coordinates to radii in polar coordinates; The correlation of each point in the normalized original time series is quantified, that is, the sum angle formula in the trigonometric function is used to calculate and form a Gram matrix, which is expressed as: In the formula, φ i (i=1,2,…,n) represents the angle of the i-th point in the normalized original time series transformed to the polar coordinates, and I is the unit row vector.
Citation Information
Cited By
Wind turbine generator fault diagnosis data processing method based on machine learning algorithm
CN121256425A