Rolling bearing fault diagnosis method based on PGStarNet-MFE model
By embedding a multi-scale feature extraction module into the StarNet network and designing a linear weighted loss function based on the improved Big Cane Rat optimization algorithm, the problem of insufficient feature representation in the rolling bearing fault diagnosis model is solved, achieving efficient multi-level feature integration and representation, and improving the accuracy and convergence speed of fault diagnosis.
Patent Information
- Application Number
- CN202511025642.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to achieve multi-level and comprehensive feature representation, and have poor ability to integrate and express complex features, resulting in poor performance of rolling bearing fault diagnosis models in multi-class classification tasks.
The PGStarNet-MFE model is adopted, which improves feature extraction and classification capabilities by embedding a multi-scale feature extraction module into the StarNet network and designing a linear weighted loss function that improves the design of the large cane rat optimization algorithm.
It significantly improves the model's multi-level feature representation capability and diagnostic efficiency, enhances the accuracy and convergence speed of rolling bearing fault diagnosis, and meets the real-time requirements of industrial scenarios.
Smart Images

Figure CN120976562A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault diagnosis technology, and specifically to a method for diagnosing rolling bearing faults based on the PGStarNet-MFE model. Background Technology
[0002] Rolling bearings, as a key component of rotating machinery, have long been widely used in various fields such as automobile manufacturing, energy production, and aerospace. However, due to their typically harsh operating environments, often facing complex conditions such as continuous vibration, extreme temperature differences, and high-intensity loads, the failure rate of rolling bearings is significantly increased. Once a rolling bearing fails, it not only triggers a chain reaction of damage to the mechanical system and causes direct economic losses, but may also lead to major safety accidents due to mechanical loss of control, threatening personnel lives. Therefore, developing efficient bearing condition monitoring and fault diagnosis technologies is of great significance for ensuring stable equipment operation and improving the safety of industrial production.
[0003] In recent years, with the rapid development of artificial intelligence and big data technologies, deep learning models have demonstrated significant advantages in the field of intelligent fault diagnosis of rolling bearings due to their powerful local perception and pattern recognition capabilities. In particular, intelligent diagnostic models based on convolutional neural networks can effectively achieve end-to-end intelligent diagnosis from raw data to fault categories, and have gradually become a research hotspot. Among them, the StarNet network achieves efficient fusion of cross-subspace features through its core star operation. This operation uses element-wise multiplication to nonlinearly integrate abstract information from different feature subspaces, giving it both high performance and low latency advantages within a compact network architecture. Simultaneously, it can acquire high-dimensional and nonlinear feature spaces in a low-dimensional space and perform computations. Its unique network architecture and feature fusion method significantly reduce computational resource consumption. However, StarNet still suffers from insufficient initial feature extraction in the multi-fault classification scenario of rolling bearings. It only performs single-level feature learning through the initial convolutional layer, making it difficult to achieve multi-level and comprehensive feature representation. At the same time, when the cross-entropy loss function is used for optimization in the classification layer, the inherent correlation between categories is not fully considered, resulting in a discrete distribution of the features learned by the model. This severely restricts its ability to integrate and express complex features, ultimately reducing the model's performance in multi-class classification tasks.
[0004] The above problems urgently need to be solved. To address this, a rolling bearing fault diagnosis method based on the PGStarNet-MFE model is proposed. Summary of the Invention
[0005] The technical problem to be solved by this invention is: how to solve the problem that existing technologies have difficulty in achieving multi-level and comprehensive feature expression and poor ability to integrate and express complex features, and provides a rolling bearing fault diagnosis method based on the PGStarNet-MFE model.
[0006] like Figure 6 As shown, the present invention solves the above-mentioned technical problems through the following technical solution, and the present invention includes the following steps:
[0007] S1: Acquire raw signal
[0008] Collect one-dimensional vibration acceleration signals of rolling bearings under different fault conditions;
[0009] S2: Building the dataset
[0010] The one-dimensional vibration acceleration signal is converted into a two-dimensional time-frequency graph by continuous wavelet transform. The two-dimensional time-frequency graph is used to construct a rolling bearing fault diagnosis dataset, which is then divided into a training set and a test set.
[0011] S3: Network Construction
[0012] Using StarNet as the baseline model, a multi-scale feature extraction module is embedded into StarNet, and a linear weighted loss function designed based on the improved large cane rat optimization algorithm is integrated into the classification layer of StarNet to obtain the PGStarNet-MFE network.
[0013] S4: Model Training
[0014] The PGStarNet-MFE network is trained using the training set, and the optimal weights obtained during the training process are saved as the final model parameters, thereby obtaining the rolling bearing fault diagnosis model, namely the PGStarNet-MFE model.
[0015] S5: Fault Diagnosis
[0016] The test set data is input into the rolling bearing fault diagnosis model to perform end-to-end fault diagnosis and output the fault diagnosis results.
[0017] Furthermore, in step S1, the different fault state types of rolling bearings include healthy, inner ring fault, outer ring fault, rolling element fault, and mixed faults involving the inner ring, outer ring, and rolling elements.
[0018] Furthermore, in step S3, the multi-head feature extraction module is embedded into the stem layer of the StarNet network for multi-scale feature extraction of the two-dimensional time-frequency map.
[0019] Furthermore, in step S3, the PGStarNet-MFE network extracts features from the two-dimensional time-frequency map through a multi-scale feature extraction module. This multi-scale feature extraction module includes four independent branches: a first convolutional branch, a second convolutional branch, a third convolutional branch, and a global average pooling branch. In the first convolutional branch, small-scale output features are extracted through a 1×1 convolutional layer. In the second convolutional branch, medium-scale output features are extracted through a 3×3 convolutional layer. In the third convolutional branch, large-scale output features are extracted through a 3×3 convolutional layer. In the global average pooling branch, the global context information is combined with local convolutional operations by calculating the global average value of each channel in the feature map to obtain the output features. The output features of each branch are concatenated along the channel dimension to form fused features, which are then transformed through a 4×4 convolutional layer to obtain multi-scale features.
[0020] Furthermore, in step S3, the linear weighted loss function L LF Specifically as follows:
[0021] L LF =L CE +λL t
[0022] Among them, L CE Let L be the cross-entropy loss function. t Let λ be the triplet loss function, and λ be the dynamic weight coefficient. The optimization is performed iteratively by improving the Big Cane Mouse optimization algorithm.
[0023] Furthermore, the cross-entropy loss function L CE The expression is as follows:
[0024]
[0025] Where N is the number of samples, y im y′ is the true label of the Mth class of sample i. im It is the predicted probability of the Mth class of sample i;
[0026] Triple loss function L t The expression is as follows:
[0027] L t =max(0,d(a,p)-d(a,n)+ε)
[0028] Where d(a,p) and d(a,n) represent the Euclidean distances between sample a and positive sample p and negative sample n, respectively, and ε is the margin parameter.
[0029] Furthermore, the improved Big Cane Rat optimization algorithm incorporates Piecewise chaotic mapping into the original Big Cane Rat optimization algorithm, thereby generating an ergodic chaotic sequence from the initial population through Piecewise chaotic mapping. The specific process of optimizing the dynamic weight coefficient λ is as follows:
[0030] S31: Initialization Phase
[0031] Randomly generate a population X of large cane rats. i,j It is the random position of the i-th large cane rat in the j-th dimension of the population, as shown in the following formula:
[0032]
[0033] X i,j =X k+1 ×(UB j -LB j )+LB j
[0034] Among them, X k ∈[0,1], its initial value X is a random number that follows a uniform distribution; q is a parameter that controls the mapping effect, and UB and LB are the upper and lower bounds, respectively, used to limit the location range of the giant cane rat population;
[0035] Calculate the fitness value of the global giant cane rat and search the spatial boundary;
[0036] S32: Determine if the maximum number of iterations has been reached. If the maximum number of iterations has been reached, output the position of the large cane rat with the highest fitness value as the dynamic weight coefficient λ. If the maximum number of iterations has not been reached, evaluate the food abundance and generate a random number. Determine if the random number is less than the set stage switching parameter ρ. Enter the corresponding exploration or development stage to update the position of the large cane rat, obtain the new position of the large cane rat, and recalculate the fitness value of each large cane rat. Obtain the position of the large cane rat with the highest fitness value and output it as the dynamic weight coefficient λ. The position of the large cane rat with the highest fitness value is the optimal position of the large cane rat.
[0037] Furthermore, in step S32, the formula for calculating the fitness value of the optimal large cane rat is as follows:
[0038]
[0039] Where X represents all the large cane rats in the current population, and f() is the objective function;
[0040] In step S32, the formula for calculating the stage switching parameter ρ is as follows:
[0041]
[0042] Among them, C iter It is the current iteration number, Max iter It represents the maximum number of iterations.
[0043] Furthermore, in step S32, during the exploration phase, the formula for calculating the new location of the giant cane rat is as follows:
[0044]
[0045] Among them, X i,j new Let X represent the new position of the i-th large cane rat in the j-th dimension. k,j Indicates the position of the optimal individual;
[0046] During the development phase, the formula for calculating the new location of the giant cane rat is as follows:
[0047] X i,j new =X i,j +C×(X k,j -r×X i,j )
[0048] Where C represents a random number defined within the spatial boundary; r is used to simulate the reinforcing effect of abundant food sources on foraging behavior.
[0049] Furthermore, in step S32, during the development phase, if the fitness value of any large cane rat exceeds that of the current best individual, the optimal position is updated and other individuals are guided to migrate; otherwise, a movement strategy is adopted to deviate from the optimal position. The specific movement strategy is as follows:
[0050]
[0051] Among them, F i new F is the fitness value of the optimal large cane rat. i X is the current fitness value, α is the coefficient for the reduction of food sources; β is the coefficient that prompts the algorithm to migrate to other high-value areas during the development phase, and the high-value area is the food-rich area where the optimal cane rat is located; i,j Indicates the current position of the giant cane rat, X. k,j It is the optimal position of the large cane rat in the j-th dimension.
[0052] Compared with existing technologies, this invention has the following advantages: The rolling bearing fault diagnosis method based on the PGStarNet-MFE model innovatively incorporates a multi-head feature extraction (MFE) module, which can effectively achieve multi-level and comprehensive feature extraction and expression. At the same time, it adopts a linear weighted loss function based on an improved large cane rat optimization algorithm to replace the cross-entropy loss function, which effectively solves the problem of inter-class and intra-class distance optimization faced by the cross-entropy loss function when processing two-dimensional fault image data. It overcomes the non-adaptability of the traditional cross-entropy loss function to complex fault features and significantly improves the convergence speed and diagnostic efficiency of the model. Attached Figure Description
[0053] Figure 1 This is a schematic diagram illustrating the implementation process of the rolling bearing fault diagnosis method based on the PGStarNet-MFE model in an embodiment of the present invention.
[0054] Figure 2 These are waveform diagrams of different fault types of rolling bearings in embodiments of the present invention;
[0055] Figure 3 This is a schematic diagram of the continuous wavelet transform time-frequency transformation process in an embodiment of the present invention;
[0056] Figure 4 This is a schematic diagram of the structure of the MFE module in the PGStarNet-MFE model in an embodiment of the present invention;
[0057] Figure 5 This is a flowchart illustrating the design of a linear weighted loss function using an improved large cane rat optimization algorithm in an embodiment of the present invention.
[0058] Figure 6 This is a schematic diagram of the overall process of the rolling bearing fault diagnosis method based on the PGStarNet-MFE model of the present invention. Detailed Implementation
[0059] The embodiments of the present invention are described in detail below. These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. However, the scope of protection of the present invention is not limited to the following embodiments.
[0060] Example 1
[0061] like Figure 1 As shown, this embodiment provides a technical solution: a rolling bearing fault diagnosis method based on the PGStarNet-MFE model, comprising the following steps:
[0062] S1: Acquire one-dimensional vibration acceleration signals of rolling bearings under healthy conditions and various fault types (e.g., ... Figure 2 (As shown).
[0063] S2: Convert the original one-dimensional vibration data of the rolling bearing into a two-dimensional time-frequency graph by continuous wavelet transform, create a rolling bearing fault diagnosis dataset, and divide it into a training set and a test set.
[0064] S3: The multi-scale feature extraction module (MFE) is embedded into the initial feature extraction layer (stem layer) of the StarNet network to perform multi-level feature extraction on the two-dimensional time-frequency map; at the same time, the linear weighted loss function (PGLF) designed based on the improved cane rat optimization algorithm is integrated into the classification layer of the StarNet network to achieve accurate classification prediction.
[0065] S4: The PGStarNet-MFE network is trained using the training set, and the network weights are dynamically updated based on the training process; the optimal weights obtained during the training process are saved as the final model parameters to construct a rolling bearing fault diagnosis model; the rolling bearing condition is evaluated using this diagnosis model, and the fault diagnosis result is finally output.
[0066] Specifically, in step 1, four types of faults were collected: normal (healthy state), inner race fault, outer race fault, and rolling element fault. The healthy state was considered a special fault. The sample length was 1024, and the overlap sampling rate was 0.5. Four folders were created, named after each type of fault. Then, wavelet transform was used to perform time-frequency transformation on the collected four types of fault samples, generating 224×224×3 time-frequency samples. The specific transformation process is as follows: Figure 3 As shown, the final generated time-frequency sample image is saved to the corresponding file.
[0067] In this embodiment, the four types of two-dimensional time-frequency images of faults obtained in step 2 are used as the rolling bearing fault diagnosis dataset. The total number of samples in the dataset is 4000, which are divided into a training set (3200 samples) and a test set (800 samples) in an 8:2 ratio. The specific sample distribution is shown in Table 2.
[0068] Table 2 Distribution of Rolling Bearing Failure Samples
[0069] Fault type Label Sample size Number of training set samples Number of test set samples healthy 0 1000 800 200 Inner ring fault 1 1000 800 200 Outer ring fault 2 1000 800 200 Rolling element failure 3 1000 800 200
[0070] In this embodiment, step S3 embeds the Multi-Head Feature Extraction (MFE) module into the StarNet network to achieve multi-level feature extraction and feature selection of the two-dimensional time-frequency graph. The StarNet-MFE network completes the initial feature extraction through the MFE module. The MFE module completes the multi-level feature extraction task of the input features by fusing convolutional operations at different scales with global average pooling operations. Specifically, the MFE module uses 1×1 convolutional layers to extract small-scale spatial information, integrating feature maps while preserving spatial details, thereby enhancing the model's feature representation ability; at the same time, it captures medium-scale and large-scale feature information through 3×3 and 5×5 convolutional layers, ensuring that the network can comprehensively cover feature representations at different scales. In addition, the global average pooling branch (GAP) combines global context information with local convolutional operations by calculating the global average value of each channel in the feature map, further improving the robustness and feature representation ability of the model. This multi-scale feature extraction and fusion strategy not only enhances the network's ability to perceive local details and global structure, but also provides richer and more discriminative feature representations for complex fault diagnosis tasks.
[0071] It should be noted that small-scale features focus on local micro-features at a single time-frequency point, such as sudden transient signals (impacts, pulses) or high-frequency noise; medium-scale features cover local time-frequency regions, balancing the local context of time and frequency dimensions, and mainly focusing on local harmonic signal features; large-scale features cover a wide range of time-frequency regions, effectively capturing global features.
[0072] like Figure 4 As shown, when the MFE module extracts and learns features from the time-frequency map of a two-dimensional rolling bearing fault, the input feature map is X∈R. C×H×W Where C represents the channel size, and H and W represent the height and width of the input sample, respectively. The MFE module implements multi-level feature extraction through a multi-branch parallel structure. Specifically, the input feature map X is processed through four independent branches, including a 1×1 convolution branch, a 3×3 convolution branch, a 5×5 convolution branch, and a global average pooling branch. The output feature of each branch is denoted as F. i (Where i = 1, 2, 3, 4), corresponding to feature information at different scales. Then, the output features of each branch are concatenated along the channel dimension to form a fused feature. To further adjust the size of the feature map and enhance the expressive power of the multi-scale features, the concatenated features are transformed through a 4×4 convolutional layer, ultimately obtaining the multi-scale feature X. c The model utilizes the cascaded convolutional structure of the MFE module for small-scale, medium-scale, and large-scale feature extraction, and combines global context with local context to improve the model's robustness and representational ability. The specific processing steps are as follows:
[0073] F i =Conv(k i ×k i ,d i i = 1, 2, 3, 4
[0074] X c =Conv[σ(F1(X));σ(F2(X));σ(F3(X));g a (F4(X))∈R C×H×W ]
[0075] Among them, F i X represents the output feature of the i-th branch. c Represents the final multi-scale features, Conv(·) represents the convolution operation, and k i Represents the kernel size, σ(·) is the Rectified Linear Unit (ReLU) activation function, and g a (·) represents the global average pooling operation.
[0076] The parameters of each layer in the MFE module are shown in Table 2.1 below:
[0077] Table 2.1 Parameters of each layer of the MFE module
[0078] Layer name kernel size Expansion rate Step length 1×1Conv 1×1 1 0 3×3Conv 3×3 2 2 5×5Conv 5×5 3 6 GAP - - -
[0079] In this embodiment, in step S3, the feature data extracted and filtered by the StarNet-MFE network is fed into a classification layer (Softmax+PGLF) incorporating a linear weighted loss function designed with an improved Big Cane Mouse optimization algorithm for classification prediction. The linear loss function (PGLF) in the improved Big Cane Mouse optimization algorithm dynamically balances the inter-class distance penalty term and the intra-class distance constraint term, explicitly modeling inter-class separability and intra-class compactness. This effectively solves the problem of co-optimizing inter-class separability and intra-class compactness, thereby improving the model's classification performance for complex data distributions. The PGLF algorithm flow is as follows: Figure 5 As shown, the specific construction process is as follows:
[0080] The linearly weighted loss function is constructed by linearly weighting the cross-entropy loss function and the triplet loss function. The specific construction steps are as follows:
[0081] L LF =L CE +λL t
[0082] Among them, L CE L is the cross-entropy loss function used to optimize the inter-class distance. t λ is the triplet loss function used to optimize intra-class distance, and λ is the dynamic weight coefficient, which ranges from 0 to 1.
[0083] Cross-entropy loss function L CE The expression is as follows:
[0084]
[0085] Where N is the number of samples, y im Let y' be the true label of the Mth class of sample i (1 indicates belonging to this class, 0 indicates not belonging to this class). im It is the predicted probability of the Mth class of sample i;
[0086] Triple loss function L t The expression is as follows:
[0087] L t =max(0,d(a,p)-d(a,n)+ε)
[0088] Where d(a,p) and d(a,n) represent the Euclidean distances between sample a and positive sample p and negative sample n, respectively, and ε is the margin parameter used to represent the threshold distance.
[0089] The dynamic weight coefficient λ is iteratively optimized using an improved Big Cane Rat optimization algorithm. This algorithm incorporates a Piecewise chaotic mapping into the Big Cane Rat optimization algorithm, generating an ergodic chaotic sequence from the initial population through a mapping mechanism. The specific optimization process is as follows:
[0090] (1) Initialization phase. Randomly generate a population of large cane rats X, X i,j It is the random position of the i-th large cane rat in the j-th dimension of the population, as shown in the following formula:
[0091]
[0092] X i,j =X k+1 ×(UB j -LB j )+LB j
[0093] Among them, X k ∈[0,1], its initial value X is a random number that follows a uniform distribution; q is a parameter that controls the mapping effect, and q = 0.4 was selected through multiple experiments; UB and LB are the upper and lower bounds, respectively, which restrict the location range of the large cane rat population and ensure that the algorithm is carried out within a reasonable search space.
[0094] The parameter ρ controls the switching between the exploration (global search) and development (local search) phases of the algorithm. The initial value of ρ requires rigorous parameter analysis to ensure a balance between exploration and development. Experiments have shown that an initial value of 0.5 for ρ is used to achieve high performance in multi-round iterative optimization. The mathematical expression for the value of ρ is as follows:
[0095]
[0096] Among them, C iter It is the current iteration number, Max iter It represents the maximum number of iterations.
[0097] (2) Exploration Phase. During this phase, the giant cane rat migrates between multiple shelters within its territory to forage, leaving its path. This behavior simulates the global exploration process in the algorithm, helping to discover multiple potential optimal solutions. The mathematical expression for the giant cane rat's new location is as follows:
[0098]
[0099] Among them, X i,j new Let X represent the new position of the i-th large cane rat in the j-th dimension. k,j This indicates the position of the optimal individual.
[0100] (3) Development stage. During the mating season, the population updates the search space positions of the remaining individuals based on the location information of the optimal large cane rat. The specific implementation process is shown in the following formula:
[0101] X i,j new =X i,j +C×(X k,j -r×X i,j )
[0102] Where C represents a random number defined within the problem space boundary, used to simulate the random distribution characteristics of scattered food sources and shelters; r is used to simulate the reinforcing effect of abundant food sources on foraging behavior, prompting the algorithm to explore high-value areas more focusedly during the development phase and accelerate convergence.
[0103] Once the development phase begins, if the objective function (fitness) value of a certain large cane rat surpasses that of the current best individual, the optimal position is updated and other individuals are guided to migrate; otherwise, a movement strategy is adopted to deviate from the optimal position (meaning the large cane rat being evaluated will deviate from the optimal position). The specific movement process is as follows:
[0104]
[0105] Among them, F inew The objective function value of the optimal large cane rat, F i X is the current value of the objective function; α is the coefficient for the reduction of food sources; β is the coefficient that prompts the algorithm to migrate to other high-value areas during the breeding season (development phase), where the high-value area is the food-rich area where the optimal cane rat is located; i,j Indicates the current position of the giant cane rat, X. k,j It is the optimal large cane rat in the j-th dimension.
[0106] The migration condition during the exploration phase is as follows: an individual migrates to the new location only if the objective function value of the new location increases; otherwise, it remains in its original location. This mechanism achieves a dynamic balance between global exploration and local exploitation by simulating the biological behavior of large cane rats foraging intensively during the rainy season.
[0107] It should be noted that, in Figure 5 The specific process for assessing food richness is as follows:
[0108] The first step is to calculate the simulated food source coefficient r to evaluate the food availability in the current area. The formula for calculating the coefficient r is as follows:
[0109]
[0110] The second step is to calculate the food source reduction coefficient α, which represents the reduction in food consumption in the current area. When resources in the current area are exhausted, the algorithm is forced to find new food sources or shelters by increasing the exploration intensity (e.g., increasing randomness), thus avoiding getting trapped in local optima. The formula for calculating the coefficient α is as follows:
[0111] α=2×r×rand-r
[0112] Here, rand refers to a random number between 0 and 1;
[0113] The third step is to calculate the coefficient β that prompts the algorithm to migrate to other high-value areas during the development phase. After the resources in the current area are depleted, migration to higher-value areas, i.e., to the location of dominant mice with richer food, optimizes the local search capability of the solution by enhancing the following of dominant locations. The formula for calculating the coefficient β is as follows:
[0114] β=2×r×μ-r
[0115] Where μ is a random number between 1 and 4, used to simulate the number of offspring produced by each female cane rat per year.
[0116] To verify the effectiveness of this invention, this embodiment utilizes a BVT-5 bearing fault vibration measuring instrument (fault diagnosis experimental platform) to collect signal data from rolling bearings. Radial vibration acceleration signals under normal conditions and radial vibration acceleration signals under early pitting conditions of the outer ring, inner ring, and rolling elements are collected. In this experiment, the motor speed is set to 1800 r / min, and the sampling frequency is 10240 Hz. The obtained vibration accelerations are converted into two-dimensional time-frequency graphs using wavelet transform, which are used as the rolling bearing fault dataset. The dataset has a total of 4000 samples, divided into a training set (3200 samples) and a test set (800 samples) in an 8:2 ratio. The specific sample distribution is shown in Table 2 above.
[0117] To verify the effectiveness of this invention, the proposed PGStarNet-MFE model was tested using the aforementioned dataset. Performance comparison experiments were conducted with existing mainstream networks such as VGG16, EfficientNet, GoogleNet, DenseNet, and StarNet. The selected models cover mainstream feature extraction paradigms (such as stacked convolutions, multi-scale fusion, dense connections, and lightweight design), systematically verifying the comprehensive performance improvement of PGStarNet-MFE in terms of feature representation depth, computational efficiency, and multi-scale modeling. Meanwhile, models such as VGG16 and EfficientNet have been widely deployed in industrial equipment monitoring systems, and the comparative experimental results directly reflect the replacement value of the new model in engineering scenarios. Specific results are shown in Table 3.
[0118] Table 3 Comparison of Fault Diagnosis Performance of Different Rolling Bearing Models
[0119]
[0120] The experimental results in Table 3 demonstrate that the proposed PGStarNet-MFE model exhibits significant advantages in rolling bearing fault diagnosis: it surpasses the comparison models (StarNet 97.6%, DenseNet 95.7%) with an accuracy of 99.7%, verifying the role of the multi-scale feature fusion mechanism in improving diagnostic accuracy. Furthermore, the model converges in just 14 training rounds, taking 1790 seconds, a 53.74% improvement over StarNet's convergence speed, reflecting the synergistic effect of the lightweight network structure and the improved optimization algorithm. This model achieves near-industrial-grade reliability (>99.5%) while meeting real-time requirements, providing an innovative solution to the inefficiencies and poor accuracy of traditional deep learning models in engineering scenarios.
[0121] Example 2
[0122] To verify the cross-domain adaptability and engineering applicability of this invention, a multi-condition system was constructed using the Xi'an Jiaotong University dataset. The experimental platform integrated key components such as an ABB AC servo motor (5.5kW power), a Schneider variable frequency speed control system, and a Kistler triaxial vibration sensor (sampling frequency 25.6kHz) to rigorously simulate the bearing degradation process in an industrial setting. The experimental setup covered three typical load conditions: (1) Condition A: 2100r / min, 12kN; (2) Condition B: 2250r / min, 11kN; (3) Condition C: 2400r / min, 10kN. Vibration signals of five types of defects—cage damage (CF), inner ring spalling (IF), outer ring crack (OF), outer ring composite fault (IOF), and mixed fault (MF)—were collected under each condition to construct a multimodal database containing 9000 samples. The original signals were converted into a 224×224×3 pixel time-frequency diagram using the continuous wavelet transform method. The dataset was divided into training and test sets in an 8:2 ratio, as shown in Table 4. The results of the validation experiments are shown in Table 5.
[0123] Table 4 Distribution of Rolling Bearing Failure Samples
[0124]
[0125] Table 5 shows the experimental results comparing the results with other methods on the Xi'an Jiaotong University rolling bearing dataset.
[0126]
[0127] Table 5 shows the experimental analysis based on the Xi'an Jiaotong University rolling bearing dataset, demonstrating that the proposed PGStarNet-MFE network framework exhibits significant advantages in fault diagnosis tasks. As shown in Table 4, the model achieves a fault identification accuracy of 99.5% on the test set, representing performance improvements of 10.0%, 14.4%, 5.8%, 2.1%, and 1.2% respectively compared to the benchmark models VGG16 (89.5%), EfficientNet (85.1%), DenseNet (93.7%), GoogleNet (97.4%), and StarNet (98.3%). Regarding training efficiency, the proposed method exhibits superior convergence characteristics, achieving stable convergence in approximately 15 training epochs, significantly reducing the convergence epochs of the comparison models StarNet (25 epochs), EfficientNet (33 epochs), GoogleNet (28 epochs), DenseNet (31 epochs), and VGG16 (30 epochs). Meanwhile, compared to the baseline model StarNet, the training iteration time is reduced by 48.81%. This accelerated convergence characteristic indicates that the architecture has engineering application advantages in terms of feature extraction efficiency and gradient optimization mechanism. Comparative analysis of experimental data further confirms that the PGStarNet-MFE network, through the synergistic effect of the multi-head feature extraction module (MFE) and the linear loss function based on the improved Big Cane Rat optimization algorithm, not only improves the accuracy of fault mode identification but also enhances the model's generalization ability under complex conditions.
[0128] In summary, the PGStarNet-MFE model of this invention demonstrates its superior performance in highly challenging tasks, fully validating its application potential in fault diagnosis under various real-world operating conditions. Overall, the excellent performance of this model across all tasks highlights its advanced design philosophy.
[0129] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for fault diagnosis of rolling bearings based on the PGStarNet-MFE model, characterized in that, Includes the following steps: S1: Acquire raw signal Collect one-dimensional vibration acceleration signals of rolling bearings under different fault conditions; S2: Building the dataset The one-dimensional vibration acceleration signal is converted into a two-dimensional time-frequency graph by continuous wavelet transform. The two-dimensional time-frequency graph is used to construct a rolling bearing fault diagnosis dataset, which is then divided into a training set and a test set. S3: Network Construction Using StarNet as the baseline model, a multi-scale feature extraction module is embedded into StarNet, and a linear weighted loss function designed based on the improved large cane rat optimization algorithm is integrated into the classification layer of StarNet to obtain the PGStarNet-MFE network. S4: Model Training The PGStarNet-MFE network is trained using the training set, and the optimal weights obtained during the training process are saved as the final model parameters, thereby obtaining the rolling bearing fault diagnosis model, namely the PGStarNet-MFE model. S5: Fault Diagnosis The test set data is input into the rolling bearing fault diagnosis model to perform end-to-end fault diagnosis and output the fault diagnosis results.
2. The rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 1, characterized in that, In step S1, the different fault state types of rolling bearings include healthy, inner ring fault, outer ring fault, rolling element fault, and mixed faults involving the inner ring, outer ring, and rolling elements.
3. The rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 1, characterized in that, In step S3, the multi-head feature extraction module is embedded into the stem layer of the StarNet network for multi-scale feature extraction of the two-dimensional time-frequency map.
4. The rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 3, characterized in that, In step S3, the PGStarNet-MFE network extracts features from the two-dimensional time-frequency map through a multi-scale feature extraction module. The multi-scale feature extraction module includes four independent branches: a first convolution branch, a second convolution branch, a third convolution branch, and a global average pooling branch. In the first convolution branch, small-scale output features are extracted through a 1×1 convolutional layer. In the second convolution branch, medium-scale output features are extracted through a 3×3 convolutional layer. In the third convolution branch, large-scale output features are extracted through a 3×3 convolutional layer. In the global average pooling branch, the global context information is combined with the local convolution operation by calculating the global average value of each channel in the feature map to obtain the output features. The output features of each branch are concatenated along the channel dimension to form fused features, which are then transformed through a 4×4 convolutional layer to obtain multi-scale features.
5. The rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 4, characterized in that, In step S3, the linear weighted loss function L LF Specifically as follows: THE LF =L CE +λL t Among them, L CE Let L be the cross-entropy loss function. t Let λ be the triplet loss function, and λ be the dynamic weight coefficient. The optimization is performed iteratively by improving the Big Cane Mouse optimization algorithm.
6. The rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 5, characterized in that, Cross-entropy loss function L CE The expression is as follows: Where N is the number of samples, y im y′ is the true label of the Mth class of sample i. im It is the predicted probability of the Mth class of sample i; Triple loss function L t The expression is as follows: L t =max(0,d(a,p)-d(a,n)+ε) Where d(a,p) and d(a,n) represent the Euclidean distances between sample a and positive sample p and negative sample n, respectively, and ε is the margin parameter.
7. The rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 5, characterized in that, The improved Big Cane Rat optimization algorithm incorporates Piecewise chaotic mapping into the original Big Cane Rat optimization algorithm, thereby generating an ergodic chaotic sequence from the initial population through Piecewise chaotic mapping. The specific process of optimizing the dynamic weight coefficient λ is as follows: S31: Initialization Phase Randomly generate a population X of large cane rats. i,j It is the random position of the i-th large cane rat in the j-th dimension of the population, as shown in the following formula: X i,j =X k+1 ×(UB j -LB j )+LB j Among them, X k ∈[0,1], its initial value X is a random number that follows a uniform distribution; q is a parameter that controls the mapping effect, and UB and LB are the upper and lower bounds, respectively, used to limit the location range of the giant cane rat population; Calculate the fitness value of the global giant cane rat and search the spatial boundary; S32: Determine if the maximum number of iterations has been reached. If the maximum number of iterations has been reached, output the position of the large cane rat with the highest fitness value as the dynamic weight coefficient λ. If the maximum number of iterations has not been reached, evaluate the food abundance and generate a random number. Determine if the random number is less than the set stage switching parameter ρ. Enter the corresponding exploration or development stage to update the position of the large cane rat, obtain the new position of the large cane rat, and recalculate the fitness value of each large cane rat. Obtain the position of the large cane rat with the highest fitness value and output it as the dynamic weight coefficient λ. The position of the large cane rat with the highest fitness value is the optimal position of the large cane rat.
8. The rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 7, characterized in that, In step S32, the formula for calculating the fitness value of the optimal cane rat is as follows: Where X represents all the large cane rats in the current population, and f() is the objective function; In step S32, the formula for calculating the stage switching parameter ρ is as follows: Among them, C iter It is the current iteration number, Max iter It represents the maximum number of iterations.
9. A rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 8, characterized in that, In step S32, during the exploration phase, the formula for calculating the new location of the giant cane rat is as follows: Among them, X i,j new Let X represent the new position of the i-th large cane rat in the j-th dimension. k,j Indicates the position of the optimal individual; During the development phase, the formula for calculating the new location of the giant cane rat is as follows: X i,j new =X i,j +C×(X k,j -r×X i,j ) Where C represents a random number defined within the spatial boundary; r is used to simulate the reinforcing effect of abundant food sources on foraging behavior.
10. A rolling bearing fault diagnosis method based on the PGStarNet-MFE model according to claim 9, characterized in that, In step S32, during the development phase, if the fitness value of any large cane rat exceeds that of the current best individual, the optimal position is updated and other individuals are guided to migrate; otherwise, a movement strategy is adopted to deviate from the optimal position. The specific movement strategy is as follows: Among them, F i new F is the fitness value of the optimal large cane rat. i X is the current fitness value, α is the coefficient for the reduction of food sources; β is the coefficient that prompts the algorithm to migrate to other high-value areas during the development phase, and the high-value area is the food-rich area where the optimal cane rat is located; i,j Indicates the current position of the giant cane rat, X. k,j It is the optimal position of the large cane rat in the j-th dimension.