Motor bearing fault diagnosis method and system for time sequence monitoring data loss

CN122839083APending Publication Date: 2026-09-29TAIHANG LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611329263.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-31
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0010]本发明的目的在于针对现有电机轴承故障诊断技术在时序监测数据缺失条件下存在的诊断精度低且随缺失率骤降、数据补全易丢失故障关键时序特征、对缺失模式适配性差、轻量化与高精度难以兼顾等技术问题,提供一种面向时序监测数据缺失的电机轴承故障诊断方法及系统

Benefits of technology

1.模型适配性强,可覆盖不同缺失模式:本发明通过教师-学生双轨模型联动训练机制,结合离线知识迁移与多态特征对比学习双重约束,让学生模型在训练中充分习得不同缺失率、不同缺失模式下的特征对齐规律。无需为特定缺失场景单独训练模型,一次训练即可稳定适配不同缺失率及随机间歇性缺失、突发式连续缺失、跨模态协同缺失等工业现场常见模式,解决现有技术缺失模式变化即精度暴跌的痛点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839083A_ABST
    Figure CN122839083A_ABST
Patent Text Reader

Abstract

The application discloses a motor bearing fault diagnosis method and system for time series monitoring data loss, and belongs to the technical field of motor fault diagnosis. The application constructs a missing mask matrix corresponding to the time dimension and the feature dimension of time series data, and accurately marks the missing position and the missing rate; based on complete data, a multi-state missing sample cluster covering different missing rates and various missing modes is generated through random masking; a teacher-student dual-track model linkage training framework is constructed, semantic alignment of missing data and complete data is realized through triple constraints of hidden feature migration loss, performance migration loss and feature comparison learning loss; after training, the student model is executed for structured pruning optimization, redundant channels are removed, and the recovery accuracy is fine-tuned. The application does not need data completion, can adapt to any missing rate and various missing modes through one-time training, the pruned model is lightweight, the fault diagnosis accuracy is high, and can be deployed on industrial edge devices to realize real-time diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of motor fault diagnosis and industrial time-series data processing technology, specifically to a method and system for diagnosing motor bearing faults due to missing time-series monitoring data. It is applicable to online monitoring and fault diagnosis scenarios for industrial motor bearings and can be deployed on edge computing devices, industrial gateways, and cloud monitoring platforms. Background Technology

[0002] As a core transmission component of industrial motors, the operating condition of motor bearings directly determines the motor's reliability, operational stability, and service life. If early bearing failures are not diagnosed and addressed promptly, they can easily lead to motor jamming, shutdown, or even complete machine damage, causing serious industrial production losses and safety hazards. Therefore, real-time and accurate fault diagnosis of motor bearings is one of the core aspects of intelligent operation and maintenance of industrial equipment.

[0003] With the development of industrial sensor technology and the Internet of Things (IoT) technology, collecting time-series monitoring data of motor bearings using sensors such as acceleration, temperature, and current, and achieving end-to-end fault diagnosis based on deep learning technology, has become the mainstream technical direction for motor bearing fault diagnosis. Compared with traditional signal processing methods, deep learning-based diagnostic methods can automatically extract fault features from time-series monitoring data without the need for manually designed feature extractors. This allows for a more comprehensive and accurate capture of the time-series evolution patterns and feature correlations of bearing faults, and has been widely researched and applied in industrial settings.

[0004] However, in actual industrial applications, the timing monitoring data of motor bearings inevitably suffers from missing information due to multiple factors, including sensor performance, the on-site working environment, and data transmission. The missing timing monitoring data manifests primarily as: random, intermittent missing data caused by poor sensor contact, power supply interruptions, or hardware failures; sudden, continuous missing data caused by network fluctuations or transmission interruptions in the industrial environment; and condition-dependent missing data caused by harsh operating conditions such as high motor loads and strong vibrations. These missing data patterns are irregular, the data loss rate can vary randomly over time, and the location and duration of the missing data are uncertain. This is a common problem in industrial settings that cannot be completely resolved through simple hardware upgrades or maintenance.

[0005] To address the issue of missing time-series monitoring data, various corresponding solutions for motor bearing fault diagnosis have emerged in existing technologies. However, all of them have significant limitations and cannot adapt to the missing conditions in industrial settings. Specifically, these shortcomings are reflected in the following three aspects:

[0006] Traditional data completion methods have poor adaptability and are prone to losing key fault features. Existing technologies often use traditional methods such as interpolation, Gaussian imputation, and matrix completion to complete missing time-series monitoring data before inputting the completed data into the diagnostic model. These methods are only suitable for static scenarios with low missing rates and lack effective adaptability to time-varying missing patterns. Furthermore, the completion process is mostly based on the statistical regularity of the data and does not take into account the temporal characteristics of motor bearing faults. This makes it very easy to distort or lose key features of bearing faults, such as pulse signals, frequency band drift, and amplitude abrupt changes caused by the fault. The completed data can lead to deviations in subsequent diagnostic models and even cause misdiagnosis or missed diagnosis.

[0007] Existing deep learning diagnostic methods require separate training, resulting in high deployment and maintenance costs. Some deep learning diagnostic methods design and train models separately for different missing rates and missing patterns, which can improve diagnostic accuracy under specific missing conditions to some extent. However, this approach requires building multiple diagnostic models, which not only increases the training cost and computing power consumption of the models, but also significantly increases the difficulty of model deployment, updating, and maintenance in industrial settings. At the same time, the missing patterns in industrial settings are unpredictable, making it impossible to achieve full coverage of all missing conditions, thus limiting the actual diagnostic effectiveness.

[0008] Existing missing data adaptive diagnostic methods lack robustness, with accuracy plummeting as the missing data rate increases. A few techniques attempt to incorporate missing data adaptation mechanisms into diagnostic models, but these are mostly designed for static missing data scenarios, failing to fully consider the temporal correlation and time-varying nature of missing data in motor bearing time-series monitoring data. They also lack semantic alignment mechanisms between missing and complete data, resulting in highly sensitive feature representations learned by the model to the missing data rate. When the missing data rate in industrial settings increases, diagnostic accuracy drops drastically, leading to the problem of "diagnostic failure due to missing data," failing to meet the stable diagnostic requirements under missing data conditions in industrial settings.

[0009] In summary, current technical solutions for motor bearing fault diagnosis have not effectively addressed the problems of low diagnostic accuracy, poor robustness, and high deployment costs caused by missing time-series monitoring data. There is an urgent need for a motor bearing fault diagnosis method and system that can adapt to missing scenarios, does not require separate model training for different missing modes, and can retain key time-series characteristics of bearing faults during missing data processing. This would improve the real-time performance, accuracy, and robustness of motor bearing fault diagnosis in industrial settings and meet the actual needs of intelligent operation and maintenance of industrial equipment. Summary of the Invention

[0010] The purpose of this invention is to address the technical problems of existing motor bearing fault diagnosis technologies under conditions of missing time-series monitoring data, such as low diagnostic accuracy that drops sharply with the missing rate, easy loss of key temporal features of faults during data completion, poor adaptability to missing patterns, and difficulty in balancing lightweight design and high accuracy. This invention provides a method and system for motor bearing fault diagnosis under conditions of missing time-series monitoring data. By employing a design combining a heterogeneous dual-track semantic alignment learning mechanism and post-pruning optimization, accurate semantic alignment between missing and complete data is achieved, allowing for adaptation to arbitrary missing rates and multiple missing patterns with a single training iteration. Simultaneously, a dedicated lightweight pruning stage further compresses the student model, doubly ensuring its adaptability to the real-time inference requirements of industrial edge computing devices and industrial gateways, minimizing deployment and application costs while maintaining diagnostic accuracy.

[0011] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for diagnosing motor bearing faults due to missing timing monitoring data, comprising the following steps: Step 1: Preprocessing of time-series monitoring data and construction of missing mask matrix.

[0012] This stage forms the foundation of the entire process and is also compatible with offline training and online diagnostic stages. Its core function is to generate a mask matrix that accurately represents the missing data state.

[0013] During offline training: Multi-dimensional time-series monitoring data of motor bearings in industrial sites, including vibration, temperature, and current, are collected and acquired. Preprocessing operations such as noise reduction, standardization, and time dimension alignment are then performed on the time-series monitoring data to eliminate interference from environmental noise and sensor differences. The noise reduction process employs an improved wavelet thresholding method, using wavelet coefficient thresholds... Wavelet coefficients larger than the threshold are subjected to soft thresholding, while wavelet coefficients smaller than the threshold are set to zero. , For a single sample time step, The data consists of time-series monitoring data of any dimension. The standardization process employs Z-score transformation applied to each dimension of the data to ensure that the mean of each dimension is 0 and the standard deviation is 1.

[0014] A missing data mask matrix is ​​constructed based on data acquisition status markers to accurately mark the missing locations and missing rates of data at each time point and in each dimension. At the same time, the missing patterns of the complete data are expanded by random masking to supplement samples with different missing rates and missing patterns, ensuring that the mask matrix covers common patterns in industrial settings such as random intermittent missing data, sudden continuous missing data, and conditionally dependent missing data.

[0015] During online diagnostics: For single-sample time-series data with or without missing data collected in real time from the industrial site, the same preprocessing operations as in the training phase are performed. Based on the real-time acquisition status, a real-time single-sample missing mask matrix is ​​constructed, marking the missing information of each sample one-to-one, providing an accurate basis for subsequent feature extraction. Each missing mask matrix corresponds one-to-one with the time dimension and feature dimension of the time-series data, ensuring the accurate transmission of missing information.

[0016] Step 2: Construction of multimodal missing sample clusters.

[0017] This stage mainly involves offline training, the core of which is to build a sample cluster covering the full missing rate to provide positive / negative sample support for feature comparison learning. This stage is not required in the online diagnostic stage.

[0018] Based on the preprocessed complete time-series monitoring data and the missing mask matrix, the complete data within the same time window are masked with different missing rates using a random masking method to generate polymorphic missing sample clusters. The masking with different missing rates covers missing rates of 20%, 40%, 60%, and 80%, encompassing random intermittent missing patterns, bursty continuous missing patterns, and conditionally dependent missing patterns.

[0019] Incomplete data with different missing rates within the same time window are defined as positively correlated sample pairs, while incomplete data from different time windows are considered as negatively dissimilar sample pairs. That is, the positively correlated sample pairs are defined as samples with different missing rates within the same time window that are positively correlated with each other, and the negatively dissimilar sample pairs are defined as any missing samples from different time windows that are negatively correlated with each other. At the same time, complete data samples are retained as semantic alignment benchmarks, ultimately forming a polymorphic missing sample cluster that includes single-modal missing samples, cross-dimensional fused missing samples, and complete data samples, providing a sample foundation for subsequent dual-track model linkage training.

[0020] Step 3: Construct a teacher-student dual-track model linkage training framework.

[0021] This stage is the core component. Through the design of a heterogeneous dual-track semantic alignment learning mechanism, the student model learns the ability to adapt to missing information under the guidance of the teacher model. After training, only the student model is retained for subsequent optimization and online diagnosis.

[0022] The heterogeneous backbone network configuration adopts a heterogeneous design of a high-precision complex teacher model architecture and a lightweight student model architecture, which takes into account both feature extraction accuracy and edge deployment compatibility.

[0023] (1) The teacher model employs a complex, high-precision temporal backbone network with deeper layers and wider channels. It focuses on extracting the complete fault semantic feature space of different health states of motor bearings from complete data and is used only for semantic benchmark construction during the offline training phase. The teacher model uses a deep temporal convolutional network with an input layer, four convolutional layers, and a fully connected layer. The convolutional kernel size is 3×1, the stride is 1, the padding is 1, the number of channels is 8→256→512→512→256, the dropout is 0.1, and the activation function is ReLU. The training process uses complete data samples as input. After training, all parameters are fixed as a benchmark for semantic alignment.

[0024] (2) The student model adopts a lightweight temporal backbone network with shallow layers and narrow channels, possessing low parameter and low computational cost characteristics, laying the foundation for edge deployment. The student model adopts a lightweight temporal convolutional network, the network structure of which includes an input layer, three depthwise separable convolutional layers and a fully connected layer, with a convolutional kernel size of 3×1, stride of 1, padding=1, number of channels from 8→64→128→64, number of groups = number of channels, dropout=0.1, and activation function of GELU. The training process takes the polymorphic missing sample cluster as input, and uses the complete feature representation and diagnostic results output by the teacher model as transferable knowledge. Through hidden feature transfer loss and performance transfer loss, the difference between the output of the student model and the output of the teacher model is constrained (such as mean squared error loss), ensuring that the features extracted by the student model from the missing data are semantically consistent with the features of the complete data.

[0025] (3) Polymorphic feature contrast learning constraint: For the feature output of the student model for samples with different missing rates, calculate the feature similarity of positively correlated sample pairs and the feature difference of negatively correlated sample pairs. Through feature contrast loss, constrain the feature output of the student model for samples with different missing rates to achieve semantic alignment, so that the model can adapt to any missing rate without separate training.

[0026] (4) The total loss function for training the student model includes model prediction loss, hidden feature transfer loss, performance transfer loss, and feature contrast loss, avoiding the information forgetting problem in multi-stage training. The formula is as follows:

[0027] in, To predict losses, characterize the difference between the student model's diagnostic results and the true labels, and adapt to the 5-class classification task of motor bearings (normal, inner ring fault, outer ring fault, rolling element fault, cage fault). To hide the feature transfer loss, we characterize the differences between the hidden layer features of the student model and the hidden layer features of the teacher model; The performance transfer loss characterizes the difference between the diagnostic output of the student model and the diagnostic output of the teacher model. Feature contrast loss is used to characterize the contrast loss of features between samples with different missing rates. , , , These are the weighting coefficients for each loss.

[0028] Furthermore, the feature contrast loss The calculation formula is:

[0029] in, This represents the number of samples within the batch. For positive feature output of samples, For the first The feature vector of each sample is generated by the student model. For the first batch The feature vector of each sample For cosine similarity, This refers to the temperature parameter.

[0030] Step 4: Structured pruning and optimization of the student model.

[0031] This stage is a lightweight optimization phase for offline training. The core is to further compress the student model size and reduce the amount of computation while ensuring diagnostic accuracy, and adapt to the computing power constraints of edge devices.

[0032] First, based on the parameters after the student model has been trained, the importance score of each channel in the convolutional and fully connected layers is calculated. An L1 regularization combined with gradient contribution weighting is used to calculate the importance score of each channel. This method simultaneously considers the magnitude of the channel weights and the channel's contribution to the model loss, avoiding the accidental deletion of core feature channels. The channel importance score... The calculation formula is:

[0033] in, , To balance the weighting coefficients of the two scores, The number of convolution kernel weights, Indicates the first Channel parameters, Indicates the first d The first in the passage Each convolutional kernel weight, This is the gradient of the channel parameter with respect to the total loss. This represents the maximum gradient across all channels. The importance score for each channel. The higher the value, the greater the contribution of the channel to fault feature extraction.

[0034] All channels are scored by importance. The channels are sorted in descending order, and redundant channels ranked lower in importance are removed according to a preset pruning rate. Only the remaining core feature channels are retained, while the overall structure of the model is preserved, thus avoiding model failure caused by pruning.

[0035] After pruning, some parameters of the model will change, which may lead to a slight decrease in diagnostic accuracy. Therefore, it is necessary to perform a light fine-tuning on the pruned model and continue training it at a learning rate lower than that used in the training phase to avoid drastic fluctuations in parameters. The fine-tuning data should be core samples from the sample cluster. The fine-tuning can help recover the accuracy loss that may have been caused by pruning and ensure the stability of the diagnostic performance of the pruned model.

[0036] Step 5: Fault diagnosis reasoning under missing conditions.

[0037] This stage does not require constructing polymorphic missing sample clusters. The core is to directly process the target test samples using a pruned and optimized lightweight student model. Real-time collected multi-dimensional time-series monitoring data of the motor bearing, after preprocessing, is input into the pruned and optimized student model along with the corresponding real-time missing mask matrix. Based on the missing data adaptation capability learned during training, the model directly extracts robust fault semantic features from the missing data without any completion operations. Through internal feature adaptive adjustment, the interference of missing data on semantic expression can be offset. Then, using robust fused features as input, the lightweight student model outputs the diagnostic results of the motor bearing.

[0038] The student model uses Softmax classification logic to transform the extracted fault semantic features into probability distributions for five states. The diagnostic results include classification and identification of five states: normal state, inner race fault, outer race fault, rolling element fault, and cage fault. The entire fault diagnosis reasoning process is completed end-to-end without the need for a teacher model, and the inference latency of the ultra-lightweight student model is significantly reduced, meeting the real-time requirements of industrial online monitoring.

[0039] Furthermore, it also includes a reasoning reliability verification step: verifying the missing rate of the current sample, and marking it as low reliability if the missing rate is greater than a preset threshold (e.g., 70%); verifying the cosine similarity between the fused features of the current sample and the feature template of the same type of fault sample, and judging it as medium reliability if the similarity is less than a preset threshold (e.g., 0.7); verifying the consistency of the diagnostic results of multiple consecutive time windows (e.g., 3), and judging it as high reliability if the results are consistent.

[0040] Secondly, the present invention provides a motor bearing fault diagnosis system for cases of missing timing monitoring data, used to implement the above method, comprising: The data acquisition module is used to acquire multi-dimensional time-series monitoring data of the motor bearing. The data acquisition module includes a vibration sensor, a torque sensor, and corresponding signal acquisition circuitry, used to acquire vibration and torque signals of the motor bearing in real time during operation.

[0041] The preprocessing module performs preprocessing operations on the time-series monitoring data and constructs a missing mask matrix based on the data acquisition status markers. The preprocessing operations include noise reduction and normalization, and the missing mask matrix corresponds one-to-one with the time dimension and feature dimension of the time-series monitoring data.

[0042] The training module is used to generate polymorphic missing sample clusters based on the preprocessed complete time-series monitoring data and the missing mask matrix, and to train the student model based on a teacher-student dual-track model linkage training framework. The training module includes a teacher model training submodule, a student model training submodule, and a contrastive learning constraint submodule.

[0043] The pruning optimization module is used to perform structured pruning optimization on the student model after training. It calculates the importance score of each channel in the convolutional layer and fully connected layer in the model, sorts each channel in descending order according to the importance score, removes the channels ranked lower according to the preset pruning rate, retains the remaining channels, and continues to train the pruned model with a learning rate lower than the learning rate in the training phase.

[0044] The diagnostic inference module is used to input the real-time collected time-series monitoring data along with the corresponding real-time mask matrix into the pruned and optimized student model, extract fault semantic features, and output diagnostic results. The diagnostic inference module also includes a reliability verification submodule, used to verify the diagnostic results for missing rate, feature consistency, and inference stability.

[0045] The modules mentioned above are connected in sequence. The data stream flows from the data acquisition module through the preprocessing module, the training module (offline stage) or the diagnostic reasoning module (online stage), and the pruning and optimization module (offline stage), and finally the diagnostic reasoning module outputs the diagnostic results.

[0046] Compared with the prior art, the present invention has at least the following beneficial effects: 1. Strong model adaptability, covering different missing data patterns: This invention employs a teacher-student dual-track model linkage training mechanism, combining offline knowledge transfer and polymorphic feature comparison learning with dual constraints. This allows the student model to fully learn feature alignment rules under different missing rates and missing data patterns during training. There is no need to train a separate model for specific missing data scenarios; a single training session can stably adapt to common industrial scenarios such as different missing rates, random intermittent missing data, sudden continuous missing data, and cross-modal collaborative missing data, solving the pain point of existing technologies where accuracy drops drastically when missing data patterns change.

[0047] 2. High Model Diagnostic Accuracy: This invention leverages the strong feature extraction capabilities of a high-precision teacher model to mine more comprehensive and nuanced semantic features of faults from complete data, providing a more reliable and accurate semantic benchmark for the student model. Through dual transfer learning of hidden features and diagnostic results, the student model achieves high-precision feature representation consistent with the teacher model. Compared to traditional lightweight models that sacrifice accuracy for efficiency, this invention maintains excellent diagnostic accuracy while remaining lightweight.

[0048] 3. Outstanding Lightweight Performance, Adaptable to Edge Deployment: This invention employs a dual lightweight design of a lightweight student model and post-optimized structured pruning, ensuring edge compatibility from both the source and subsequent optimization perspectives. The student model adopts a shallow, narrow-channel lightweight architecture, inherently possessing low parameter count and low computational cost. The post-pruning stage, through channel importance assessment and structured pruning, eliminates redundant channels without sacrificing core accuracy, further reducing the model's parameter count and computational cost, ultimately forming an ultra-lightweight student model that significantly reduces deployment costs and computing power consumption in industrial settings. Attached Figure Description

[0049] Figure 1 This is an overall flowchart of the motor bearing fault diagnosis method for missing timing monitoring data according to an embodiment of the present invention; Figure 2 This is a diagram illustrating the loss relationship in the teacher-student dual-track model linkage training in an embodiment of the present invention. Detailed Implementation

[0050] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0051] The present invention will be described below through specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0052] Example 1 like Figure 1 , Figure 2 As shown, this embodiment provides a method for diagnosing motor bearing faults due to missing timing monitoring data, including the following steps: Step 1: Preprocessing of time-series monitoring data and construction of missing mask matrix The input data used in this embodiment were all collected from a motor bearing fault simulation test bench. Based on the core requirements of motor bearing fault diagnosis, seven sets of vibration signals and one set of torque signals were selected as core monitoring parameters to reflect the operating status of the motor bearing. The input data is 8-dimensional time-series monitoring data. The seven sets of vibration signals (denoted as V1-V7) were collected from different monitoring points of the motor bearing to capture abnormal vibrations during bearing operation. The one set of torque signals (denoted as T) was used to monitor changes in the motor output torque, reflecting the impact of motor load changes on bearing faults. The two types of signals work together to achieve comprehensive monitoring of the bearing's health status.

[0053] Considering the frequency characteristics of motor bearing fault pulse signals and the practicality of industrial field data acquisition, the sampling frequency f is set. s =12800Hz, which can fully capture the detailed features of the fault pulse and avoid the loss of fault features due to the low sampling frequency. The time length of a single sample is set to 0.04 seconds, corresponding to 512 time steps. The calculation formula is: 512 = 12800Hz × 0.04s. This time length can ensure that each sample contains the complete fault pulse cycle, providing sufficient data support for subsequent fault feature extraction.

[0054] The original 8-dimensional time-series monitoring data may contain spike noise and outliers caused by sensor malfunctions, power grid interference, etc. Such outliers can interfere with the accuracy of subsequent feature extraction and model training; therefore, data cleaning is necessary first. This embodiment uses... Criteria for removing outliers, applicable to time-series monitoring data of any dimension. If its value satisfies If it is an outlier, it is considered an outlier and the mean of the data in that dimension is used. Replacement is used to eliminate the impact of outliers on subsequent processes and ensure data reliability. After data cleaning, further noise reduction and standardization processing is performed on the cleaned data.

[0055] Because the original time-series monitoring data contains a large amount of environmental noise and the dimensions of the data are inconsistent, problems such as gradient vanishing and feature weight imbalance may occur during model training. Therefore, it is necessary to preprocess the acquired multi-dimensional time-series monitoring data. The specific preprocessing operations are as follows: Noise Reduction: This embodiment employs an improved wavelet thresholding noise reduction method. Compared to traditional wavelet thresholding noise reduction, this method effectively preserves key features such as fault pulses while suppressing noise interference. The db4 wavelet is selected as the base wavelet, and the decomposition level is set to 4 levels. The wavelet coefficient threshold calculation formula is as follows:

[0056] in, For a single sample time step, This is a noise estimate. For time-series monitoring data of any dimension, calculating using the median method avoids the influence of extreme values ​​on noise estimation and ensures the rationality of the threshold setting. The wavelet coefficient processing rules are as follows:

[0057] in, These are wavelet coefficients. These are wavelet coefficients after soft thresholding or being set to zero. This rule achieves the dual goals of noise suppression and fault feature preservation by performing soft thresholding on wavelet coefficients greater than the threshold, retaining the corresponding wavelet coefficients, and setting wavelet coefficients less than the threshold to zero.

[0058] Standardization: To eliminate the differences in the dimensions of the data and ensure the weight balance of each feature dimension during model training, Z-score transformation is performed on the 8-dimensional data. The specific formula is as follows:

[0059] in, For the first Time step number Dimensional data ( =1-7 corresponds to V1-V7, =8 corresponds to T). For the first Time step number Dimensional standardized data, , The first This transformation sets the mean and standard deviation of the data in each dimension to 0 and 1, unifying the data units and providing standardized input data for subsequent model training. The missing mask matrix is ​​constructed as follows: To accurately characterize the missing data state and facilitate the model's subsequent learning of the missing data's features, this embodiment constructs a binary missing data mask matrix. Furthermore, based on real-world missing data patterns in industrial settings, the missing data samples are expanded to ensure the model can adapt to various missing data scenarios. An 8-dimensional binary matrix Mask∈{0,1} is used. 5 ¹² x8Mask[i,j]=1 indicates that the j-th dimension of data at time step i is valid, and Mask[i,j]=0 indicates that the data in that dimension is missing at that time point. The missing mask matrix corresponds one-to-one with the time dimension and feature dimension of the time series monitoring data, which can accurately convey the missing location and missing rate information of each dimension of data, providing a basis for subsequent feature extraction and fusion.

[0060] Offline missing sample expansion: To ensure the model can adapt to various missing patterns in industrial settings, samples with different missing rates are generated using random masking based on preprocessed complete 8-dimensional time-series monitoring data. The missing rates cover 20%, 40%, 60%, and 80%, encompassing three common missing patterns in industrial settings: random intermittent missing (mask positions are randomly distributed, simulating missing due to poor sensor contact), sudden continuous missing (data missing at 20-50 consecutive time points, simulating missing due to network transmission interruption), and condition-dependent missing (when the torque signal is greater than 80 N·m, the missing rate of the vibration signal increases by 30%, simulating missing due to sensor failure under high load conditions). This expansion method ensures that the sample cluster can fully cover real-world missing scenarios in industry, improving the robustness of the model.

[0061] Step 2: Construction of multimodal missing sample clusters The core purpose of constructing multi-modal missing sample clusters during the offline training phase is to provide high-quality training samples for the teacher-student dual-track model training. By constructing positive correlation sample pairs and negative difference sample pairs, it supports polymorphic feature comparison learning and ensures that the student model can learn the feature alignment rules under different missing rates and missing modes. The specific implementation process is as follows: Time window division: The time window is divided into units of 512 time steps. The length of the window is consistent with the length of a single sample time. This ensures that each window contains complete fault pulse features and avoids fault feature breakage caused by too short a window or feature redundancy caused by too long a window.

[0062] Sample pair construction: Positive association sample pairs are defined as samples with missing 8-dimensional data rates of 20%, 40%, 60%, and 80% respectively within the same time window. That is, for time window T, positive association samples are considered to be paired with each other. k Sample S k (20%), S k (40%), S k (60%), S k (80%) are positive pairs, originating from the same real-world scenario and sharing consistent semantic features. These can be used to guide student models in learning feature alignment under different missing rates. Negative difference pairs are defined as samples from different time windows (e.g., T). k With T mAny missing samples (k ≠ m) are negative pairs of each other. These samples originate from different operating scenarios and have significant differences in semantic features. They can be used to constrain the student model to distinguish the feature differences of different scenarios.

[0063] Step 3: Construct a teacher-student dual-track model linkage training framework The teacher-student dual-track model collaborative training is based on a heterogeneous dual-backbone network design. The teacher model provides accurate semantic benchmarks, and combined with the dual constraints of knowledge transfer and polymorphic feature contrastive learning, it guides the student model to learn adaptive capabilities to missing information, while taking into account both diagnostic accuracy and lightweight requirements. The specific implementation is as follows: (1) Design of heterogeneous backbone network A heterogeneous design is adopted, combining a high-precision, complex teacher model with a lightweight student model. The teacher model focuses on extracting accurate fault semantic features, providing a reliable semantic benchmark for the student model; the student model focuses on lightweight design, laying the foundation for edge deployment. The network structures and core parameters of both are shown in Table 1 below: Table 1. Comparison of network structure and core parameters between the teacher model and the student model.

[0064] (2) Loss function and training parameters like Figure 2 As shown, to ensure that the student model can achieve semantic alignment with the teacher model and adapt to missing scenarios, a total loss function is designed that includes prediction loss, hidden feature transfer loss, performance transfer loss and feature comparison loss. By reasonably setting the weights of each loss, the accuracy and convergence speed of the model are balanced.

[0065] The formula for the total loss function is as follows:

[0066] Specific calculations for each loss: The predicted loss is adapted for a 5-class classification task of motor bearings (normal, inner race fault, outer race fault, rolling element fault, cage fault) to constrain the consistency between the diagnostic results of the student model and the true labels. The calculation formula is as follows:

[0067] in, This represents the batch sample size. For unique hot tags (if the first The sample belongs to the first Class of faults, then =1, otherwise 0). Predicting the first student model The sample belongs to the first The probability of a class.

[0068] The hidden feature transfer loss is used to ensure that the hidden features extracted by the student model are consistent with those extracted by the teacher model, thereby achieving knowledge transfer at the feature level. The calculation formula is as follows:

[0069] in, This represents the batch sample size. For the hidden layer feature dimension, For the teacher model The first sample 3D hidden features, For the student model The first sample Hidden features.

[0070] Performance transfer loss is used to ensure that the diagnostic output of the student model remains consistent with that of the teacher model, thereby achieving knowledge transfer at the diagnostic result level. The calculation formula is as follows:

[0071] in, For the teacher model The first sample Diagnostic output of the class, For the student model The first sample The diagnostic output of the class.

[0072] Feature contrast loss is used to constrain the feature output of the student model for samples with different missing rates, achieving semantic alignment between samples with different missing rates. The calculation formula is as follows:

[0073] in, This represents the number of samples within the batch. For positive feature output of samples, For the first The feature vector of each sample is generated by the student model. For the first batch The feature vector of each sample Cosine similarity is used to calculate the degree of similarity between the features of two samples. Its formula is: , For temperature parameters, =0.07, used to adjust the similarity weights to avoid gradient vanishing.

[0074] Training parameters: The Adam optimizer was used, with β1=0.9, β2=0.999, and weight decay of 1e-5. The initial learning rate was 1e-4. A cosine annealing scheduling strategy was adopted, and the learning rate was gradually reduced as the number of training epochs increased to improve the model's convergence accuracy. The total number of training epochs was 60. An early stopping strategy was also set. If the accuracy of the validation set did not improve for 5 consecutive epochs, training was stopped to avoid model overfitting.

[0075] Step 4: Structured Pruning and Optimization of Student Model Structured pruning optimization of the student model is a key step in achieving edge deployment. Its core objective is to further reduce the number of parameters and computational cost of the student model, eliminate redundant channels, and improve the inference speed of the model while ensuring that the diagnostic accuracy remains basically unchanged. The specific implementation process is as follows: (1) Channel importance assessment: The importance score of each channel is calculated by combining L1 regularization with gradient contribution weighting. This method can simultaneously consider the magnitude of the channel weight and the contribution of the channel to the model loss, avoiding the accidental deletion of core feature channels. The calculation formula is as follows:

[0076] in, , To balance the weighting coefficients of the two scores, The number of convolution kernel weights, Indicates the first Channel parameters, Indicates the first d The first in the passage Each convolutional kernel weight, This is the gradient of the channel parameter with respect to the total loss. This represents the maximum gradient across all channels. The importance score for each channel. The higher the value, the greater the contribution of the channel to fault feature extraction.

[0077] (2) Structured pruning: all channels are scored according to their importance. The channels are sorted in descending order, and redundant channels with lower importance are removed according to the preset pruning rate, leaving only the remaining core feature channels. At the same time, the overall structure of the model is preserved, avoiding model failure caused by pruning.

[0078] (3) Fine-tuning after pruning: Some parameters of the model will change after pruning, which may lead to a slight decrease in diagnostic accuracy. Therefore, it is necessary to perform light fine-tuning on the pruned model. The fine-tuning learning rate is lower than the learning rate during the training phase to avoid drastic fluctuations in parameters. The fine-tuning data is selected from the core samples in the sample cluster. The fine-tuning restores the accuracy loss that may be caused by pruning and ensures the stability of the diagnostic performance of the pruned model.

[0079] Step 5: Fault Diagnosis Reasoning under Missing Conditions The purpose of fault diagnosis reasoning under missing conditions is to directly process potentially missing time-series monitoring data collected in real time at industrial sites using a pruned and optimized lightweight student model, extract robust fault semantic features, and achieve accurate and rapid diagnosis of motor bearing faults, avoiding diagnostic performance distortion caused by missing data. The specific implementation process is as follows: (1) Feature extraction: The real-time preprocessed time-series monitoring data and the corresponding real-time missing mask matrix are input into the pruned and optimized student model. Based on the missing adaptive capability learned during the training phase, the model compensates for the impact of missing dimensions on semantic expression through an internal feature adaptive adjustment mechanism. At the same time, high-dimensional fault semantic features are extracted. These features can accurately characterize the equipment operating status and have strong robustness to missing data.

[0080] (2) Fault diagnosis reasoning: The standard Softmax classification logic is adopted. The fused features (i.e. the extracted fault semantic features) are input into the fully connected classification layer of the pruned student model. The output of the classification layer is transformed into the probability distribution of five states (normal state, inner ring fault, outer ring fault, rolling element fault, cage fault) through the Softmax function, so as to realize the direct classification and identification of the samples.

[0081] (3) Inference reliability verification: In order to avoid misdiagnosis due to missing data, an inference reliability verification mechanism is designed to verify the reliability from three dimensions: First, the missing rate verification. If the current sample missing rate is >70%, then “low reliability” will be marked when outputting the diagnosis result. Second, feature consistency verification. The cosine similarity between the current sample fusion feature and the feature template of the same type of fault sample will be compared. If the similarity is <0.7, it will be judged as “medium reliability”. Third, inference stability verification. If the sample diagnosis results are consistent for three consecutive time windows, it will be judged as “high reliability”.

[0082] Example 2 This embodiment provides a motor bearing fault diagnosis system for missing time-series monitoring data, used to implement the above method. The system includes a data acquisition module, a preprocessing module, a training module, a pruning optimization module, and a diagnostic inference module, which are connected in sequence and work together.

[0083] The data acquisition module includes vibration sensors, torque sensors, and corresponding signal acquisition circuits to obtain multi-dimensional time-series monitoring data of the motor bearing. Vibration sensors (such as accelerometers) are deployed at different monitoring points on the motor bearing to collect seven sets of vibration signals; torque sensors are used to collect the motor's output torque signal. The analog signals output by each sensor are converted into digital signals by an analog-to-digital converter and include a unified timestamp to ensure time synchronization of the multiple signals.

[0084] The preprocessing module receives the raw multi-dimensional time-series monitoring data output from the data acquisition module and performs preprocessing operations such as noise reduction, standardization, and time dimension alignment. Noise reduction employs an improved wavelet thresholding method, using db4 as the basis wavelet and a 4-level decomposition to denoise the vibration signal. Standardization uses Z-score transformation to ensure that the mean of each dimension of the data is 0 and the standard deviation is 1. The preprocessing module also constructs a binary missing data mask matrix based on the data acquisition status markers, accurately marking the missing locations and missing rates of each time node and each dimension of the data. The missing data mask matrix corresponds one-to-one with the time dimension and feature dimension of the time-series monitoring data.

[0085] The training module operates offline, receiving complete temporal monitoring data and a missing mask matrix from the preprocessing module. First, it generates polymorphic missing sample clusters using random masks, covering missing rates of 20%, 40%, 60%, and 80%, as well as three typical industrial missing patterns. Internally, the training module comprises a teacher model training submodule, a student model training submodule, and a contrastive learning constraint submodule. The teacher model training submodule uses complete data samples as input to train a deep temporal convolutional network and fixes its parameters; the student model training submodule uses polymorphic missing sample clusters as input to train a lightweight temporal convolutional network; and the contrastive learning constraint submodule calculates the feature contrast loss for positively correlated sample pairs and negatively dissimilar sample pairs. These three submodules work collaboratively, using a total loss function... Jointly optimize student model parameters.

[0086] The pruning and optimization module operates after the training module completes, performing structured pruning and optimization on the trained student model. The module first calculates the importance score of each channel using L1 regularization combined with gradient contribution weighting. Then, the channels are sorted in descending order of importance score, and redundant channels are removed according to the preset pruning rate. Finally, the pruned model is trained again with a learning rate lower than that of the training phase, and minor adjustments are made to restore accuracy.

[0087] The diagnostic inference module operates online, receiving real-time collected time-series monitoring data and the corresponding real-time missing mask matrix. It then inputs the pruned and optimized student model, extracts fault semantic features, and outputs diagnostic results. The diagnostic inference module also includes a reliability verification submodule, which evaluates the confidence level of the diagnostic results from three dimensions: missing rate, feature consistency, and inference stability, outputting the fault category, health index, and reliability level.

[0088] The system's workflow is divided into an offline training phase and an online diagnostic phase. In the offline training phase, the data stream sequentially passes through the data acquisition module, preprocessing module, training module, and pruning optimization module to complete the training and lightweight optimization of the student model. In the online diagnostic phase, the data stream sequentially passes through the data acquisition module, preprocessing module, and diagnostic inference module to complete real-time fault diagnosis output.

[0089] Experimental Results and Analysis To verify the effectiveness, adaptability, and engineering practicality of the method of this invention, a comparative experiment was conducted using a dataset constructed from measured data from a motor bearing fault simulation test bench. The experiment focused on evaluating the diagnostic accuracy, model lightweighting index, and identification performance of the method under different missing rates. The experimental results and analysis are as follows: (1) Experimental dataset The dataset contains 8-dimensional time-series monitoring data, including 7 vibration channels and 1 torque channel. It collects data samples under four states: inner ring fault, outer ring fault, rolling element fault, and cage fault, as well as normal state data samples, which can be used to verify the diagnostic performance of the method for different fault types.

[0090] (2) Experimental results The experiment compared and tested the diagnostic performance and the lightweight nature of the model. The specific experimental results are as follows: Table 2 below shows a comparison of the average diagnostic accuracy of different methods at various missing rates: Table 2. Comparison of average diagnostic accuracy (%) of different methods at various missing rates

[0091] Table 3 below compares the lightweight model with real-time performance metrics: Table 3 Comparison of lightweight and real-time performance metrics for each model

[0092] (3) Results Analysis Experimental results demonstrate that this invention possesses significant advantages across multiple dimensions. Designed for scenarios with missing data in multi-time-series monitoring, this invention adapts well to various missing data scenarios. After pruning, the model achieves an average diagnostic accuracy of 91.9%, maintaining 85.8% diagnostic accuracy even with a high missing data rate of 50%, demonstrating outstanding data adaptability. Simultaneously, the model achieves a balance between high accuracy and lightweight design, with only 0.15M parameters and an inference latency of 0.2ms, directly meeting the deployment requirements of industrial edge devices and exhibiting excellent engineering practicality.

[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for diagnosing motor bearing faults due to missing time-series monitoring data, characterized in that, Includes the following steps: Step 1: Obtain multi-dimensional time-series monitoring data of motor bearings, perform preprocessing operations on the time-series monitoring data, and construct a missing mask matrix based on the data acquisition status markers. The missing mask matrix corresponds one-to-one with the time dimension and feature dimension of the time-series monitoring data, and is used to mark the missing location and missing rate of each time node and each dimension of data. Step 2: Based on the preprocessed complete time-series monitoring data and the missing mask matrix, the complete data is masked with different missing rates using a random masking method to generate polymorphic missing sample clusters. The polymorphic missing sample clusters include positive correlation sample pairs with different missing rates under the same time window, negative difference sample pairs under different time windows, and complete data samples. Step 3: Construct a teacher-student dual-track model linkage training framework: Adopt a heterogeneous backbone network configuration of teacher model and student model, train the teacher model with complete data samples and fix its parameters, and train the student model with the polymorphic missing sample cluster. During training, the differences between the student model output and the teacher model output are constrained by hidden feature transfer loss and performance transfer loss. At the same time, the feature contrast loss is used to constrain the feature output of the student model for samples with different missing rates to achieve semantic alignment. Step 4: Perform structured pruning optimization on the student model after training, calculate the importance score of each channel in the convolutional layer and fully connected layer in the model, sort each channel in descending order according to the importance score, remove the channels ranked lower according to the preset pruning rate, retain the remaining channels, and continue to train the pruned model with a learning rate lower than the learning rate in the training stage. Step 5: After preprocessing, the real-time collected multi-dimensional time-series monitoring data of the motor bearing, along with the corresponding real-time missing mask matrix, is input into the pruned and optimized student model. The student model directly extracts the fault semantic features and outputs the diagnostic results of the motor bearing.

2. The method according to claim 1, characterized in that, The preprocessing described in step one includes noise reduction and normalization: the noise reduction process employs an improved wavelet thresholding method, using wavelet coefficient thresholds... Wavelet coefficients larger than the threshold are subjected to soft thresholding, while wavelet coefficients smaller than the threshold are set to zero. , For a single sample time step, The data is time-series monitoring data of any dimension; the standardization process uses Z-score transformation to perform on the data of each dimension separately, so that the mean of the data of each dimension is 0 and the standard deviation is 1.

3. The method according to claim 1, characterized in that, The masking process described in step two covers missing rates of 20%, 40%, 60%, and 80%, encompassing random intermittent missing patterns, bursty continuous missing patterns, and conditionally dependent missing patterns. The positively correlated sample pairs are defined as samples with different missing rates within the same time window that are positively correlated with each other, and the negatively correlated sample pairs are defined as any missing samples from different time windows that are negatively correlated with each other.

4. The method according to claim 1, characterized in that, The teacher model described in step three uses a deep temporal convolutional network, with a network structure including an input layer, four convolutional layers, and a fully connected layer, and the number of channels is 8→256→512→512→256; the student model uses a lightweight temporal convolutional network, with a network structure including an input layer, three depthwise separable convolutional layers, and a fully connected layer, and the number of channels is 8→64→128→64.

5. The method according to claim 1, characterized in that, The total loss function for training the student model in step three is: in, To predict the loss, the difference between the student model's diagnostic results and the true labels is characterized; To hide the feature transfer loss, we characterize the differences between the hidden layer features of the student model and the hidden layer features of the teacher model; The performance transfer loss characterizes the difference between the diagnostic output of the student model and the diagnostic output of the teacher model. Feature contrast loss is used to characterize the contrast loss of features between samples with different missing rates. , , , These are the weighting coefficients for each loss.

6. The method according to claim 5, characterized in that, The formula for calculating the feature contrast loss is: in, This represents the number of samples within the batch. For positive feature output of samples, For the first The feature vector of each sample is generated by the student model. For the first batch The feature vector of each sample For cosine similarity, This refers to the temperature parameter.

7. The method according to claim 1, characterized in that, The importance score of the channel described in step four The calculation formula is: in, , To balance the weighting coefficients of the two scores, The number of convolution kernel weights, Indicates the first Channel parameters, Indicates the first d The first in the passage Each convolutional kernel weight, This is the gradient of the channel parameter with respect to the total loss. This represents the maximum value of the gradient across all channels.

8. The method according to claim 1, characterized in that, The diagnostic results in step five include the classification and identification of five states: normal state, inner ring fault, outer ring fault, rolling element fault, and cage fault. The student model uses Softmax classification logic to transform the extracted fault semantic features into probability distributions of the five states, thereby realizing end-to-end fault diagnosis reasoning.

9. The method according to claim 1, characterized in that, Step five also includes a reasoning reliability verification step: verify the missing rate of the current sample, and if the missing rate is greater than a preset threshold, it is marked as low reliability; verify the cosine similarity between the fused features of the current sample and the feature template of the same type of fault sample, and if the similarity is less than a preset threshold, it is judged as medium reliability; verify the consistency of the diagnostic results of multiple consecutive time windows, and if the results are consistent, it is judged as high reliability.

10. A motor bearing fault diagnosis system for missing time-series monitoring data, used to implement the method as described in any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire multi-dimensional time-series monitoring data of the motor bearing; The preprocessing module is used to perform preprocessing operations on the time-series monitoring data and construct a missing mask matrix based on the data acquisition status markers; The training module is used to generate polymorphic missing sample clusters based on the preprocessed complete time-series monitoring data and the missing mask matrix, and to train the student model based on the teacher-student dual-track model linkage training framework. The pruning optimization module is used to perform structured pruning optimization on the student model after training. The diagnostic reasoning module is used to input the real-time collected time-series monitoring data along with the corresponding real-time mask matrix into the pruned and optimized student model, extract fault semantic features, and output diagnostic results.