Gearbox gear fault detection method and system based on unsupervised domain self-adaption
Through unsupervised domain adaptation methods, utilizing feature extraction and distribution alignment techniques, combined with long short-term memory networks and multi-head self-attention mechanisms, the problems of data imbalance and scarce annotations in gearbox gear fault detection are solved, and high-precision fault detection and identification are achieved.
Patent Information
- Application Number
- CN202510701977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-26
AI Technical Summary
Existing gearbox gear fault detection methods face the problems of data imbalance and scarce annotations in practical applications, resulting in poor performance of traditional supervised learning methods in fault diagnosis tasks, especially in the difficulty of accurately identifying complex fault modes in real environments.
An unsupervised domain adaptation method is adopted to obtain labeled source domain data and unlabeled target domain data, perform feature extraction and distribution alignment, combine long short-term memory network and multi-head self-attention mechanism, dynamically adjust margins and weights, optimize model loss, and train the target fault detection model.
High-precision fault detection is achieved under unsupervised conditions, reducing dependence on manually labeled data, improving the ability to identify complex fault modes, and reducing engineering costs.
Smart Images

Figure CN120705729A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of fault detection, and in particular to a method and system for fault detection of a transmission gear based on unsupervised domain adaptation. Background Art
[0002] Gearboxes adjust the speed and torque of the output shaft through a combination of a series of gears and are widely used in automobiles, industrial machinery, wind power generation equipment and other fields. However, during long-term operation, the gears inside the gearbox may suffer from faults such as wear, cracks, broken teeth and bearing damage. If these faults cannot be detected or diagnosed in time, they may lead to equipment performance degradation or even catastrophic failure. Traditional gearbox gear fault detection and diagnosis methods mainly include methods based on vibration signal analysis, frequency domain analysis, and machine learning or deep learning. However, these methods have challenges in practical applications. Existing deep learning diagnosis methods usually assume that training samples and test samples have the same distribution, and the training data should contain rich label information. However, in practical applications, real fault data is difficult to obtain and annotations are scarce. At the same time, normal state data often dominates, resulting in serious data imbalance. This makes traditional supervised learning methods face many challenges in fault diagnosis tasks. Summary of the Invention
[0003] The present disclosure provides a gearbox gear fault detection method, system, electronic device and storage medium based on unsupervised domain adaptation to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, a method for fault detection of a transmission gear based on unsupervised domain adaptation is provided, the method comprising: acquiring sample data, the sample data comprising labeled source domain data and unlabeled target domain data; inputting the sample data into a fault detection model for feature extraction to obtain a target feature vector; performing distribution alignment between the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data in a feature space; inputting the high-dimensional feature vector of the sample data after distribution alignment into a classifier for fault category prediction to obtain a classification result; dynamically adjusting the margins and weights of different categories according to the classification result; determining a model loss of a fault detection model, the model loss comprising distribution alignment loss, classification loss, and margin-aware weighted loss; iteratively training the fault detection model based on the model loss to minimize the loss of the fault detection model to obtain a target fault detection model, wherein the target fault detection model is used for fault detection.
[0005] In one possible implementation manner, the inputting the sample data into a fault detection model for feature extraction to obtain a target feature vector includes: inputting the sample data into a set of stacked one-dimensional convolutional layers to extract local time-frequency features of the sample data; inputting the local time-frequency features into a bidirectional long short-term memory network to obtain time series features of the sample data; inputting the time series features into a multi-head self-attention mechanism to obtain time series dynamic features of the sample data; and inputting the time series dynamic features into a flattening layer for flattening to obtain a target feature vector.
[0006] In one possible implementation manner, the target feature vector corresponding to the source domain data is distributedly aligned with the target feature vector corresponding to the target domain data in the feature space, including: determining the maximum mean difference of the global alignment based on the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data through a high-dimensional feature mapping function; dividing the source domain data into multiple source domain subdomains according to the label of the source domain data, each source domain subdomain corresponding to a fault type; determining the category center of each source domain subdomain; determining the pseudo label of the target domain data according to the distance from the target domain data to the category center of each source domain subdomain; and aligning the target domain data according to the pseudo label. The domain data is divided into multiple target domain subdomains, and the number of the target domain subdomains is the same as the number of the source domain subdomains; for the source domain subdomains and the target domain subdomains corresponding to the same fault type, the maximum mean difference of the subdomain alignment is determined based on the target feature vector corresponding to the source domain subdomain data and the target feature vector corresponding to the target domain subdomain data through a high-dimensional feature mapping function; the target maximum mean difference is determined according to the maximum mean difference of the global alignment and the maximum mean difference of each subdomain alignment; the minimum value of the target maximum mean difference is determined to align the target feature vector corresponding to the source domain data with the target feature vector corresponding to the target domain data in the feature space distribution.
[0007] In one possible implementation, the dynamic adjustment of the margins and weights of different categories based on the classification results includes: determining the marginal coefficient corresponding to each category based on the number of samples corresponding to each category; determining the cosine similarity of each category based on the characteristic vector of the sample data and the weight vector of each category; scaling the angle corresponding to the cosine similarity according to the marginal coefficient to obtain the cosine similarity after angle scaling; determining the predicted probability of each category based on the cosine similarity after angle scaling to achieve marginal adjustment of different categories; determining the frequency corresponding to each category based on the number of labeled sample data contained in each category; and determining the category weight of each category based on the frequency corresponding to each category to achieve weight adjustment of different categories.
[0008] In one possible implementation, after acquiring the sample data, the method further includes preprocessing the sample data.
[0009] In one possible implementation manner, the method further includes: acquiring data to be detected collected by a transmission sensor; inputting the data to be detected into the target fault detection model, and determining a target category of the data to be detected.
[0010] According to a second aspect of the present disclosure, a fault detection system for a transmission gear based on unsupervised domain adaptation is provided, characterized in that the system comprises: an acquisition module for acquiring sample data, wherein the sample data comprises labeled source domain data and unlabeled target domain data; a feature extraction module for inputting the sample data into a fault detection model for feature extraction to obtain a target feature vector; a distribution alignment module for performing distribution alignment between the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data in a feature space; a prediction module for inputting the high-dimensional feature vector of the sample data after distribution alignment into a classifier for fault category prediction to obtain a classification result; an adjustment module for dynamically adjusting the margins and weights of different categories according to the classification result; a determination module for determining the model loss of the fault detection model, wherein the model loss comprises distribution alignment loss, classification loss and margin-aware weighted loss; and a training module for iteratively training the fault detection model based on the model loss to minimize the loss of the fault detection model and obtain a target fault detection model, wherein the target fault detection model is used for fault detection.
[0011] In one possible implementation, the feature extraction module is specifically used to input the sample data into a set of stacked one-dimensional convolutional layers to extract local time-frequency features of the sample data; input the local time-frequency features into a bidirectional long short-term memory network to obtain time series features of the sample data; input the time series features into a multi-head self-attention mechanism to obtain time series dynamic features of the sample data; input the time series dynamic features into a flattening layer for flattening operation to obtain a target feature vector.
[0012] In one embodiment, the distribution alignment module is specifically configured to determine a maximum mean difference of global alignment based on a target feature vector corresponding to the source domain data and a target feature vector corresponding to the target domain data through a high-dimensional feature mapping function; divide the source domain data into multiple source domain subdomains according to the label of the source domain data, each source domain subdomain corresponding to a fault type; determine a category center of each source domain subdomain; determine a pseudo label of the target domain data based on a distance from the target domain data to the category center of each source domain subdomain; divide the target domain data into multiple target domain subdomains according to the pseudo label, the number of the target domain subdomains being the same as the number of the source domain subdomains; for source domain subdomains and target domain subdomains corresponding to the same fault type, determine a maximum mean difference of subdomain alignment based on a target feature vector corresponding to the source domain subdomain data and a target feature vector corresponding to the target domain subdomain data through a high-dimensional feature mapping function; determine a target maximum mean difference based on the maximum mean difference of the global alignment and the maximum mean difference of each subdomain alignment; and determine a minimum value of the target maximum mean difference so that the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data are distributed aligned in the feature space.
[0013] In one possible implementation, the adjustment module is specifically used to determine the marginal coefficient corresponding to each category based on the number of samples corresponding to each category; determine the cosine similarity of each category based on the characteristic vector of the sample data and the weight vector of each category; scale the angle corresponding to the cosine similarity according to the marginal coefficient to obtain the cosine similarity after angle scaling; determine the predicted probability of each category based on the cosine similarity after angle scaling to achieve marginal adjustment of different categories; determine the frequency corresponding to each category based on the number of labeled sample data contained in each category; determine the category weight of each category based on the frequency corresponding to each category to achieve weight adjustment of different categories.
[0014] In one embodiment, the system further includes: a preprocessing module, configured to preprocess the sample data after the sample data is acquired.
[0015] In one embodiment, the system further includes: a detection module configured to obtain the data to be detected collected by the transmission sensor; input the data to be detected into the target fault detection model, and determine a target category of the data to be detected.
[0016] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.
[0020] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the present disclosure.
[0021] The present invention discloses a fault detection method, system, electronic device and storage medium for a transmission gear based on unsupervised domain adaptation. The method first obtains sample data, which includes labeled source domain data and unlabeled target domain data; inputs the sample data into a fault detection model for feature extraction to obtain a target feature vector; distributes and aligns the target feature vector corresponding to the source domain data with the target feature vector corresponding to the target domain data in a feature space; inputs the high-dimensional features of the sample data after distribution alignment into a classifier to predict the fault category and obtain a classification result; dynamically adjusts the margins and weights of different categories according to the classification result; determines the model loss of the fault detection model, which includes distribution alignment loss, classification loss and margin-aware weighted loss; iteratively trains the fault detection model based on the model loss to minimize the loss of the fault detection model and obtain a target fault detection model, which is used for fault detection.
[0022] This method uses sample data to train the fault detection model. When extracting features, the combination of long-short-term memory networks and a multi-head self-attention mechanism effectively captures spatiotemporal features, enabling the model to more accurately identify complex fault patterns. It also optimizes class imbalance, preventing traditional methods from overlooking minority faults and improving overall diagnostic performance. Through an unsupervised subdomain adaptive mechanism, efficient migration from simulation data to actual device data is achieved. High-precision diagnosis can be maintained even with significant differences in data distribution, achieving high accuracy under unsupervised conditions and reducing reliance on manually labeled data. Finally, the target fault detection model is obtained, which can be used to detect gearbox fault data and determine the corresponding fault type. This allows for rapid and accurate determination of the corresponding fault type, reducing engineering costs.
[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0025] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0026] Figure 1 The schematic diagram of the implementation process of a gearbox gear fault detection method based on unsupervised domain adaptation in the embodiment of the present disclosure is shown. Figure 1 ;
[0027] Figure 2 The schematic diagram of the implementation process of a gearbox gear fault detection method based on unsupervised domain adaptation in the embodiment of the present disclosure is shown. Figure 2 ;
[0028] Figure 3 A schematic diagram showing feature extraction of sample data according to an embodiment of the present disclosure is shown;
[0029] Figure 4 The schematic diagram of the implementation process of a gearbox gear fault detection method based on unsupervised domain adaptation in the embodiment of the present disclosure is shown. Figure 3 ;
[0030] Figure 5 The schematic diagram of the implementation process of a gearbox gear fault detection method based on unsupervised domain adaptation in the embodiment of the present disclosure is shown. Figure 4 ;
[0031] Figure 6 A module schematic diagram of a gearbox gear fault detection system based on unsupervised domain adaptation according to an embodiment of the present disclosure is shown;
[0032] Figure 7 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0033] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.
[0034] Figure 1 The schematic diagram of the implementation process of a gearbox gear fault detection method based on unsupervised domain adaptation in the embodiment of the present disclosure is shown. Figure 1,include:
[0035] Step 101: Obtain sample data, where the sample data includes labeled source domain data and unlabeled target domain data.
[0036] Step 102: Input the sample data into the fault detection model to extract features to obtain a target feature vector.
[0037] First, sample data is obtained. The sample data in this application includes labeled source domain data and unlabeled target domain data. Labeled source domain data refers to high-fidelity simulation data with identified fault types or fault data from devices in actual applications. Unlabeled target domain data refers to fault data from devices in actual operation with unknown fault types. The sample data is then input into a fault detection model, which first extracts features from the sample data to obtain a target feature vector for the sample data. This target feature vector includes the target feature vector corresponding to the source domain data and the target domain data.
[0038] Step 103: align the target feature vector corresponding to the source domain data with the target feature vector corresponding to the target domain data in the feature space.
[0039] There are significant differences in feature distribution between source domain data and target domain data, which leads to a decrease in migration performance. Therefore, this application aligns the target feature vectors corresponding to the source domain data with the target feature vectors corresponding to the target domain data in the feature space to reduce the distribution differences between the source domain data and the target domain data in the feature space, making the feature distributions of the source domain data and the target domain data more consistent. Furthermore, the distribution alignment of this application includes global alignment and subdomain alignment, which can enable the fault recognition model to learn more robust domain-invariant features, avoid overfitting the source domain, and adapt to the target domain data.
[0040] Step 104 : Input the high-dimensional feature vector of the sample data after distribution alignment into a classifier to perform fault category prediction and obtain a classification result.
[0041] After feature extraction and distribution alignment, the high-dimensional features of all sample data are input into the classifier module for category prediction. The classifier typically consists of a convolutional layer, a fully connected layer, and a softmax output layer, and is trained using a standard cross-entropy loss function. Fault types are determined based on actual application requirements and typically include typical gearbox fault categories such as normal state, early wear, gear fracture, tooth root cracks, and meshing anomalies.
[0042] Step 105 : Dynamically adjust the margins and weights of different categories based on the classification results.
[0043] There is a serious class imbalance problem in the actual industrial environment. For example, the number of normal sample data is far greater than the fault sample data. Therefore, the minority class fault sample data is easily ignored during training. Therefore, this application can dynamically adjust the margins and weights of different categories according to the classification results during the training of the fault detection model. First, a larger discrimination margin is set for the minority class samples so that the minority class samples can obtain a larger classification space in the feature space. Secondly, the weights are dynamically adjusted according to the distance from the decision boundary of the sample data domain. Samples close to the boundary or easily confused receive higher training attention, and samples far from the boundary and easy to classify receive lower training attention. By dynamically adjusting the margins and weights of different categories, the recognition ability of minority classes can be improved, the discrimination ability of complex samples near the decision boundary can be enhanced, and the classification bias caused by data imbalance can be alleviated.
[0044] Step 106 : Determine the model loss of the fault detection model. The model loss includes distribution alignment loss, classification loss, and margin-aware weighted loss.
[0045] Step 107 , iteratively train the fault detection model based on the model loss to minimize the fault detection model loss, thereby obtaining a target fault detection model. The target fault detection model is used for fault detection.
[0046] During the training process of the fault detection model, the model loss of the fault detection model is determined. This model loss includes distribution alignment loss, classification loss, and margin-aware weighted loss. Distribution alignment loss is the loss incurred when aligning the target feature vector corresponding to the source domain data with the target feature vector corresponding to the target domain data in the feature space. Classification loss is the loss incurred when the classifier predicts the fault category. Margin-aware weighted loss is the loss incurred when dynamically adjusting the margins and weights of different categories. The fault detection model is iteratively trained based on these model losses. When the loss of the fault detection model is minimized, the target fault detection model is obtained. The target fault detection model can then be used to detect the data to be detected and determine the fault category.
[0047] This method is used to train the initial fault detection model with sample data to obtain the target fault detection model. The unlabeled fault data is then input into the target fault detection model for detection to determine the fault type corresponding to the fault data. This can achieve high-precision prediction of the fault type, reduce dependence on manually labeled data, and reduce engineering costs.
[0048] In one embodiment, Figure 2 As shown, the sample data is input into the fault detection model for feature extraction to obtain the target feature vector, including:
[0049] Step 201: input the sample data into a set of stacked one-dimensional convolutional layers to extract the local time-frequency features of the sample data;
[0050] Step 202: Input the local time-frequency features into a bidirectional long short-term memory network to obtain the time series features of the sample data;
[0051] Step 203: Input the time series features into the multi-head self-attention mechanism to obtain the time series dynamic features of the sample data;
[0052] Step 204: Input the temporal dynamic features into the flattening layer for flattening to obtain a target feature vector.
[0053] First, the original sample data is input into a set of stacked one-dimensional convolutional layers. Figure 3 This is a schematic diagram of feature extraction of sample data in an embodiment of the present application, as shown in FIG. Figure 3 As shown in the figure, it includes 4 one-dimensional convolution layers to extract local time-frequency features and local time-frequency features. Through the convolution operation, it is possible to capture short-term dynamic changes in the signal. Each convolution layer is connected to a batch normalization layer and a ReLU activation function. Normalization and nonlinear activation help stabilize the training process and enhance feature expression capabilities. In the process of convolution processing through stacked one-dimensional convolution layers, pooling layers are inserted at intervals. Figure 3 The pooling layer used in
[15] is the maximum pooling layer to reduce the dimension of the feature sequence, highlight the most representative local features, and reduce the computational burden.
[0054] The local time-frequency features, after convolution and pooling, are then fed into a bidirectional long short-term memory (BiLSTM) network. BiLSTM can simultaneously model both forward and reverse temporal dependencies in a sequence, fully understanding both historical and future contextual information in the signal. This bidirectional modeling approach helps capture the long-term dynamic evolution of gear fault signals, significantly improving the system's ability to characterize complex, time-varying fault patterns.
[0055] The time series features output by the BiLSTM are fed into a multi-head self-attention (MHSA) mechanism. The MHSA module uses multiple parallel attention heads to learn the complex dependencies between signal time steps from different subspaces. This mechanism dynamically assigns different importance weights to different time steps in the feature sequence, highlighting key dynamic features while effectively suppressing irrelevant noise and redundant information, further improving the quality of feature representation and discriminative capabilities.
[0056] Finally, the flattening layer flattens the temporal dynamic features output by the MHSA module and converts them into a fixed-length vector representation to obtain the target feature vector. This fixed-length representation not only integrates local time-frequency features, temporal dynamic features, and global dependencies, but also has good transferability and discriminability.
[0057] Through the organic integration of convolutional feature extraction, bidirectional time series modeling and global attention mechanism, the key characteristics of gearbox gear fault signals can be fully explored, effectively improving the accuracy and robustness of fault retrieval and classification.
[0058] In one embodiment, Figure 4 As shown, the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data are distributed and aligned in the feature space, including:
[0059] Step 401: determining the maximum mean difference of global alignment based on the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data through a high-dimensional feature mapping function;
[0060] Step 402: Divide the source domain data into multiple source domain subdomains according to the labels of the source domain data, where each source domain subdomain corresponds to a fault type.
[0061] Step 403, determining the category center of each source domain subdomain;
[0062] Step 404: Determine the pseudo label of the target domain data based on the distance between the target domain data and the category center of each source domain subdomain;
[0063] Step 405: Divide the target domain data into multiple target domain subdomains according to the pseudo labels, where the number of target domain subdomains is the same as the number of source domain subdomains.
[0064] Step 406: For the source domain subdomain and the target domain subdomain corresponding to the same fault type, the maximum mean difference of the subdomain alignment is determined by a high-dimensional feature mapping function based on the target feature vector corresponding to the source domain subdomain data and the target domain subdomain data.
[0065] Step 407 , determining a target maximum mean difference based on the maximum mean difference of the global alignment and the maximum mean difference of each subdomain alignment;
[0066] Step 408 : Determine the minimum value of the target maximum mean difference so that the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data are aligned in the feature space.
[0067] In the gearbox gear fault detection task, there are significant differences in feature distribution between the source domain and the target domain. This application optimizes the domain adaptation strategy under unsupervised conditions to make the feature distribution of the source domain and the target domain more consistent, thereby improving the generalization performance of the model in the target domain.
[0068] In order to reduce the difference in the overall feature distribution between the source domain and the target domain, the local maximum mean discrepancy (LMMD) is first used to perform global distribution alignment. LMMD is a distance metric based on feature mean, which is used to quantify the consistency of the embedded feature distribution between two domains. Let the target feature vector set corresponding to the source domain data be The target feature vector set corresponding to the target domain data is The maximum mean difference (LMMD) of global alignment determined by a high-dimensional feature mapping function can be expressed as: in is a high-dimensional feature mapping function.
[0069] Since different fault categories have finer-grained differences in feature distribution, global alignment alone cannot guarantee a one-to-one correspondence between the features of each fault category, and may even cause category confusion. Therefore, the present invention further performs feature alignment at the subdomain level based on global alignment. The specific process is as follows: First, the source domain data is divided into multiple source domain subdomains according to the label of the source domain, and each source domain subdomain corresponds to a fault type. Based on the divided source domain subdomains, the category center of the source domain subdomain is determined based on the source domain data in the source domain subdomain. The distance from the target domain data to the category center of each source domain subdomain is calculated, and the fault type of the source domain subdomain corresponding to the category center closest to it is determined as the pseudo label of the target domain data. The target domain data is divided into multiple target domain subdomains according to the pseudo label of the target domain data, and the number of target domain subdomains is the same as the number of source domain subdomains. For the source domain subdomain and target domain subdomain corresponding to each fault type, the maximum mean difference of the subdomain alignment is determined by the high-dimensional feature mapping function. The maximum mean difference LMMD of the subdomain alignment corresponding to the fault category can be expressed as: in and They represent the number of samples of the cth fault type in the source domain subdomain and the target domain subdomain respectively.
[0070] The maximum mean differences corresponding to all subdomain alignments are accumulated, and then the accumulated maximum mean differences of the subdomain alignments are weighted with the maximum mean difference of the global alignment to obtain the target maximum mean difference. The formula is: Among them, λ1 and λ2 are the weight coefficients corresponding to the maximum mean difference of global alignment and the maximum mean difference of subdomain alignment respectively. Finally, the minimum value of the target maximum mean difference is determined and the target maximum mean difference is minimized to improve the robustness of the model. By minimizing LMMD global , which can effectively reduce the overall mean difference between source domain data and target domain data in the feature space and alleviate the negative transfer effect caused by global distribution mismatch.
[0071] This process is implemented based on an unsupervised mechanism. At the global level, the mean difference between the feature vectors of the source and target domain data is calculated to reduce the risk of negative transfer caused by overall distribution discrepancies. A pseudo-labeling mechanism is introduced at the category level to soft-assign the target domain data based on the similarity between its features and the category center of the source domain data, forming a pseudo-subdomain structure of the target domain. For each category of source and target domain data, the local mean difference is calculated, constraining the model to learn consistent representations under subdomain conditions. This joint alignment strategy not only ensures the consistency of the overall distribution of the source and target domain data, but also fine-tunes the matching of subdomain features at the type level, preserving type discriminability to the greatest extent and avoiding confusion between different fault modes. The global and local alignment losses are jointly optimized in a weighted manner to ensure that the model effectively aligns category semantics while maintaining inter-class discrimination. Ultimately, a representation with strong cross-domain adaptability is obtained, improving the accuracy and robustness of the model for gearbox gear fault identification in actual working conditions, providing a reliable solution for intelligent fault diagnosis in low-label environments in industrial applications.
[0072] In one embodiment, Figure 5 As shown, the margins and weights of different categories are dynamically adjusted according to the classification results, including:
[0073] Step 501, determining the marginal coefficient corresponding to each category based on the number of samples corresponding to each category;
[0074] Step 502, determining the cosine similarity of each category based on the feature vector of the sample data and the weight vector of each category;
[0075] Step 503: scaling the angle corresponding to the cosine similarity according to the marginal coefficient to obtain the cosine similarity after angle scaling;
[0076] Step 504: Determine the predicted probability of each category based on the cosine similarity after angle scaling to achieve margin adjustment for different categories;
[0077] Step 505 , determining the frequency corresponding to each category based on the number of labeled sample data contained in each category;
[0078] Step 506: Determine the category weight of each category based on the frequency corresponding to each category to adjust the weights of different categories.
[0079] In the gearbox gear fault identification, there are problems such as difficulty in distinguishing minority class samples and fuzzy boundaries between different fault types. By adjusting the decision boundary based on category awareness, it is possible to guide the fault detection model to strengthen its learning ability for edge samples and, thus, effectively improve the classification accuracy and boundary discrimination. Specifically, if the number of samples in the cth class is N c , then the corresponding marginal coefficient is defined as:
[0080] This function indicates that the minority class should obtain a larger decision margin, thereby increasing the distance from other types in the feature space and reducing the overlap of judgments. In practical applications, the output can be determined as the cosine similarity of each type: Among them, a pred is the feature vector of the sample to be classified, W c is the weight vector of category c. To introduce marginal constraints, the angle θ c Scaling by category marginal coefficient is: θ c ‘ =M c ·θ c ; Based on the scaled angle θ c ‘ Determine its corresponding cosine similarity cosθ c ‘ ; Then calculate the adjusted fault category score based on this: z c =‖W c ‖·||a pred ||·cosθ' c +b c Finally, by normalizing the scores of the fault categories, the predicted probability of each fault category is obtained: By dynamically adjusting the margins of different categories, the discrimination boundary of the minority class is significantly stretched, reducing the risk of confusion.
[0081] In order to further improve the model's attention to minority class faults, MAAWM introduces an adaptive class weight adjustment strategy during the training phase. Suppose the frequency of class c is: Where I(·) is an indicator function, which is used to count the number of labels belonging to category c in the sample data. Based on this, the category weight is determined as: When there are fewer labeled sample data in a certain fault category, the weight corresponding to the fault category is relatively increased to avoid excessive bias towards high-frequency categories in model training.
[0082] The specific weighted cross entropy loss function is expressed as: Then average the sample loss in the small batch to get the final loss expression: In this way, the model's perception of minority class samples is improved, and the model is guided to strengthen its learning ability for rare class faults during training, so that the model can still maintain high precision, high robustness and good generalization ability in multi-category imbalanced fault diagnosis tasks.
[0083] In one embodiment, after acquiring the sample data, the method further includes preprocessing the sample data.
[0084] It is understandable that the sample data mainly includes vibration signals or acoustic signals. Due to the complexity of actual working conditions, these original signals often contain a variety of interference information, such as noise, trend drift and non-stationary components, so the sample data needs to be preprocessed. The preprocessing process of sample data in this application includes four steps: first, denoising is performed by wavelet threshold filtering or bandpass filtering to filter out high-frequency noise and low-frequency background interference; second, detrending processing is performed using methods such as sliding average to remove baseline drift components in the signal; then, the signal is standardized using the normalization method to unify the amplitude scale between samples and improve the consistency of model training; finally, according to the requirements of the model input format, the time series signal is segmented into equal lengths, and a sliding window strategy is often used to divide long sequences into fixed-length segments.
[0085] In one embodiment, the method further includes: acquiring data to be detected collected by a transmission sensor; inputting the data to be detected into a target fault detection model, and determining a target category of the data to be detected.
[0086] After training the fault detection model as described above to obtain the target fault detection model, the target fault detection model has good fault detection capabilities. Therefore, it is possible to obtain the data to be detected collected by the transmission sensors, namely, fault data collected by various types of transmission sensors such as triaxial accelerometers and acoustic sensors in real-world operating environments. This data can be input into the target fault detection model, which then outputs the target category corresponding to the data to be detected.
[0087] It is understandable that after the target fault detection model is used to detect the data to be detected and the target category is determined, the present application can perform a visual analysis of the detection results, including drawing a confusion matrix to evaluate the classification accuracy, generating a precision-recall curve to evaluate the minority class recognition ability, and applying dimensionality reduction technology to visualize the embedded features to assist users in understanding the judgment boundaries and feature distribution characteristics of the model under different working conditions.
[0088] Figure 6A module schematic diagram of a transmission gear fault detection system based on unsupervised domain adaptation according to an embodiment of the present disclosure is shown.
[0089] See also Figure 6 According to a second aspect of the present disclosure, a gearbox gear fault detection system based on unsupervised domain adaptation is provided, characterized in that the system includes: an acquisition module 601 for acquiring sample data, wherein the sample data includes labeled source domain data and unlabeled target domain data; a feature extraction module 602 for inputting the sample data into a fault detection model for feature extraction to obtain a target feature vector; a distribution alignment module 603 for performing distribution alignment between the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data in a feature space; a prediction module 604 for inputting the high-dimensional feature vector of the sample data after distribution alignment into a classifier for fault category prediction to obtain a classification result; an adjustment module 605 for dynamically adjusting the margins and weights of different categories based on the classification result; a determination module 606 for determining a model loss of the fault detection model, wherein the model loss includes distribution alignment loss, classification loss, and margin-aware weighted loss; and a training module 607 for iteratively training the fault detection model based on the model loss to minimize the loss of the fault detection model and obtain a target fault detection model, wherein the target fault detection model is used for fault detection.
[0090] In one embodiment, the feature extraction module 602 is specifically used to input the sample data into a set of stacked one-dimensional convolutional layers to extract local time-frequency features of the sample data; input the local time-frequency features into a bidirectional long short-term memory network to obtain time series features of the sample data; input the time series features into a multi-head self-attention mechanism to obtain time series dynamic features of the sample data; input the time series dynamic features into a flattening layer for flattening operation to obtain a target feature vector.
[0091] In one embodiment, the distribution alignment module 603 is specifically configured to determine a maximum mean difference of global alignment based on a target feature vector corresponding to the source domain data and a target feature vector corresponding to the target domain data through a high-dimensional feature mapping function; divide the source domain data into multiple source domain subdomains according to the label of the source domain data, each source domain subdomain corresponding to a fault type; determine a category center of each source domain subdomain; determine a pseudo label of the target domain data based on a distance from the target domain data to the category center of each source domain subdomain; divide the target domain data into multiple target domain subdomains based on the pseudo label, the number of the target domain subdomains being the same as the number of the source domain subdomains; for source domain subdomains and target domain subdomains corresponding to the same fault type, determine a maximum mean difference of subdomain alignment based on a target feature vector corresponding to the source domain subdomain data and a target feature vector corresponding to the target domain subdomain data through a high-dimensional feature mapping function; determine a target maximum mean difference based on the maximum mean difference of global alignment and the maximum mean difference of each subdomain alignment; and determine a minimum value of the target maximum mean difference so that the target feature vector corresponding to the source domain data is distributed aligned with the target feature vector corresponding to the target domain data in the feature space.
[0092] In one embodiment, the adjustment module 605 is specifically used to determine the marginal coefficient corresponding to each category based on the number of samples corresponding to each category; determine the cosine similarity of each category based on the characteristic vector of the sample data and the weight vector of each category; scale the angle corresponding to the cosine similarity according to the marginal coefficient to obtain the cosine similarity after angle scaling; determine the predicted probability of each category based on the cosine similarity after angle scaling to achieve marginal adjustment of different categories; determine the frequency corresponding to each category based on the number of labeled sample data contained in each category; determine the category weight of each category based on the frequency corresponding to each category to achieve weight adjustment of different categories.
[0093] In one embodiment, the system further includes a pre-processing module 608, configured to pre-process the sample data after the sample data is acquired.
[0094] In one embodiment, the system further includes: a detection module 609 for acquiring data to be detected collected by a transmission sensor; inputting the data to be detected into the target fault detection model, and determining a target category of the data to be detected.
[0095] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0096] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0097] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0098] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0099] The computing unit 701 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as a method for detecting a fault in a transmission gear. For example, in some embodiments, a method for detecting a fault in a transmission gear can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method for detecting a fault in a transmission gear described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute a transmission gear fault detection method in any other appropriate manner (for example, by means of firmware).
[0100] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0101] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0104] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0105] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0106] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0107] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.
[0108] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A gearbox gear fault detection method based on unsupervised domain adaptation, characterized in that: The method comprises: Acquire sample data, where the sample data includes labeled source domain data and unlabeled target domain data; Inputting the sample data into a fault detection model to extract features to obtain a target feature vector; Aligning the target feature vector corresponding to the source domain data with the target feature vector corresponding to the target domain data in a feature space; Inputting the high-dimensional feature vector of the sample data after distribution alignment into the classifier to predict the fault category and obtain a classification result; Dynamically adjusting margins and weights of different categories based on the classification results; determining a model loss of a fault detection model, wherein the model loss comprises a distribution alignment loss, a classification loss, and a margin-aware weighted loss; The fault detection model is iteratively trained based on the model loss to minimize the fault detection model loss, thereby obtaining a target fault detection model, which is used for fault detection.
2. The method according to claim 1, characterized in that The step of inputting the sample data into a fault detection model to extract features to obtain a target feature vector includes: Inputting the sample data into a set of stacked one-dimensional convolutional layers to extract local time-frequency features of the sample data; Inputting the local time-frequency features into a bidirectional long short-term memory network to obtain the time series features of the sample data; Inputting the time series features into a multi-head self-attention mechanism to obtain the time series dynamic features of the sample data; The temporal dynamic features are input into the flattening layer for flattening operation to obtain the target feature vector.
3. The method according to claim 1, characterized in that The aligning the target feature vector corresponding to the source domain data with the target feature vector corresponding to the target domain data in a feature space includes: Based on the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data, the maximum mean difference of the global alignment is determined through a high-dimensional feature mapping function; Divide the source domain data into multiple source domain subdomains based on the source domain data's labels, with each source domain subdomain corresponding to a fault type. Determine the category center of each source domain subdomain; Determine a pseudo label for the target domain data based on the distance between the target domain data and the category center of each source domain subdomain; Dividing the target domain data into a plurality of target domain subdomains according to the pseudo labels, wherein the number of the target domain subdomains is the same as the number of the source domain subdomains; For the source domain subdomain and target domain subdomain corresponding to the same fault type, the maximum mean difference of subdomain alignment is determined through a high-dimensional feature mapping function based on the target feature vector corresponding to the source domain subdomain data and the target domain subdomain data; determining a target maximum mean difference according to the maximum mean difference of the global alignment and the maximum mean difference of each subdomain alignment; The minimum value of the target maximum mean difference is determined so that the target feature vector corresponding to the source domain data and the target feature vector corresponding to the target domain data are aligned in feature space distribution.
4. The method according to claim 1, wherein The dynamically adjusting the margins and weights of different categories according to the classification results includes: According to the number of samples corresponding to each category, determine the marginal coefficient corresponding to each category; According to the feature vector of the sample data and the weight vector of each category, the cosine similarity of each category is determined; Scaling the angle corresponding to the cosine similarity according to the marginal coefficient to obtain the cosine similarity after angle scaling; The predicted probability of each category is determined based on the cosine similarity after angle scaling to achieve marginal adjustments for different categories; Determine the frequency of each category based on the number of labeled sample data contained in each category; The category weight of each category is determined based on the frequency corresponding to each category to achieve weight adjustment of different categories.
5. The method according to claim 1, wherein After acquiring the sample data, the method further includes preprocessing the sample data.
6. The method according to claim 1, wherein The method further comprises: Obtain the data to be tested collected by the gearbox sensor; The data to be detected is input into the target fault detection model to determine the target category of the data to be detected.
7. A gearbox gear fault detection system based on unsupervised domain adaptation, characterized in that: The system comprises: An acquisition module is used to acquire sample data, where the sample data includes labeled source domain data and unlabeled target domain data; A feature extraction module is used to input the sample data into a fault detection model to extract features and obtain a target feature vector; A distribution alignment module, configured to distribute and align the target feature vector corresponding to the source domain data with the target feature vector corresponding to the target domain data in a feature space; A prediction module, configured to input the high-dimensional feature vector of the sample data after distribution alignment into a classifier to perform fault category prediction and obtain a classification result; an adjustment module for dynamically adjusting margins and weights of different categories based on the classification results; A determination module, configured to determine a model loss of a fault detection model, wherein the model loss includes a distribution alignment loss, a classification loss, and a margin-aware weighted loss; The training module is used to iteratively train the fault detection model based on the model loss to minimize the loss of the fault detection model and obtain a target fault detection model, which is used for fault detection.
8. The system according to claim 7, characterized in that The feature extraction module is specifically used to Inputting the sample data into a set of stacked one-dimensional convolutional layers to extract local time-frequency features of the sample data; Inputting the local time-frequency features into a bidirectional long short-term memory network to obtain the time series features of the sample data; Inputting the time series features into a multi-head self-attention mechanism to obtain the time series dynamic features of the sample data; The temporal dynamic features are input into the flattening layer for flattening operation to obtain the target feature vector.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.