Adaptive feature adjustment optimization method for HDMS wear fault diagnosis

Through the adaptive feature adjustment optimization method (AFAO), the decision boundaries are broadened and the weight is adaptively adjusted, and the misdiagnosis problem caused by long-tail distribution in HDMS wear fault diagnosis is solved, improving the identification accuracy of fault types and the learning effect of classifiers.

CN120372477AActive Publication Date: 2025-07-25SICHUAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510406540.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The prior art faces misdiagnosis caused by long-tail distribution data sets in HDMS wear fault diagnosis. The resampling method is time-consuming and relies on professional knowledge, and the weight allocation is unreasonable. The traditional boundary adjustment method is static and cannot adapt to classifier changes, resulting in confusion in the classification of fault types.

Method used

By extending the clustering of training sample features, compressing the clustering of test samples, using pre-trained models to evaluate the similarity of fault types, adaptively adjust sample point boundaries and weights, and adopting the adaptive feature adjustment optimization method (AFAO) to broaden the decision boundaries and allocate reasonable weights to improve the classifier learning effect.

Benefits of technology

It significantly improves the recognition accuracy of tail fault categories, reduces the probability of misclassification, enhances the classifier's learning ability on various fault types, and adapts to the differentiation of complex fault types in industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372477A_ABST
    Figure CN120372477A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of fluid dynamic pressure sealing fault detection, and discloses a self-adaptive feature adjustment optimization method for HDMS wear fault diagnosis, which comprises the following steps: step 1, pre-training an HDMS wear fault diagnosis network model, and extracting original features; calculating the similarity between the categories, and constructing a similarity window according to the similarity; 2, performing amplification processing on the original features by adopting feature cluster compression; 3, introducing a similarity window boundary value into the cross entropy loss function to form an SWM loss function, and then adjusting the SWM loss function through a sample weight to construct an SWRM loss function; and 4, initializing the network model, and gradually optimizing model parameters through multi-round iteration. According to the method, the learning effect of the classifier on different fault types is balanced by increasing the density of the original feature clusters and combining the similarity features with the frequency of the samples; model confusion from similar HDMS samples is reduced, and accuracy of diagnosis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fluid dynamic press seal fault detection, and particularly relates to an adaptive feature adjustment and optimization method for HDMS wear fault diagnosis. Background Art

[0002] In nuclear power units, due to its small contact wear and long service life, the hydrodynamic mechanical seal (HDMS) has become the preferred solution for the shaft seal of the reactor coolant pump. One of the main failure modes of HDMS is the wear of the sealing surface. In actual engineering, the operating reliability requirements for HDMS are extremely high, so it is very difficult to obtain fault samples. In addition, due to artificial active maintenance intervention, it is also very difficult to obtain rare fault types caused by extreme operating conditions. Therefore, the wear fault diagnosis dataset of HDMS often presents a long-tail distribution.

[0003] When the sample data exhibits a long-tailed distribution, deep neural networks (DNNs) may lead to misdiagnosis. To improve the performance of DNNs on long-tailed distribution datasets, Bowyer et al. proposed the Synthetic Minority Over-sampling Technique (SMOTE) to rebalance the dataset. Liu et al. proposed an exploratory resampling method to undersample each class of samples in each sample batch. Lin et al. proposed focal loss, which reduces the loss weight of correctly classified samples by adding a weighting factor to the standard cross-entropy (CE) loss function. Cui et al. proposed the Class-Balanced Cross Entropy (CBCE) loss and the Class-Balanced Focal (CBF) loss to incorporate the effective number of different classes. Ren et al. proposed a meta-sampler to resample imbalanced datasets and proposed the Balanced Softmax Cross Entropy (BSCE) loss to adapt to the label distribution change between training and testing. Jitkrittum et al. proposed using embeddings and logit margins to adjust the output of the classifier and suppress long-tailed samples from crossing the decision boundary. These methods have been introduced into the field of fault diagnosis to address the intelligent fault detection problem of different devices under long-tailed distributions. Santos et al. combined SMOTE and the bagging method to successfully detect faults in wind turbine gearboxes. Zhou et al. proposed the Adaptive Vine Copula-based Dependence Description (AVCDD) to select the most informative data from imbalanced datasets in the Industrial Internet of Things (IIoT), successfully rebalancing the training set. Xi et al. proposed a squeeze-and-excitation initialized residual network model with focal loss to pay more attention to difficult-to-classify faults in transmission lines. Zhao et al. proposed a new weighted feature parameter information reconstruction method and applied it to the fault detection of aero-engine main bearings. This method successfully reduced the interference of noise components to the overall signal and effectively highlighted the fault feature information. Li et al. also proposed a federated learning verification strategy that assigns weights to different clients collecting fault sample data in the IIoT, rather than weighting different fault classes conventionally. Their method successfully achieved excellent performance in the inter-turn short circuit fault of permanent magnet synchronous motors. The method of adjusting the logit boundary has also been applied to the field of fault diagnosis. Liu et al. proposed a soft-boundary HyperDisk (HD) tensor machine for the intelligent fault diagnosis of rotating machinery. By introducing weights, the HD boundary details they proposed were enhanced to obtain a soft boundary, which can better approximate the true class region and improve the anti-interference ability against outliers and noise samples. Yan et al. designed a boundary-aware regularization method to impose significant regularization penalties on the fault data boundary in a digital-twin-assisted imbalanced fault diagnosis framework. Despite these advances, there are still several challenges in applying these strategies to the HDMS fault diagnosis task with a long-tailed distribution.

[0004] First, it is quite challenging to directly adopt resampling methods to generate a balanced dataset for HDMS fault diagnosis. Most existing studies are based on empirical selection and manual parameter adjustment. These methods require a large amount of expertise and result in a time-consuming debugging process. Second, directly applying weight-based methods to balance the classifier's recognition of various HDMS fault types has flaws. Traditional methods only consider sample frequency when assigning weights. This unreasonable weight assignment causes the classifier to overemphasize rare fault categories, thus reducing the overall recognition accuracy. Third, the classic method of adjusting the boundary through logistic regression is static and cannot adapt to changes in the classifier. Fourth, in actual industrial HDMS fault scenarios, fault types often resemble each other. As Figure 1 shown, categories (a) and (b) both have spikes at the leftmost end; categories (c) and (d) are almost indistinguishable; category (e) is highly similar to categories (c) and (d) at the highest peak. However, few algorithms take these similarities into account, leading to confusion in fault type classification. Summary of the Invention

[0005] The object of the present invention is to provide an adaptive feature adjustment and optimization method for HDMS wear fault diagnosis. This method broadens the decision boundary by expanding the clustering of training sample features; compresses the clustering of test samples, reducing the possibility of actual sample points crossing the decision boundary; then uses a pre-trained model to evaluate the similarity between different fault types; adaptively adjusts the boundaries of sample points based on the similarity according to the weight parameters of the classifier, and assigns different weights to each fault type according to the sample frequency, enhancing the learning effect of the classifier on various fault types; to improve its long-tail learning ability.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] An adaptive feature adjustment and optimization method for HDMS wear fault diagnosis, comprising the following steps:

[0008] Step 1: Pre-train the HDMS wear fault diagnosis network model on the training set using the cross-entropy loss function and extract the original features; then use t-SNE to calculate the similarity between categories, and construct a similarity window according to the similarity. The size of the similarity window is S;

[0009] Step 2: Amplify the original features using feature cluster compression; the amplification factor C is the number of categories, γ is a hyperparameter controlling the compression degree, and c ∈ (1, 2,..., C) is the category;

[0010] Step 3: Introduce the similarity window boundary value into the cross-entropy loss function to form the SWM loss function, and then adjust the SWM loss function through sample weights to form the SWRM loss function;

[0011] Step 4: Initialize the HDMS wear fault diagnosis network model, and set the total number of training rounds T, the SWR loss enabled round T0, and the FCC enabled round T1. Gradually optimize the model parameters of the HDMS wear fault diagnosis model through multiple rounds of iteration;

[0012] Gradually optimizing the parameters of the HDMS wear fault diagnosis model through multiple rounds of iteration includes:

[0013] Step 4-1: Randomly shuffle the data of the samples in the training set before each training to prevent the model from overfitting a specific order, and traverse the training round t;

[0014] Step 4-2: Judge the size of the current round t and the SWR loss enabled round T0. When t < T0, calculate the similarity window boundary m c , and start the SWM loss function; when t ≥ T0, calculate the similarity window reweighting w c , and then recalculate the similarity window boundary m c , and start the SWRM loss function;

[0015] Step 4-3: Judge the size of the current round t and the FCC enabled round T1. When t < T1, the original features extracted by the convolutional layer are used as product features. When t ≥ T1, the extracted original features are subjected to feature amplification processing to generate product features; Calculate the output and net output of the l-th convolutional layer through the forward propagation of the product features, z (l) = W (l) a (l-1) + b (l) a (l) = f l (z (l) ), z (l) is the net output of the l-th layer, W (l) is the weight matrix from the (l - 1)-th layer to the l-th layer, a (l-1) and a (l) are the outputs of the (l - 1)-th layer and the l-th layer, b (l) is the bias from the (l - 1)-th layer to the l-th layer, and f l (.) is the activation function of the l-th layer;

[0016] Step 4-4: Use the corresponding loss function to calculate the gradient and backpropagate to update the model parameters.

[0017] Furthermore, the HDMS wear fault diagnosis network model adopts one of VGG11, ResNet18, ResNet34, ResNet50, ResNext50, and DenseNet121.

[0018] Furthermore, the SWM loss function is constructed based on calculating the similarity window boundary by the classifier weights and the categories within the similarity window;

[0019] The similarity window boundary W k and W c are the classifier weights of categories k and c respectively, β is a hyperparameter controlling the category similarity, and S is the size of the similarity window;

[0020] The SWM loss function

[0021] Furthermore, the similarity window reweighting n k 、n c are the number of samples in categories k and c, and n max is the maximum number of samples in all categories;

[0022] The SWRM loss function is constructed by adjusting the SWR loss function with sample weights

[0023]

[0024] The present invention significantly improves the recognition accuracy of the tail fault categories through the feature space recalibration and similarity constraint mechanism. Description of the Drawings

[0025] Figure 1 is a signal sample diagram of five types of HDMS wear fault types.

[0026] Figure 2 is a schematic diagram of the AFAO method for diagnosing HDMS wear faults.

[0027] Figure 3 is a feature visualization diagram after feature compression.

[0028] Figure 4 is a schematic diagram of the similarity window operation mechanism.

[0029] Figure 5 is a radar chart of the comparison of all indicators of each long-tail strategy.

[0030] Figure 6 is a line chart of the influence of hyperparameters γ, β, and S on the accuracy of the model.

[0031] Figure 7 It is a line chart showing the influence of hyperparameters T0 and T1 on the accuracy of the model. Specific implementation manners

[0032] As Figure 2 shown, an adaptive feature adjustment and optimization method for HDMS wear fault diagnosis provided in this embodiment includes the following steps:

[0033] Step 1: Pre-train the HDMS wear fault diagnosis model on the training set using the cross-entropy loss function, extract the original features in the training set, then use t-SNE to calculate the similarity between categories, and construct a similarity window according to the similarity;

[0034] The training set is a long-tail distributed HDMS fault data set, including head (high-frequency faults) and tail (sparse faults) categories. The training set D = {(x i , c)|i ∈ (1, 2,..., N), c ∈ (1, 2,..., C)}, where x i is the i-th sample in the training set, c is the category corresponding to x i , N is the total number of samples in the training set, and C is the total number of categories.

[0035] The cross-entropy loss function a i is the predicted value of x i , a i,j is the j-th component of the predicted value a i , j ∈ {1, 2,..., C}; a i,c is the component corresponding to the predicted value category c of x i .

[0036] The HDMS wear fault diagnosis model adopts one of VGG11, ResNet18, ResNet34, ResNet50, ResNext50, and DenseNet121.

[0037] For each category c, traverse all training samples, extract features through the pre-trained model, then calculate the mean of all features to obtain the prototype features; then use t-SNE to project all the prototype features into a one-dimensional space, generating C points on the number axis. Categories with high similarity will be closer on this axis. As Figure 3 shown, it focuses on retaining the local structure and is suitable for visualizing the clustering relationship of high-dimensional data.

[0038] Pre-training and similarity evaluation provide structural guidance for subsequent feature compression (FCC) and dynamic weighting (SWRM) by extracting prototype features and quantifying category similarity, effectively solving the problems of category confusion and sparsity in long-tail data.

[0039] Step 2: Use feature cluster compression to amplify the original features to address the sparsity problem of tail class features in long-tail data. By amplifying the training features, the decision boundary is extended, making the original features more compact during testing and reducing the misclassification probability.

[0040] The amplified feature f i M = τ c × f i O , f i O is the original feature; τ c is the amplification factor, representing the amplification multiple corresponding to the class c to which the sample x i belongs.

[0041] After extracting sample features in the convolutional layer, the clusters in the N-dimensional space become sparser. From the opposite perspective, the distance between the original feature points is reduced by τ times, and compared with the multiplied features, the density of the original features increases by τ to the power of N. After training on these features multiplied by τ, the classifier maps the samples to clusters with a larger spatial hypervolume in the feature space and simultaneously draws a more accurate decision hyperplane. As Figure 3 shown, the visualization of the original features and the multiplied features for class 0 and class 1. Among them, 3(a) are the original sample points of class 0 and class 1, and the distribution of class 1 is sparser than that of class 0; 3(b) The multiplied features can broaden the decision boundary during the training phase, making the boundary clearer. 3(c) Compared with the multiplied features, the original features seem to be compressed, making it more difficult for them to cross the decision boundary drawn by the multiplied features. The original features are used to test the performance of the final model, while the features multiplied by τ are only input into the classifier during training to train its parameters. We can easily observe that the clusters of the features multiplied by τ in the feature space are broadened, thus making the decision boundary clearer.

[0042] This embodiment uses arithmetic compression to calculate the amplification factor C is the number of classes, γ is a hyperparameter controlling the compression degree, and c ∈ (1, 2,..., C) is the index of the class.

[0043] By dynamically amplifying the tail features, forcing the classifier to learn a more robust decision boundary during training, thereby enhancing the discrimination ability of sparse classes during testing and balancing the misclassification region and feature density, is a key technical innovation in long-tail learning.

[0044] Step 3: Introduce a similarity window boundary value into the cross-entropy loss function to form the SWM loss function, and then adjust the cross-entropy loss function through sample weights to form the SWR loss function. Combine the SWM loss function and the SWR loss function into the SWRM loss function;

[0045] Calculate the similarity window boundary based on the classifier weights and the classes within the similarity window. The similarity window boundary W k and W c are the classifier weights for classes k and c respectively, β is a hyperparameter for controlling class similarity, and S is the size of the similarity window.

[0046] In this embodiment, the ability of the model to distinguish between common fault types and rare fault types is enhanced by adjusting the prediction value; the SWM loss function reflects the distribution-aware boundary value loss; a fixed boundary is not applicable to HDMS diagnosis because the classifier lacks adaptability, and in this embodiment, the SWM loss function can be adaptively adjusted to encourage learning with high discrimination ability.

[0047] The SWM loss function introduces the similarity window boundary m c into the cross-entropy loss to enhance the decision-making interval for tail classes.

[0048] Combine the class label frequency with the similarity window to assign different weights to each class. The similarity window reweighting n k 、n c are the number of samples in classes k and c, and n max is the maximum number of samples in all classes.

[0049] By adjusting the SWR loss function with sample weights, the imbalance between majority classes and minority classes can be alleviated, and the model can be better trained by paying more attention to minority classes, so as to distinguish different fault types; combining the above, the final SWRM loss function is obtained in this embodiment

[0050] The SWRM loss function provides a larger margin for tail classifiers with smaller weight norms, and a larger margin means better generalization ability. It assigns smaller weights to classes that are easily classified correctly and larger weights to classes that are difficult to classify, enabling the model to pay more attention to minority classes. In addition, the adaptive margin calculation for dynamic classifiers can precisely control the distance between each class and the decision boundary, thereby further improving performance.

[0051] Step 4: Initialize the model parameters of the HDMS wear fault diagnosis model, and set the hyperparameters γ for controlling the compression degree, the similarity window size S, the hyperparameter β for class similarity, the total number of training rounds T, the round T0 when the SWR loss is enabled, and the round T1 when the FCC is enabled; gradually optimize the model parameters of the HDMS wear fault diagnosis model through multiple rounds of iteration; the model parameters mainly include the weight W and bias b of the classifier.

[0052] The steps for gradually optimizing the parameters of the HDMS wear fault diagnosis model through multiple rounds of iteration include:

[0053] Step 3-1: Randomly shuffle the data of the samples in the training set before each training to prevent the model from overfitting a specific order, and traverse the training round t;

[0054] Step 3-2: Judge the size relationship between the current round t and the round T0 when the SWR loss is enabled. When t < T0, calculate the similarity window boundary m c , and start the SWM loss function; when t ≥ T0, calculate the similarity window reweighting w c , then recalculate the similarity window boundary m c , and start the SWRM loss function;

[0055] Step 3-3: Judge the size relationship between the current round t and the round T1 when the FCC is enabled. When t < T1, the original features extracted by the convolutional layer are used as product features. When t ≥ T1, the extracted original features are subjected to feature amplification processing to generate product features; calculate the output and net output z of the l-th convolutional layer through forward propagation of the product features (l) = W (l) a (l-1) + b (l) , a (l) = f(z (l) ), z (l) is the net output of the l-th layer, W (l) is the weight matrix from the (l - 1)-th layer to the l-th layer, a (l-1) , a (l) are the outputs of the (l - 1)-th layer and the l-th layer, b (l) is the bias from the (l - 1)-th layer to the l-th layer, and f l (.) is the activation function of the l-th layer;

[0056] Step 3-4: Use the corresponding loss function to calculate the gradient and perform backpropagation to update the model parameters. Starting from the loss value, calculate the gradients of each parameter; use an optimizer (such as SGD, Adam) to update the parameters according to the gradients.

[0057] To simulate HDMS faults in a real industrial scenario and collect five HDMS operating states: normal, 1 radial scratch, 2 radial scratches, 4 radial scratches, and 9 pits. The sample sizes are 3333 (normal), 1054 (1 scratch), 333 (2 scratches), 593 (4 scratches), and 1894 (9 pits) respectively, and the length of each sample is 1024. The ratio of the largest class sample size to the smallest class sample size is 10, i.e., the imbalance factor (IF) is 10. According to the sample size, we reorder the classes as 0 - 4.

[0058] The network models adopted include: VGG11, ResNet18, ResNet34, ResNet50

[19] , ResNext50, DenseNet12; all networks are trained for 100 epochs with an initial learning rate of 0.01, except ResNext50 which uses 0.1. The learning rate decays by 0.1 at the 50th, 80th, and 90th epochs, and pre-training is carried out in the first five epochs. The batch size is 64, and the optimization is Stochastic Gradient Descent (SGD) with a momentum of 0.9 and a weight decay of 0.0005. The random seed is set to 42. The experiments are conducted using the PyTorch toolbox on a server equipped with an NVIDIA RTX4060 GPU (the PC method runs on an NVIDIA RTX4090 GPU). The Top-1 accuracy is used to evaluate the fault diagnosis performance of the model. To obtain the best results, the hyperparameters of FCC and SWRM are set to different values through experiments.

[0059] Table 1 Hyperparameters of the AFAO method

[0060]

[0061] To evaluate the effectiveness and competitiveness of the method (AFAO) proposed in this embodiment, its performance is compared with a variety of long-tail learning methods, namely: (1) Cross-Entropy (CE) loss; (2) Focal loss; (3) Balanced Meta Softmax Cross-Entropy (BSCE); (4) Class-Balanced Cross-Entropy (CBCE); (5) Class-Balanced Focal Loss (CBF); (6) Cyclic Focal Loss (CF); (7) Probability Contrastive Learning (PC). The comparison results of various methods on different backbone networks are shown in Table 2. The AFAO method proposed in this embodiment shows the best performance on all backbone networks in the long-tail high-dimensional multi-label fault diagnosis task.

[0062] Table 2 Comparison results of various methods on different backbone networks

[0063]

[0064]

[0065] As Figure 5 shown, under each backbone network, the AFAO method is almost closest to the periphery, indicating its strong overall competitiveness. Except on ResNext50, it has the highest recall rate, meaning that the AFAO algorithm has a lower probability of missing detections in high-dimensional multimodal fault diagnosis of long tails. At the same time, the high precision ensures that our method will not over-monitor faults and can accurately determine the fault type, which helps subsequent maintenance work. The high F1 score and AUROC also support this point.

[0066] Given that industrial signals usually contain noise, which will contaminate the original data and make high-dimensional multi-scale (HDMS) fault diagnosis difficult, the ability of the model to resist noise interference is evaluated through experiments; Gaussian noise (mean 0, variance 0.1) is added to each item in the original dataset, and various long-tail learning methods are evaluated. The backbone network uses ResNet18, and the hyperparameters are the same as those in Table 1. The results are shown in Table 3.

[0067] Table 3 Comparison of the AFAO method with other long-tail methods in learning with noisy signals

[0068]

[0069]

[0070] After introducing Gaussian noise, the accuracy of all methods has decreased. However, the AFAO method of this embodiment still has the highest accuracy, recall rate, and F1 value. This verifies that in the actual HDMS industrial scenario, in the face of noisy detection samples, the method of this embodiment can more comprehensively check the fault situation.

[0071] This embodiment proposes a novel adaptive feature adjustment and optimization method for high-dimensional multi-scale (HDMS) wear fault diagnosis. First, FCC compresses the original features in the dimensional space to increase the clustering density to reduce the possibility of tail classes crossing the decision boundary. Second, the SWM loss function adjusts the distance of sample points to the decision boundary to further reduce this possibility. Third, the SWM loss function considers the similarity of fault classes and assigns reasonable class weights to form the final SWRM loss function, further improving the model performance. Applying it to the long-tail HDMS fault dataset, the experimental results confirm that the intelligent diagnosis performance has been effectively improved, demonstrating its superiority in HDMS wear fault diagnosis.

[0072] The impact of the five hyperparameters involved on the HDMS fault identification task are γ, S, β, T1, and T0. Only three models, the backbone network VGG11, ResNet18, and DenseNet121, are applied. When studying specific hyperparameters, the remaining hyperparameters are set to the values in Table 1; only the accuracy (%) of each model is evaluated to compare their performance.

[0073] The influence of hyperparameter γ: Increasing γ will also increase the misclassification area; the values of γ in {0.0001, 0.001, 0.0015, 0.01, 0.015, 0.1} were tested to examine the influence of the error area size on the model performance. Figure 6 (a) It can be seen that the optimal γ value varies depending on the backbone network. Contrary to theoretical expectations, very small γ values can reduce accuracy. This is because a smaller degree of compression will cause fewer boundary points to return to the decision boundary, and the harm caused by this will outweigh the benefit of reducing the misclassified area. Finally, it is found that γ∈{0.01,0.015,0.1} are three good choices for the model.

[0074] The influence of the hyperparameter S: how many categories should the similarity window contain; the options of S are limited because the similarity window should cover 2 categories and no more than all categories. The cases with window sizes of 2, 3, 4, and 5 were investigated. The results show that for DenseNet121 and ResNet18, the method of this embodiment is robust to this hyperparameter, such as Figure 6 (b). However, VGG11 is particularly sensitive to this hyperparameter.

[0075] Effect of hyperparameter β: Similar to hyperparameter γ, we set β∈{0.2,0.5,1,1.2,1.5,2,2.2,2.5} to study its effect. Figure 6 It can be seen that the curves of these hyperparameters have similar shapes. As the value of β increases, the model performance improves slightly at first, but then starts to decline after reaching a certain point. This indicates that too small or too large a value of β has limited performance improvement, which may be due to under-learning or over-learning of tail fault types for which knowledge is scarce.

[0076] When to introduce FCC in the training stage; to study the impact of introducing FCC at different training stages. The three backbone networks were trained for 100 rounds, and FCC was introduced at the 0th, 10th, 20th, 30th, 40th and 50th rounds. Figure 7 The results in (a) show that introducing FCC earlier leads to better performance, while introducing it later makes the model performance worse.

[0077] When to start using the SWR strategy; determine the best time to introduce the reweighting strategy by adjusting T0. Check this by setting the T0 values from the 30th round to the 90th round, with an increment of 10. Figure 7 The results in (b) show that introducing SWR too early harms performance, indicating that early weight allocation weakens the representation learning of the head classes.

[0078] The above are only the preferred embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any modification and replacement based on the technical solutions and inventive concepts provided by the present invention should be covered within the protection scope of the present invention.

Claims

1. An adaptive feature adjustment and optimization method for HDMS wear fault diagnosis, characterized in that It includes the following steps: Step 1: Pre-train the HDMS wear fault diagnosis network model on the training set using the cross-entropy loss function, and extract the original features; then use t-SNE to calculate the similarity between classes, and construct a similarity window according to the similarity. The size of the similarity window is S; Step 2: Amplify the original features using feature cluster compression; amplification factor C is the number of classes, γ is a hyperparameter controlling the degree of compression, and c ∈ (1, 2, …, C) is the class Step 3: Construct the SWM loss function by introducing the similarity window boundary value into the cross-entropy loss function, and then construct the SWRM loss function by adjusting the SWM loss function with sample weights; Step 4: Initialize the HDMS wear fault diagnosis network model, and set the total number of training rounds T, the SWR loss enabled round T0, and the FCC enabled round T1. Gradually optimize the model parameters of the HDMS wear fault diagnosis model through multiple rounds of iteration; Gradually optimizing the parameters of the HDMS wear fault diagnosis model through multiple rounds of iteration includes: Step 4-1: Randomly shuffle the data of the samples in the training set before each training to prevent the model from overfitting a specific order, and traverse the training round t; Step 4-2: Determine the size relationship between the current round t and the SWR loss enabling round T0. When t < T0, calculate the similarity window boundary m c , and activate the SWM loss function; when t ≥ T0, calculate the similarity window reweighting w c , and then recalculate the similarity window boundary m c , and activate the SWRM loss function; Step 4-3: Determine the magnitude relationship between the current round t and the FCC activation round T1. When t < T1, the original features extracted by the convolutional layer are used as product features. When t ≥ T1, the extracted original features are subjected to feature amplification processing to generate product features; the output and net output of the l-th convolutional layer are calculated by forward propagation through the product features, z (l) = W (l) a (l-1) + b (l) , a (l) = f l (z (l) ), z (l) is the net output of the l-th layer, W (l) is the weight matrix from the (l-1)-th layer to the l-th layer, a (l-1) , a (l) are the outputs of the (l-1)-th layer and the l-th layer, b (l) is the bias from the (l-1)-th layer to the l-th layer, f l (.) is the activation function of the l-th layer; Step 4-4: Use the corresponding loss function to calculate the gradient and backpropagate to update the model parameters.

2. An adaptive feature adjustment and optimization method for HDMS wear fault diagnosis according to claim 1, characterized in that The HDMS wear fault diagnosis network model adopts one of VGG11, ResNet18, ResNet34, ResNet50, ResNext50, DenseNet121.

3. An adaptive feature adjustment and optimization method for HDMS wear fault diagnosis according to claim 1, characterized in that, The SWM loss function is constructed based on the classifier weights and the similarity of classes within the similarity window; The similarity window boundary W k and W c are the classifier weights for classes k and c respectively, β is a hyperparameter controlling class similarity, and S is the similarity window size; The SWM loss function 4. An adaptive feature adjustment and optimization method for HDMS wear fault diagnosis according to claim 3, characterized in that The similarity window reweighting n k and n c are the number of samples in classes k and c, and n max is the maximum number of samples in all classes; Construct the SWRM loss function by adjusting the SWR loss function with sample weights

Citation Information

Patent Citations

  • Gear fault diagnosis method based on local time sequence self-similarity and bag-of-word model

    CN115165342A

  • Fluid dynamic pressure sealing ring wear fault detection method based on dynamic graph residual convolution

    CN115600138A

  • Unmanned aerial vehicle control surface intelligent fault diagnosis method and device under unbalanced small sample

    CN117171681A

  • Incremental device fault diagnosis method

    WO2024060381A1