Numerical control machine tool feeding system fault diagnosis method based on labeling deviation immune loss

By employing an embedded fault diagnosis framework that incorporates annotation deviation immunity loss and modal channel attention modules in the CNC machine tool feed system, the problem of decreased diagnostic accuracy caused by annotation deviations in historical fault data is solved, achieving highly efficient fault diagnosis.

CN120951044APending Publication Date: 2025-11-14SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511072032.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Occasional labeling deviations exist in the historical fault data of CNC machine tool feed systems, which significantly reduces the accuracy of intelligent fault diagnosis methods.

Method used

A fault diagnosis method based on label bias immune loss is adopted. Through modal channel attention module and embedded fault diagnosis framework, the gradient calculation of samples during training is adaptively adjusted by label bias immune loss function. Combined with noise reduction interactive sample processing strategy, an embedded fault diagnosis model is constructed.

Benefits of technology

It significantly improves the accuracy of fault diagnosis under labeled deviation data, realizes the global optimal solution for fault diagnosis of the feed system, and enhances the flexibility and diagnostic accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951044A_ABST
    Figure CN120951044A_ABST
Patent Text Reader

Abstract

The invention discloses a numerical control machine tool feeding system fault diagnosis method based on label deviation immune loss, and relates to the field of fault diagnosis, comprising the following steps: collecting operation data of a numerical control machine tool feeding system, and processing the data to form a data set; slicing the data in the data set, and inputting the sliced data into a modal channel attention module for feature fusion and weight distribution; constructing an embedded fault diagnosis framework, and obtaining data used for training the embedded fault diagnosis framework by using a noise-free interaction sample processing strategy; the standard cross entropy loss is replaced by the labeled deviation immune loss, and the embedded fault diagnosis framework is trained; and inputting the new sample into a trained fault diagnosis model in the embedded fault diagnosis framework, and carrying out fault diagnosis on the numerical control machine tool feeding system. According to the method, the deviation immune loss is marked, the influence of the sample on gradient calculation in the training process can be adaptively adjusted according to the one-hot code of the actual label of the sample, and the accuracy of fault diagnosis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault diagnosis, and in particular to a fault diagnosis method for CNC machine tool feed systems based on annotation deviation immune loss. Background Technology

[0002] With the intelligent upgrading of the manufacturing industry, CNC machine tools, as the core of high-end equipment, directly affect product quality and production efficiency through their machining accuracy and reliability. The feed system, as a key moving component of CNC machine tools, endures high loads and high-frequency movements over long periods, making it prone to problems such as mechanical wear, guide rail loosening, and servo motor failure. If these faults are not diagnosed in time, they can lead to parts being scrapped due to exceeding tolerances, or even equipment damage and safety accidents. Therefore, developing efficient fault diagnosis methods is of great significance for ensuring the stable operation of CNC machine tools.

[0003] Traditional fault diagnosis relies primarily on expert experience and regular maintenance, which is inefficient and ill-suited for handling sudden failures. From a technological development perspective, early fault diagnosis relied mainly on physical models and signal processing techniques. While model-based methods have clear physical meaning, they require extremely high accuracy in modeling complex systems, limiting their practical application. This approach necessitates establishing an accurate physical or mathematical model and detecting faults by comparing the residuals between the actual output and the model output. Typical techniques include parameter estimation methods for real-time identification of system parameters.

[0004] Signal processing methods extract features through time-frequency analysis, which reduces the reliance on precise models to some extent, but still depends on signal feature extraction and requires manual intervention. This method collects physical signals such as vibration, noise, and temperature during equipment operation and extracts feature parameters using techniques such as spectral analysis, wavelet transform, and correlation functions. For example, abnormal high-frequency components of bearing vibration signals can be used to diagnose wear faults, or harmonic analysis of motor current signals can be used to detect broken rotor bars.

[0005] In recent years, with the development of sensing technology and artificial intelligence, data-driven intelligent diagnostic methods have gradually become a research hotspot. These methods analyze multi-source signals such as vibration, current, and temperature during equipment operation, and use machine learning algorithms to automatically identify fault modes, greatly improving the real-time performance and accuracy of diagnosis. Examples include convolutional neural networks and Transformer network architectures.

[0006] However, data collected from industrial sites often suffers from quality issues, which in turn affects the performance of intelligent diagnostic algorithms. Labeling bias is one of the most prominent challenges. In actual production environments, fault data labeling is typically done by field engineers. Due to a lack of unified labeling standards and differences in professional knowledge, different personnel may have differing opinions on the same fault. For example, vibration signals from early bearing wear may have blurred boundaries with normal conditions, easily leading to mislabeling. Furthermore, equipment operating under normal conditions for extended periods results in a scarcity of fault samples, causing severe class imbalance in the dataset.

[0007] Noise labels and sample imbalance caused by labeling bias significantly reduce the performance of intelligent diagnostic models. Noise labels mislead the model into learning incorrect feature associations, while sample imbalance causes the model to favor the majority class, leading to missed detections of rare faults. Although existing research has proposed some improvement methods, their performance in real-world industrial scenarios remains unsatisfactory. Especially in high-noise environments, traditional methods often struggle to balance noise suppression and effective feature learning. Adding to the complexity, the operating conditions of CNC machine tools are highly variable; the same fault may manifest completely differently under different loads and speeds. This dynamic characteristic further increases the difficulty of data labeling, making the labeling bias problem even more prominent. Current research on noise labeling in industrial data mainly follows three technical routes. Sample discovery and filtering methods identify and filter potential noise samples that significantly deviate from normal samples by analyzing the distribution patterns of the sample feature space, thereby achieving data cleaning. Model structure optimization methods attempt to establish a transition matrix from noise labels to true labels, automatically correcting erroneous label information during training through probabilistic modeling. The adaptive loss function method starts from the optimization objective and redesigns the form of the loss function so that the model assigns smaller gradient update weights to samples that may contain noise, thereby reducing the negative impact of mislabels on model training.

[0008] Therefore, those skilled in the art are dedicated to developing a fault diagnosis method for CNC machine tool feed systems based on annotation deviation immune loss. Summary of the Invention

[0009] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is that the historical fault data of the feeding system inevitably has occasional labeling deviations, which leads to a significant decrease in the accuracy of the intelligent fault diagnosis method.

[0010] To achieve the above objectives, this invention provides a fault diagnosis method for CNC machine tool feed systems based on annotation deviation immune loss, the method comprising the following steps:

[0011] S101: Collect the operating data of the CNC machine tool feed system, process the operating data to form a dataset, and the dataset includes real-time data and historical fault data;

[0012] S103: Slice the data in the dataset, and input the sliced ​​data into the modal channel attention module for feature fusion and weight allocation;

[0013] S105: Construct an embedded fault diagnosis framework, which includes a fault diagnosis model, and use a noise reduction interactive sample processing strategy to obtain data from the dataset for training the embedded fault diagnosis framework.

[0014] S107: The embedded fault diagnosis framework is trained using label bias immune loss instead of standard cross-entropy loss;

[0015] S109: Input the new sample into the fault diagnosis model trained in the embedded fault diagnosis framework to perform fault diagnosis on the CNC machine tool feed system.

[0016] Further, in step S101, the dataset includes the CNC system current signal, the tracking error signal, and the vibration signal collected by the vibration sensor, wherein the vibration signal includes vibration signals from 4 points, and the dataset is organized into 6 channels of raw input data.

[0017] Further, in step S103, the modal channel attention module, based on the channel attention module, introduces initialization weights and knowledge regularization terms according to prior knowledge to guide the weight allocation of each modality. The modal channel attention module is as follows:

[0018]

[0019] O j =[g1,g2,…,g j ]

[0020] Among them, O j Let O′ be a channel signal matrix consisting of j channel signals. j O″j is the intermediate matrix after channel reorganization and weight allocation of samples based on prior knowledge, and g is the output matrix after residual concatenation of the intermediate matrix and the original matrix. j Let g be the signal of the j-th group of the original matrix. j ′ represents the signal of the j-th group in the intermediate matrix, g j " is the signal of the j-th group in the output matrix. This is the weight matrix.

[0021] Further, step S103 includes the following sub-steps:

[0022] S1031: Slicing: Slice the long-time sequence of the input 6 modalities to obtain multiple samples of shape [6, 2000].

[0023] S1032: Channel Reorganization: Based on prior knowledge, the sample is reorganized into channels so that the i-th modal channel is sensitive to the i-th fault mode;

[0024] S1033: Weight Allocation: Multiply the recombined samples by the initial weight matrix W to obtain the intermediate matrix O′. j ;

[0025] S1034: Residual Linkage: Connect O′ j Add the recombined sample O j The output O″ of the modal channel attention module is obtained. j .

[0026] Further, in step S105, the fault diagnosis model in the embedded fault diagnosis framework adopts a deep learning-based fault diagnosis model. The architecture of the embedded fault diagnosis framework includes an in-situ data acquisition module, a noise reduction interactive sample processing strategy, an embedded model training module, and a fault diagnosis task module.

[0027] The in-situ data acquisition module collects multiple modal data from the CNC machine tool feed system, and after slicing, sends them to the modal channel attention module.

[0028] The noise reduction interactive sample processing strategy uses the dataset with labeling bias as the training set and divides the accurately labeled dataset into a validation set and a test set.

[0029] The embedded model training module selects any advanced fault diagnosis model as the backbone network and uses label bias immune loss instead of standard cross-entropy loss to train the fault diagnosis model in the embedded fault diagnosis framework.

[0030] The fault diagnosis task module uses the trained embedded fault diagnosis framework for new fault diagnosis tasks, and performs fault diagnosis tasks for CNC machine tool feed systems with unknown faults.

[0031] Furthermore, the fault diagnosis model in the embedded fault diagnosis framework is configured such that: the number of output layers and fault modes of the fault diagnosis model are consistent; the length of the input layer of the fault diagnosis model is consistent with the length of the slice sample; and the loss function used by the fault diagnosis model is the label bias immune loss function.

[0032] Furthermore, in the labeling bias immune loss function, adaptive coefficients are added to both the cross-entropy and inverse cross-entropy. These adaptive coefficients are determined by a custom normalized exponential function and a custom confidence factor. The labeling bias immune loss function is as follows:

[0033]

[0034] y = [y1, y2, ..., y i ,…y k ]

[0035]

[0036] Among them, L USCE Let y be the label bias immune loss function, and y be the actual label vector. i For the actual label of the i-th fault mode in a single sample, To predict the label vector, Here is the predicted label for the i-th fault mode in a single sample, where k is the total number of fault modes, i is the fault mode number, and s is the predicted label for the i-th fault mode in a single sample. i Let W be the custom normalization index for the i-th failure mode, C be the custom confidence factor, and W be the custom confidence index. ii Let be the weight value of the i-th fault mode on the i-th channel signal.

[0037] Furthermore, the gradient calculation of the label bias immune loss function during the training process is as follows:

[0038]

[0039] in, Let y be the gradient value of the label bias immune loss function, and y be the actual label vector. i For the actual label of the i-th fault mode in a single sample, To predict the label vector, Here is the predicted label for the i-th fault mode in a single sample, where k is the total number of fault modes, i is the fault mode number, and s is the predicted label for the i-th fault mode in a single sample. i Let W be the weight of the encoding position for the i-th fault mode, C be a custom confidence factor, and W be the weight of the encoding position for the i-th fault mode. ii Let be the weight value of the i-th fault mode on the i-th channel signal.

[0040] Furthermore, the customized normalized exponential function is used to calculate the one-hot encoded label confidence of a single sample and its impact on gradient calculation. The customized normalized exponential function is as follows:

[0041]

[0042] Where s is the weight of the encoding position, U-Softmax is a custom normalized exponential function, η is the scaling factor, and yi is the actual label for a single sample, and k is the total number of failure modes.

[0043] Furthermore, the customized confidence factor, by calculating the zero-norm value of the one-hot encoding in a batch, determines the confidence of the sample itself and its impact on gradient calculation. The customized confidence factor is calculated as follows:

[0044]

[0045] Among them, C i Let α be the confidence factor for the i-th failure mode, β be the lower bound of the confidence level, β be the hyperparameter that adjusts the correlation between the confidence level and the 0-norm, and N0 be the 0-norm of the sample.

[0046] In a preferred embodiment of the present invention, compared with the prior art, the present invention has the following beneficial effects:

[0047] 1. The labeling bias immune loss designed in this invention can adaptively adjust the negative impact of samples on gradient calculation during training based on the one-hot encoding of the actual labels of the samples, which significantly improves the accuracy of fault diagnosis methods under labeled bias data.

[0048] 2. The modal channel attention module designed in this invention is suitable for fault diagnosis tasks of CNC machine tool feed systems. It is plug-and-play and obtains the global optimal solution for the weight allocation of each modality, thereby significantly improving the accuracy of feed system fault diagnosis.

[0049] 3. The embedded fault diagnosis model architecture designed in this invention is compatible with existing advanced fault diagnosis models. The model can be used plug-and-play, enabling existing models to be compatible with annotation bias immune loss function and modal channel attention module, which greatly improves the fault diagnosis accuracy of the model and ensures the flexibility of model selection. It has strong flexibility and ease of use.

[0050] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description

[0051] Figure 1 This is a flowchart of a fault diagnosis method for a CNC machine tool feed system according to a preferred embodiment of the present invention;

[0052] Figure 2 This is a schematic diagram of the modal channel attention module according to a preferred embodiment of the present invention;

[0053] Figure 3 This is a schematic diagram illustrating the effect of weight matrix training with and without modal channel attention modules according to a preferred embodiment of the present invention;

[0054] Figure 4 This is a schematic diagram of an embedded fault diagnosis model architecture according to a preferred embodiment of the present invention;

[0055] Figure 5 This is a schematic diagram illustrating the principle of the annotation bias immune loss function according to a preferred embodiment of the present invention;

[0056] Figure 6 This is a schematic diagram comparing the training effects of annotation bias immune loss and standard cross-entropy loss according to a preferred embodiment of the present invention. Detailed Implementation

[0057] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0058] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.

[0059] like Figure 1 As shown, in addressing the problem that historical fault data of CNC machine tool feed systems inevitably contains occasional labeling biases, leading to a significant decrease in the accuracy of intelligent fault diagnosis methods, this invention provides a fault diagnosis method for CNC machine tool feed systems based on labeling bias immune loss. By using labeling bias immune loss, the influence of samples on gradient calculation during training can be adaptively adjusted according to the one-hot encoding of the actual labels of the samples, thereby improving the accuracy of fault diagnosis methods under labeled bias data.

[0060] The method provided in this embodiment includes the following steps:

[0061] S101: Collects the operating data of the CNC machine tool feed system, processes the operating data to form a dataset, and the processed dataset includes real-time data and historical fault data.

[0062] When performing online data acquisition on the CNC machine tool feed system, necessary historical fault datasets are obtained. The acquired datasets include CNC system current signals, tracking error signals, and vibration signals collected by vibration sensors. Specifically, the vibration signals were collected from four points using vibration sensors, and the acquired datasets were organized into 6-channel raw input data.

[0063] S103: Slice the data in the dataset and input the sliced ​​data into the modal channel attention module for feature fusion and weight allocation.

[0064] The modal channel attention module provided in this embodiment of the invention, based on the channel attention module in the prior art, introduces initial weights and knowledge regularization terms according to prior knowledge to guide the weight allocation of each modality.

[0065] The modal channel attention module can be represented as:

[0066]

[0067] O j =[g1,g2,…,g j ]

[0068] Among them, O j Let O′ be a channel signal matrix consisting of j channel signals. j O″ is the intermediate matrix after channel reorganization and weight allocation of samples based on prior knowledge. j The output matrix is ​​the result of concatenating the intermediate matrix and the original matrix residuals, g j Let g be the signal of the j-th group of the original matrix. j ′ represents the signal of the j-th group in the intermediate matrix, g j " is the signal of the j-th group in the output matrix. This is the weight matrix.

[0069] In this embodiment, the step includes the following sub-steps:

[0070] S1031: Slicing: Slice the long-time sequence of the input 6 modalities to obtain multiple samples of shape [6, 2000].

[0071] S1032: Channel Reorganization: Based on prior knowledge, the sample is reorganized into channels so that the i-th modal channel is sensitive to the i-th fault mode;

[0072] S1033: Weight Allocation: Multiply the recombined samples by the initial weight matrix W to obtain the intermediate matrix O′. j ;

[0073] S1034: Residual Linkage: Connect O′ j Add the recombined sample O j The output O″ of the modal channel attention module is obtained. j .

[0074] S105: Construct an embedded fault diagnosis framework and use a noise reduction interactive sample processing strategy to obtain data from the dataset for training the embedded fault diagnosis framework.

[0075] In this embodiment, the fault diagnosis model in the embedded fault diagnosis framework adopts a deep learning-based fault diagnosis model, which requires that the number of output layers and fault modes of the fault diagnosis model be consistent, the length of the input layer and the length of the slice samples be consistent, and the loss function used is the label bias immune loss function.

[0076] The architecture of the embedded fault diagnosis model provided in this embodiment includes an in-situ data acquisition module, a noise reduction interactive sample processing strategy, an embedded model training module, and a fault diagnosis task module.

[0077] 1) In-situ data acquisition module

[0078] The system collects various modal data from the CNC machine tool's feed system, processes them into slices, and then sends them to the modal channel attention module.

[0079] 2) Noise Reduction Interaction Sample Processing Strategy

[0080] The dataset with labeling bias is used as the training set, and the accurately labeled dataset is divided into a validation set and a test set.

[0081] 3) Embedded model training module

[0082] Any advanced fault diagnosis model can be selected as the backbone network, and the label bias immune loss can be used instead of the standard cross-entropy loss to train the fault diagnosis model in the embedded fault diagnosis framework.

[0083] 4) Fault Diagnosis Task Module

[0084] The fault diagnosis model in the trained embedded fault diagnosis framework is used for new fault diagnosis tasks to achieve fault diagnosis for CNC machine tool feed systems with unknown faults.

[0085] In this embodiment, the labeling bias immune loss function adds adaptive coefficients to both the cross-entropy and inverse cross-entropy. These adaptive coefficients are determined by a custom normalized exponential function and a custom confidence factor. The labeling bias immune loss function is as follows:

[0086]

[0087] y = [y1, y2, ..., y i ,…y k ]

[0088]

[0089] Among them, L USCE Let y be the label bias immune loss function, and y be the actual label vector. i For the actual label of the i-th fault mode in a single sample, To predict the label vector, Here is the predicted label for the i-th fault mode in a single sample, where k is the total number of fault modes, i is the fault mode number, and s is the predicted label for the i-th fault mode in a single sample. i Let W be the custom normalization index for the i-th failure mode, C be the custom confidence factor, and W be the custom confidence index. ii Let be the weight value of the i-th fault mode on the i-th channel signal.

[0090] In this embodiment, the gradient calculation of the label bias immune loss function during training is as follows:

[0091]

[0092] in, Let y be the gradient value of the label bias immune loss function, and y be the actual label vector. i For the actual label of the i-th fault mode in a single sample, To predict the label vector, Here is the predicted label for the i-th fault mode in a single sample, where k is the total number of fault modes, i is the fault mode number, and s is the predicted label for the i-th fault mode in a single sample. i Let W be the custom normalization index for the i-th failure mode, C be the custom confidence factor, and W be the custom confidence index. ii Let be the weight value of the i-th fault mode on the i-th channel signal.

[0093] A custom normalized exponential function is used to calculate the one-hot encoded label confidence of a single sample and its impact on gradient calculation.

[0094] The custom normalized exponential function is:

[0095]

[0096] Where s is the weight of the encoding position, U-Softmax is a custom normalized exponential function, η is the scaling factor, and y i is the actual label for a single sample, and k is the total number of failure modes.

[0097] Custom confidence factors determine the confidence level of a sample and its impact on gradient calculation by calculating the zero-norm value of the one-hot encoding in a batch.

[0098] The method for calculating the custom confidence factor is as follows:

[0099]

[0100] Among them, C i Let α be the confidence factor for the i-th failure mode, β be the lower bound of the confidence level, β be the hyperparameter that adjusts the correlation between the confidence level and the 0-norm, and N0 be the 0-norm of the sample.

[0101] S107: Use label bias immune loss instead of standard cross-entropy loss to train the embedded fault diagnosis model.

[0102] S109: Input the new sample into the trained embedded fault diagnosis model to perform fault diagnosis on the CNC machine tool feed system.

[0103] Compared with the prior art, the fault diagnosis method for CNC machine tool feed system based on annotation deviation immune loss provided in this embodiment of the invention has the following beneficial effects:

[0104] 1. To address the issue that historical fault data in CNC machine tool feed systems inevitably contains occasional labeling biases, leading to a significant decrease in the accuracy of intelligent fault diagnosis methods, this invention provides a labeling bias immune loss function. This function adaptively adjusts the impact of samples on gradient calculation during training based on the one-hot encoding of the actual sample labels, particularly reducing the negative impact of labeled bias samples on gradient training. Building upon the symmetric cross-entropy loss, adaptive coefficients are added to both cross-entropy and inverse cross-entropy. These adaptive coefficients are determined by a custom normalized exponential function and a custom confidence factor. Through the labeling bias immune loss function, the accuracy of fault diagnosis methods under labeled bias data is significantly improved.

[0105] 2. Addressing the issue that different modes in a CNC machine tool feed system exhibit varying sensitivities to different fault modes and influence each other, the weights of each mode during model training are prone to getting trapped in local optima. This invention provides a modal interaction channel attention module that guides the weight allocation of each mode based on prior knowledge, preventing the model training process from falling into local optima. Building upon the channel attention module, it introduces initialization weights and knowledge regularization terms based on prior knowledge. The modal interaction channel attention module obtains the globally optimal solution for the weight allocation of each mode, resulting in a globally optimal fault diagnosis rate.

[0106] 3. To address the poor generalization ability of existing intelligent fault diagnosis models, particularly their significant deterioration in performance on CNC machine tool feed systems and other equipment, this invention provides an embedded fault diagnosis model architecture. This architecture allows the use of existing advanced models while introducing a modal interaction channel attention module and a labeling bias immune loss function. Through the noise reduction interactive sample processing strategy of the modal interaction attention module, the old model can be directly replaced. The labeling bias immune loss function directly replaces the cross-entropy loss function and participates in gradient calculation. This invention achieves a significantly higher fault diagnosis accuracy for the old model embedded in the new framework than for the old model, even with labeled bias data.

[0107] The present invention will now be described in detail with reference to preferred embodiments.

[0108] The fault diagnosis method for CNC machine tool feed system based on annotation deviation immune loss provided in this embodiment of the invention includes the following steps:

[0109] Step 1: Perform online data acquisition on the CNC machine tool's feed system to obtain necessary historical fault datasets. This involves acquiring two modes of data using the CNC system: current signal and tracking error signal. Vibration signals from four points are also collected using vibration sensors, and the data is ultimately processed into a 6-channel raw input.

[0110] Step 2: Slice the input 6-channel signal and then input the sliced ​​signal into the modal channel attention module for feature fusion and weight allocation.

[0111] Based on prior knowledge, given the known sensitivity of the modes to each fault mode, the j-channel signal can be described as O j =[g1,g2,…,g j ],O′ j O″ is the intermediate matrix after channel reorganization and weight allocation of samples based on prior knowledge. j The output matrix is ​​the result of concatenating the intermediate matrix and the original matrix residuals, g j Let g be the signal of the j-th group of the original matrix. j ′ represents the signal of the j-th group in the intermediate matrix, g j " is the signal of the j-th group in the output matrix. This is the weight matrix.

[0112] Design an initial learnable weight matrix. The modal channel attention module can be described as:

[0113]

[0114] Specifically, W(a,b) represents the sensitivity of the fault mode corresponding to column b to the mode in row a. Therefore, the sum of the elements in each column of W is 1. Modes with high sensitivity will be set to a larger value γ, which is recommended to be 0.8. For example, if mode 2 is significantly more sensitive to fault mode 2 than other modes, then W(2,2) is set as the hyperparameter γ.

[0115] A diagonal regularization term is designed to guide the training of the weight matrix. When numbering fault modes, fault modes sensitive to mode 1 are designated as Fault Mode 1, and fault modes sensitive to mode 2 are designated as Fault Mode 2. This initial weight matrix is ​​set so that the diagonal is γ, and all other elements are significantly smaller than γ. Then, the diagonal elements of W are strengthened through regularization loss: this regularization term is placed after the label bias immunity loss to guide W to learn correctly.

[0116] After the above processing, the obtained weight allocation matrix is ​​the global optimal solution, and the corresponding fault diagnosis rate is also the global optimal solution.

[0117] like Figure 2 The diagram shows the structure of the modal channel attention module. Taking a CNC machine tool feed system as an example, the original data is a long-term sequence of 6 modalities. After slicing, N samples of shape [6, 2000] are obtained. The samples are reorganized into channels based on prior knowledge, making the i-th modal channel sensitive to the i-th fault mode. The reorganized output is multiplied by the initial weight matrix, and then added back to the reorganized output to obtain the module output (this step is residual connection).

[0118] Figure 3 The results of training the weight matrix with and without the modal channel attention module are significantly different, indicating that the modal channel attention module provided in this embodiment of the invention enables the weight matrix to find the global optimal solution.

[0119] Step 3: Use the noise reduction interactive sample processing strategy to build the dataset for training the model.

[0120] In the collected dataset, a small portion of clean data was selected and divided into validation and test sets in a 1:1 ratio, while the vast majority of datasets containing annotation bias were used as the training set. Due to measurement equipment and human error, low-quality annotations in real-world industrial datasets are unavoidable; this portion of the dataset containing annotation bias was used as the training set. Considering that validation and testing require a dataset free of annotation bias, a separate small dataset was collected, annotated manually and by machine, and subjected to multiple cross-validations to ensure the accuracy of the annotations. This portion of the dataset was also divided into validation and test sets in a 1:1 ratio.

[0121] Step 4: Input the training set, validation set, and test set into the model together. Just make sure that the shape of the input and output meets the customer's requirements.

[0122] The model here can be any deep learning-based fault diagnosis model. For example... Figure 4 The diagram illustrates an embedded fault diagnosis model architecture that allows existing models to be compatible with the annotation bias immune loss function and modal channel attention module, ensuring flexibility in model selection. In this embodiment, the fault diagnosis model within the framework can be freely replaced, as long as it meets the requirements of the input sample shape and the output fault mode type.

[0123] In this embodiment, the embedded fault diagnosis model architecture consists of an in-situ data acquisition part, a noise reduction interactive sample processing strategy, an embedded model training process, and a fault diagnosis task.

[0124] The in-situ data acquisition section consists of a numerical control system and a vibration sensor. The numerical control system provides two modes: current and tracking error. The vibration sensor provides vibration modes at four different locations. The data is then processed through a slice and a modal channel attention module.

[0125] Noise Reduction Interactive Sample Processing Strategy: Due to measurement equipment and human error, low-quality annotations in real-world industrial datasets are unavoidable. This portion of the dataset containing annotation bias is used as the training set. Considering that validation and testing require a dataset free of annotation bias, a separate small-scale dataset needs to be collected, annotated manually and by machine, and subjected to multiple cross-validations to ensure the accuracy of the dataset annotations. This portion of the dataset is divided into validation and test sets in a 1:1 ratio.

[0126] The embedded model training process can use any deep learning-based fault diagnosis model, as long as the number of output layers and fault modes are kept consistent, and the length of the input layer and the length of the slice samples are kept consistent. The loss function used is the label bias immune loss function.

[0127] Fault diagnosis task: The trained model can be used for new fault diagnosis tasks. For a feed system with an unknown fault, it is only necessary to collect 6 channels of signal, slice them and input them into the trained model to complete the fault diagnosis task.

[0128] Step 5: Introduce label bias immune loss to replace standard cross-entropy loss for training.

[0129] In this embodiment, the label bias immune cross-entropy loss adds adaptive coefficients to the cross-entropy and inverse cross-entropy respectively. The adaptive coefficients are determined by a custom normalized exponential function and a custom confidence factor, which can significantly improve the accuracy of the fault diagnosis method under label bias data and significantly improve the convergence effect of the training process.

[0130] The annotation bias immune loss function is:

[0131]

[0132] The gradient calculation during the training process of the label bias immune loss function is as follows:

[0133]

[0134] Among them, L USCE The annotation bias immune loss function, Let y be the gradient value of the label bias immune loss function, and y be the actual label vector. i For the actual label of the i-th fault mode in a single sample, To predict the label vector, Here is the predicted label for the i-th fault mode in a single sample, where k is the total number of fault modes, i is the fault mode number, and s is the predicted label for the i-th fault mode in a single sample. i Let W be the custom normalization index for the i-th failure mode, C be the custom confidence factor, and W be the custom confidence index. iiLet be the weight value of the i-th fault mode on the i-th channel signal.

[0135] The one-hot encoded vector of a single sample's actual label is y = [y1, y2, ..., y]. i ,…y k Each element y of ] i Each can independently take the value 0 or 1. The predicted label is a probability vector. in k indicates that there are k different failure modes.

[0136] In this embodiment, the cross-entropy has two terms. Inverse cross-entropy has two terms C(1-s) i )logy i , The coefficient of each term depends on the values ​​of S and C, with coefficient 1 being... Coefficient 2 is C(1-s) i ), coefficient 3 is C(1-s i ), coefficient 4 is

[0137] The custom normalization exponential function mainly calculates the one-hot encoded label of a single sample. The more 1s there are, the more reliable the label encoding is, and the greater its impact on gradient calculation. Conversely, the more 0s there are, the less reliable the label encoding is, and the less its impact on gradient calculation is.

[0138]

[0139] Where s is the weight of the encoding position, U-Softmax is a custom normalized exponential function, and η is a scaling factor (hyperparameter) used to adjust the weight ratio of different tag encoding positions; i The elements of vector s represent the weights at each encoding position.

[0140] The custom confidence factor mainly calculates that, in a batch calculation, the larger the 0 norm of the one-hot encoding, the lower the confidence of the sample itself, and the smaller the influence of the sample on the gradient calculation; conversely, the smaller the 0 norm of the one-hot encoding, the higher the confidence of the sample itself, and the larger the influence of the sample on the gradient calculation.

[0141]

[0142] Here, α controls the lower bound of the confidence level, β adjusts the hyperparameter of the correlation between the confidence level and the 0 norm, and N0 represents the 0 norm of the sample.

[0143] Step Six: Input the new sample into the framework of the trained model to complete the fault diagnosis.

[0144] Compared with existing technologies, the fault diagnosis method for CNC machine tool feed systems based on annotation deviation immune loss provided by this invention has the following characteristics:

[0145] 1. The labeling bias immune loss function proposed in this invention is compatible with existing deep learning frameworks (such as PyTorch and TensorFlow) and various deep learning backbone networks (CNN, Transformer, RNN). It directly replaces the standard cross-entropy loss function, requires no additional labeling and cleaning steps, and significantly reduces the impact of labeling bias of historical fault data on the model.

[0146] 2. The proposed modal channel attention module is suitable for fault diagnosis tasks of CNC machine tool feed systems. It is plug-and-play and obtains the global optimal solution for the weight allocation of each modality, thereby significantly improving the accuracy of feed system fault diagnosis.

[0147] 3. The designed embedded fault diagnosis model architecture enables existing models to be compatible with the label bias immune loss function and modal channel attention module, which significantly improves the fault diagnosis accuracy of the model while ensuring the flexibility of model selection.

[0148] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A fault diagnosis method for CNC machine tool feed systems based on annotation deviation immune loss, characterized in that, The method includes the following steps: S101: Collect the operating data of the CNC machine tool feed system, process the operating data to form a dataset, and the dataset includes real-time data and historical fault data; S103: Slice the data in the dataset, and input the sliced ​​data into the modal channel attention module for feature fusion and weight allocation; S105: Construct an embedded fault diagnosis framework, which includes a fault diagnosis model, and use a noise reduction interactive sample processing strategy to obtain data from the dataset for training the embedded fault diagnosis framework. S107: The embedded fault diagnosis framework is trained using label bias immune loss instead of standard cross-entropy loss; S109: Input the new sample into the fault diagnosis model trained in the embedded fault diagnosis framework to perform fault diagnosis on the CNC machine tool feed system.

2. The method as described in claim 1, characterized in that, In step S101, the dataset includes the CNC system current signal, the tracking error signal, and the vibration signal collected by the vibration sensor. The vibration signal includes vibration signals from four points, and the dataset is organized into six channels of raw input data.

3. The method as described in claim 2, characterized in that, In step S103, the modal channel attention module, based on the channel attention module, introduces initialization weights and knowledge regularization terms according to prior knowledge to guide the weight allocation of each modality. The modal channel attention module is as follows: O j =[g1,g2,…,g j ] Among them, O j Let g be a channel signal matrix consisting of j channel signals, O′j be an intermediate matrix after channel reorganization and weight allocation of samples based on prior knowledge, and O″j be the output matrix after residual concatenation of the intermediate matrix and the original matrix. j Let g be the signal of the j-th group of the original matrix. j ′ represents the signal of the j-th group in the intermediate matrix, g j " is the signal of the j-th group in the output matrix. This is the weight matrix.

4. The method as described in claim 3, characterized in that, Step S103 includes the following sub-steps: S1031: Slicing: Slice the long-time sequence of the input 6 modalities to obtain multiple samples of shape [6, 2000]. S1032: Channel Reorganization: Based on prior knowledge, the sample is reorganized into channels so that the i-th modal channel is sensitive to the i-th fault mode; S1033: Weight Allocation: Multiply the recombined samples by the initial weight matrix W to obtain the intermediate matrix O′. j ; S1034: Residual Linkage: Connect O′ j Add the recombined sample O j The output O of the modal channel attention module is obtained. ″j .

5. The method as described in claim 4, characterized in that, In step S105, the fault diagnosis model in the embedded fault diagnosis framework adopts a deep learning-based fault diagnosis model. The architecture of the embedded fault diagnosis framework includes an in-situ data acquisition module, a noise reduction interactive sample processing strategy, an embedded model training module, and a fault diagnosis task module. The in-situ data acquisition module collects multiple modal data from the CNC machine tool feed system, and after slicing, sends them to the modal channel attention module. The noise reduction interactive sample processing strategy uses the dataset with labeling bias as the training set and divides the accurately labeled dataset into a validation set and a test set. The embedded model training module selects any advanced fault diagnosis model as the backbone network and uses label bias immune loss instead of standard cross-entropy loss to train the fault diagnosis model in the embedded fault diagnosis framework. The fault diagnosis task module uses the trained embedded fault diagnosis framework for new fault diagnosis tasks, and performs fault diagnosis tasks for CNC machine tool feed systems with unknown faults.

6. The method as described in claim 5, characterized in that, The fault diagnosis model in the embedded fault diagnosis framework is configured as follows: the number of output layers and fault modes of the fault diagnosis model are consistent; the length of the input layer of the fault diagnosis model is consistent with the length of the slice sample; and the loss function used by the fault diagnosis model is the label bias immune loss function.

7. The method as described in claim 6, characterized in that, In the labeling bias immune loss function, adaptive coefficients are added to both the cross-entropy and inverse cross-entropy. These adaptive coefficients are determined by a custom normalized exponential function and a custom confidence factor. The labeling bias immune loss function is as follows: y=[y1,y2,…,y i ,…y k ] Among them, L USCE Let y be the label bias immune loss function, and y be the actual label vector. i For the actual label of the i-th fault mode in a single sample, To predict the label vector, Here is the predicted label for the i-th fault mode in a single sample, where k is the total number of fault modes, i is the fault mode number, and s is the predicted label for the i-th fault mode in a single sample. i Let W be the custom normalization index for the i-th failure mode, C be the custom confidence factor, and W be the custom confidence index. ii Let be the weight value of the i-th fault mode on the i-th channel signal.

8. The method as described in claim 7, characterized in that, The gradient calculation of the label bias immune loss function during training is as follows: in, Let y be the gradient value of the label bias immune loss function, and y be the actual label vector. i For the actual label of the i-th fault mode in a single sample, To predict the label vector, Here is the predicted label for the i-th fault mode in a single sample, where k is the total number of fault modes, i is the fault mode number, and s is the predicted label for the i-th fault mode in a single sample. i Let W be the weight of the encoding position for the i-th fault mode, C be a custom confidence factor, and W be the weight of the encoding position for the i-th fault mode. ii Let be the weight value of the i-th fault mode on the i-th channel signal.

9. The method as described in claim 8, characterized in that, The customized normalized exponential function is used to calculate the one-hot encoded label confidence of a single sample and its impact on gradient calculation. The customized normalized exponential function is as follows: Where s is the weight of the encoding position, U-Softmax is a custom normalized exponential function, η is the scaling factor, and y i is the actual label for a single sample, and k is the total number of failure modes.

10. The method as described in claim 9, characterized in that, The customized confidence factor is determined by calculating the zero-norm value of the one-hot encoding in a batch to determine the confidence level of the sample itself and its impact on gradient calculation. The customized confidence factor is calculated as follows: Among them, C i Let α be the confidence factor for the i-th failure mode, β be the lower bound of the confidence level, β be the hyperparameter that adjusts the correlation between the confidence level and the 0-norm, and N0 be the 0-norm of the sample.