A tool wear prediction method based on physical information neural network of Transformer
Through the physical information neural network based on the Transformer architecture, combined with multi-dimensional signal features and normalization processing, a data-mechanism dual-driven loss function is constructed, which solves the shortcomings of the existing tool wear model in multi-domain feature fusion and time dynamic modeling, and achieves accurate prediction of tool wear status and good generalization ability.
Patent Information
- Application Number
- CN202511013403.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing tool wear monitoring models fail to effectively reflect the complexity of wear when integrating multi-domain features, and have limited ability to model the time dynamics of the wear process, making it difficult to accurately reflect the temporal changes of tool status.
A physical information neural network based on the Transformer architecture is adopted, combined with multi-dimensional signal features and normalization processing, to construct a data-mechanism dual-driven loss function. The time dependence and feature interaction in the wear evolution process are enhanced through the self-attention mechanism, and the physical constraints of Archard's wear law are introduced to achieve collaborative learning of wear evolution laws and physical parameters.
It achieves accurate prediction of tool wear status, has good physical consistency and generalization ability, and improves prediction accuracy and model applicability.
Smart Images

Figure CN120524154B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tool wear prediction based on deep learning, and specifically to a tool wear prediction method based on a Transformer physical information neural network. Background Art
[0002] With the rapid development of information technology, traditional manufacturing is accelerating its transformation toward digitalization and intelligentization. As a crucial component of intelligent manufacturing, tool wear monitoring plays a key role in improving production efficiency and automation. During machining, the tool, as the actuator of a CNC machine tool, inevitably rubs against the workpiece. This wear increases with machining time, ultimately leading to tool failure. Therefore, real-time monitoring of tool status and accurate prediction of wear trends are crucial for reducing cutting downtime, optimizing resource scheduling, lowering production costs, and ensuring product quality.
[0003] Deep learning has become a key technology for tool wear monitoring. With its powerful nonlinear modeling capabilities, it can accurately map multi-frequency domain features to tool wear states, significantly reducing reliance on expert experience and manual feature engineering. Currently, convolutional neural networks and recurrent neural networks have been widely used in this field. Models based on single-domain features often ignore the interactions between features and fail to fully reflect the complexity of wear. Although hybrid models of multi-domain features have theoretical advantages, their potential has been difficult to realize in practical applications. However, two major shortcomings remain: First, most research still relies primarily on pure data-driven approaches, failing to effectively integrate the physical evolution mechanisms of tool wear, limiting the physical consistency and generalization capabilities of the models. Second, existing models have limited ability to model the temporal dynamics of the wear process, particularly in long time series and nonlinear evolution modeling, making it difficult to accurately reflect the temporal changes in tool states. Summary of the Invention
[0004] In order to solve the above technical problems, this application proposes the following technical solutions:
[0005] The present application provides a tool wear prediction method based on a physical information neural network of a Transformer, including:
[0006] After obtaining the wear signal through machine tool milling, the time domain characteristics, frequency domain characteristics and time-frequency domain characteristics of the wear signal are calculated;
[0007] The calculated multi-dimensional signal feature vector is normalized and then divided into a training set, a test set, and a validation set;
[0008] Construct a physical information neural network based on the Transformer architecture and a loss function based on data-mechanism dual drive;
[0009] The training set is input into a physical information neural network based on a Transformer architecture to perform tool wear prediction training, and a validation set is used for verification to select the optimal prediction model;
[0010] The test set is input into the optimal prediction model for prediction to obtain the prediction result of tool wear.
[0011] In a possible implementation, the step of obtaining the wear signal through milling using a machine tool and calculating the time domain characteristics, frequency domain characteristics, and time-frequency domain characteristics of the wear signal includes:
[0012] The time domain characteristics of the mean, peak, root mean square value, root amplitude, skewness, kurtosis, waveform factor, pulse factor, skewness factor, peak factor, margin factor and kurtosis factor of the wear signal, the frequency domain characteristics of the center of gravity frequency, root mean square frequency, frequency variance, and frequency standard deviation, and the time-frequency domain characteristics of the eight frequency band characteristics extracted by wavelet packet transform are calculated respectively.
[0013] In a possible implementation, normalizing the calculated multidimensional signal feature vector and dividing it into a training set, a test set, and a validation set includes:
[0014] The calculated multi-dimensional signal feature vector is proportionally mapped to the preset interval and then normalized. The calculation formula is:
[0015]
[0016] Where X represents the normalized eigenvector value, x is the calculated multidimensional signal eigenvector, represents the minimum value in the data set, Represents the maximum value in the data set.
[0017] In one possible implementation, the physical information neural network based on the Transformer architecture includes an input embedding layer, a multi-head attention module, a feedforward network layer, a residual connection layer and a normalization layer, the normalization layer includes a first normalization layer and a second normalization layer, the output end of the input embedding layer is connected to the input end of the multi-head attention module, the output end of the multi-head attention module is connected to the input end of the residual connection layer, the output end of the residual connection layer is connected to the input end of the first normalization layer, the output end of the first normalization layer is respectively connected to the input end of the feedforward network layer and the second normalization layer, and the output end of the feedforward network layer is connected to the second normalization layer.
[0018] In one possible implementation, the training set is input into a physical information neural network based on a Transformer architecture for tool wear prediction training, including:
[0019] After inputting the training set into the input embedding layer of the physical information neural network based on the Transformer architecture, the training set is embedded into a unified representation space through linear transformation;
[0020] The multi-head attention module embeds the input through linear transformation to obtain the query, key and value multi-head attention mechanism;
[0021] All attention module outputs are concatenated and linearly mapped to output;
[0022] The output of the multi-head attention module is added to the input element by element through the residual connection layer to enhance feature transfer and alleviate gradient disappearance;
[0023] The first normalization layer is used to normalize the result after residual connection, and the feedforward network layer is used to extract high-order features;
[0024] The output of the feedforward network layer is added to the input residual again and normalized, and the predicted wear value is output through the linear layer.
[0025] In one possible implementation, the calculation formula for embedding the training set into a unified representation space through linear transformation is:
[0026]
[0027]
[0028] in, is the weight matrix, , is the corresponding bias term, is the total embedding dimension, t is the normalized time variable, x is the calculated multidimensional signal feature vector, is the initial input sequence of Transformer, B is the sample batch size.
[0029] In one possible implementation, the multi-head attention module embeds the input into the query, key, and value through linear transformation. The calculation formula of the multi-head attention mechanism is:
[0030]
[0031] in, , represents the total embedding dimension of the model, dIndicates the number of heads of multi-head attention, represents the dimension of each attention head, For the The query matrix of each head, For the The transpose of the key matrix of the head, For the The value matrix of the head.
[0032] In one possible implementation, the calculation formula of the loss function based on data-mechanism dual drive is:
[0033]
[0034]
[0035]
[0036] in, , is a hyperparameter, N is the total number of acquisitions, For the The actual value of tool wear at the time of acquisition, For the The wear value predicted by the acquisition model is is the data loss function, is the mechanism loss function, a, b, c 1 , c 2 is the physical parameter to be learned, For the The predicted tool wear value of samples, For the The time point corresponding to each sample.
[0037] In one possible implementation, the use of a validation set for validation and selection of an optimal prediction model includes:
[0038] Use Adam as the optimizer and perform normalization via Batch Norm.
[0039] Set the activation function to Relu and perform specific model training. During the training process, set the training rounds, early stopping rounds, and batch processing parameters, and save the optimal prediction model during the training process.
[0040] Compared with the prior art, the present invention has the following advantages:
[0041] This application proposes a physical information neural network tool wear prediction method based on the Transformer architecture to achieve data-driven and mechanism-driven fusion modeling. Utilizing the advantages of Transformer in long-term time series modeling and multi-dimensional feature fusion, the time-frequency domain features and normalized time of multi-source sensors are used as input, and the time dependence and feature interaction in the wear evolution process are strengthened through the self-attention mechanism. At the same time, physical constraints based on Archard's wear law are introduced to construct a joint loss function containing data errors and physical residuals, thereby achieving collaborative learning of wear evolution laws and physical parameters. Compared with current prediction methods, this method has good physical consistency and generalization capabilities while ensuring prediction accuracy, providing an effective new approach for tool wear status monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A flow chart of a tool wear prediction method based on a Transformer-based physical information neural network provided in an embodiment of the present application;
[0043] Figure 2 A schematic diagram of the overall structure of a physical information neural network based on the Transformer architecture provided in an embodiment of the present application;
[0044] Figure 3 Transformer network architecture diagram provided for the embodiments of this application;
[0045] Figure 4 This is a diagram of tool wear prediction results provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The present invention will be described below with reference to the accompanying drawings and specific implementation methods.
[0047] Figure 1 A schematic diagram of a tool wear prediction method based on a physical information neural network of a Transformer provided in an embodiment of the present application is provided. Figure 1 In this embodiment, a tool wear prediction method based on a physical information neural network of a Transformer includes:
[0048] S101, after obtaining a wear signal through machine tool milling, calculate the time domain characteristics, frequency domain characteristics and time-frequency domain characteristics of the wear signal.
[0049] In this embodiment, the time domain, frequency domain and time-frequency domain features of the acquired wear signal are calculated, wherein the time domain features include 12 types: mean, peak, root mean square value, root amplitude, skewness value, kurtosis value, waveform factor, pulse factor, skewness factor, peak factor, margin factor and kurtosis factor; the frequency domain features include 4 types: centroid frequency, root mean square frequency, frequency variance and frequency standard deviation; the time-frequency domain features include 8 frequency band features extracted by wavelet packet transform. There are 24 types of time-frequency features of all signals in total, and the calculation expressions are shown in Table 1.
[0050] Table 1 Feature calculation expression
[0051]
[0052]
[0053] in, N represents the number of feature samples, Indicates the Sample values, X ( fi ) means that after Fourier transform, at point fi The complex spectrum value at fi Indicates the first discrete frequency points, E1-E8 correspond to the energy values of the 1st to 8th frequency bands respectively. Indicates the The wavelet packet coefficients of the subbands, t Represents the subband coefficient index point, Ni Indicates the length of the subband coefficients.
[0054] S102, normalizing the calculated multi-dimensional signal feature vector and dividing it into a training set, a test set, and a validation set.
[0055] In this embodiment, the input features are normalized, and the actual tool wear values are also normalized accordingly. The calculated multi-dimensional signal feature vector is proportionally mapped to the [0.1] interval and then normalized. The calculation formula is:
[0056]
[0057] Where X represents the normalized eigenvector value, x is the calculated multidimensional signal eigenvector, represents the minimum value in the data set, Represents the maximum value in the data set.
[0058] After normalization, the dataset is divided into 70% of the labeled data for training, 10% of the data for validation, and 20% of the data for the final model testing.
[0059] S103, construct a physical information neural network based on the Transformer architecture and a loss function based on data-mechanism dual drive.
[0060] See also Figure 2 In this embodiment, the proposed physical information neural network based on the Transformer architecture is mainly composed of an input embedding layer, a multi-head attention module, a feedforward network layer, a residual connection layer, and a normalization layer. The entire Transformer structure is stacked in multiple layers to learn the complex interaction between time series and features. The specific model architecture is as follows Figure 3 As shown. The normalization layer includes a first normalization layer and a second normalization layer. The output of the input embedding layer is connected to the input of the multi-head attention module. The output of the multi-head attention module is connected to the input of the residual connection layer. The output of the residual connection layer is connected to the input of the first normalization layer. The output of the first normalization layer is respectively connected to the input of the feedforward network layer and the second normalization layer. The output of the feedforward network layer is connected to the second normalization layer.
[0061] In this embodiment, a loss function based on data-mechanism dual drive is constructed, including constructing a data loss part and constructing a mechanism loss function. Among them, the loss function of the data loss part is expressed as the difference between the wear value predicted by the model and the actual wear value, and the calculation formula is:
[0062]
[0063] in, N is the total number of acquisitions, For the i The actual value of tool wear at the time of acquisition, For the i The wear value predicted by the acquisition model is is the data loss function.
[0064] In the mechanism equation driving the loss part, in order to introduce this physical process into the neural network training, the following parameterized wear evolution differential equation is constructed to describe the growth trend of tool wear over time. The calculation formula is:
[0065]
[0066] Among them, the index term Reflecting the nonlinear characteristics of the wear process, the polynomial term The time-varying characteristics of the wear rate are represented by the modeling of the corresponding force and speed changes over time, which can be compared to the effects of the fluctuation of the machining load and speed over time. a, b, c 1 , c 2is the physical parameter to be learned, For the The predicted tool wear value of samples, For the The time point corresponding to each sample.
[0067] Therefore, the final calculation formula based on the data-mechanism loss function is:
[0068]
[0069] in, , is a hyperparameter used to control the weight balance between data error and physical residual.
[0070] S104: Input the training set into a physical information neural network based on the Transformer architecture for tool wear prediction training, and use the validation set for verification to select the optimal prediction model.
[0071] In this embodiment, after the training set is input into the input embedding layer of the physical information neural network based on the Transformer architecture, the training set is embedded into a unified representation space through linear transformation. The calculation formula is:
[0072]
[0073]
[0074] in, is the weight matrix, , is the corresponding bias term, is the total embedding dimension, t is the normalized time variable, x is the calculated multidimensional signal feature vector, is the initial input sequence of Transformer, B is the sample batch size.
[0075] The multi-head attention module embeds the input through linear transformation to obtain the query, key and value multi-head attention mechanism. The calculation formula is:
[0076] in, , represents the total embedding dimension of the model, d Indicates the number of heads of multi-head attention, represents the dimension of each attention head, For the The query matrix of each head, For the The transpose of the key matrix of the head, For the The value matrix of the head.
[0077] All attention module outputs are concatenated and then linearly mapped to output.
[0078] The output of the multi-head attention module is element-wise added to the input through a residual connection layer to enhance feature transfer and mitigate gradient vanishing. The first normalization layer normalizes the residual connection results. The feedforward network uses a fully connected network to further extract high-order features. Finally, the output of the feedforward network layer is again added to the input residual and normalized, and the predicted wear value is output through a linear layer. The entire Transformer structure is stacked in multiple layers to learn the complex interactions between time series and features.
[0079] In this embodiment, the validation set is used for verification and the optimal prediction model is selected, including: using Adam as the optimizer, performing normalization operations through Batch Norm, setting the activation function to Relu, and performing specific model training. At the same time, during the training process, the training rounds are set to 4000, the early stopping rounds are set to 200, and the batch processing parameter is set to 32. The optimal prediction model during the training process is saved, and finally the network outputs the wear value on the validation set through the linear layer.
[0080] S105: Input the test set into the optimal prediction model for prediction to obtain the prediction result of tool wear.
[0081] In this embodiment, Figure 4 The comparison of the prediction results of the proposed method on the test set is shown. It can be clearly seen from the figure that the tool wear curve predicted by the proposed method is highly consistent with the actual tool wear curve. It can effectively extract deep information from the wear data and realize accurate modeling of tool wear trends, which reflects the specific good generalization ability of the model.
[0082] In the embodiments of this application, "plurality" refers to two or more. "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean that A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0083] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0084] The above description is merely a specific embodiment of the present application. Any person skilled in the art may easily conceive of variations or substitutions within the technical scope disclosed in this application, and such variations or substitutions shall be within the scope of protection of this application. The scope of protection of this application shall be subject to the scope of protection of the claims.
Claims
1. A tool wear prediction method based on a physical information neural network of Transformer, characterized in that: include: After obtaining the wear signal through machine tool milling, the time domain characteristics, frequency domain characteristics and time-frequency domain characteristics of the wear signal are calculated; The calculated multi-dimensional signal feature vector is normalized and then divided into a training set, a test set, and a validation set; Construct a physical information neural network based on the Transformer architecture and a loss function based on data-mechanism dual drive; The calculation formula of the loss function based on data-mechanism dual drive is: in, , is a hyperparameter, N is the total number of acquisitions, For the The actual value of tool wear at the time of acquisition, For the The wear value predicted by the acquisition model is is the data loss function, is the mechanism loss function, a, b, c 1 , c 2 is the physical parameter to be learned, For the The predicted tool wear value of samples, For the The time point corresponding to each sample; The training set is input into a physical information neural network based on a Transformer architecture to perform tool wear prediction training, and a validation set is used for verification to select the optimal prediction model; The test set is input into the optimal prediction model for prediction to obtain the prediction result of tool wear.
2. The tool wear prediction method based on the Transformer physical information neural network according to claim 1 is characterized in that: After the wear signal is acquired through machine tool milling, the time domain characteristics, frequency domain characteristics and time-frequency domain characteristics of the wear signal are calculated, including: The time domain characteristics of the mean, peak, root mean square value, root amplitude, skewness, kurtosis, waveform factor, pulse factor, skewness factor, peak factor, margin factor and kurtosis factor of the wear signal, the frequency domain characteristics of the center of gravity frequency, root mean square frequency, frequency variance, and frequency standard deviation, and the time-frequency domain characteristics of the eight frequency band characteristics extracted by wavelet packet transform are calculated respectively.
3. The tool wear prediction method based on the Transformer physical information neural network according to claim 1 is characterized in that: The calculated multidimensional signal feature vector is normalized and divided into a training set, a test set, and a validation set, including: The calculated multi-dimensional signal feature vector is proportionally mapped to the preset interval and then normalized. The calculation formula is: Where X represents the normalized eigenvector value, x is the calculated multidimensional signal eigenvector, represents the minimum value in the data set, Represents the maximum value in the data set.
4. The tool wear prediction method based on the Transformer physical information neural network according to claim 1 is characterized in that: The physical information neural network based on the Transformer architecture includes an input embedding layer, a multi-head attention module, a feedforward network layer, a residual connection layer and a normalization layer. The normalization layer includes a first normalization layer and a second normalization layer. The output end of the input embedding layer is connected to the input end of the multi-head attention module, the output end of the multi-head attention module is connected to the input end of the residual connection layer, the output end of the residual connection layer is connected to the input end of the first normalization layer, the output end of the first normalization layer is respectively connected to the input end of the feedforward network layer and the second normalization layer, and the output end of the feedforward network layer is connected to the second normalization layer.
5. The tool wear prediction method based on the Transformer physical information neural network according to claim 1 is characterized in that: The training set is input into a physical information neural network based on a Transformer architecture for tool wear prediction training, including: After inputting the training set into the input embedding layer of the physical information neural network based on the Transformer architecture, the training set is embedded into a unified representation space through linear transformation; The multi-head attention module embeds the input through linear transformation to obtain the query, key and value multi-head attention mechanism; All attention module outputs are concatenated and linearly mapped to output; The output of the multi-head attention module is added to the input element by element through the residual connection layer to enhance feature transfer and alleviate gradient disappearance; The first normalization layer is used to normalize the result after residual connection, and the feedforward network layer is used to extract high-order features; The output of the feedforward network layer is added to the input residual again and normalized, and the predicted wear value is output through the linear layer.
6. The tool wear prediction method based on the Transformer physical information neural network according to claim 5 is characterized in that: The calculation formula for embedding the training set into a unified representation space through linear transformation is: in, is the weight matrix, , is the corresponding bias term, is the total embedding dimension, t is the normalized time variable, x is the calculated multidimensional signal feature vector, is the initial input sequence of Transformer, B is the sample batch size.
7. The tool wear prediction method based on the Transformer physical information neural network according to claim 5 is characterized in that: The multi-head attention module embeds the input into the query, key and value through linear transformation. The calculation formula of the multi-head attention mechanism is: in, , represents the total embedding dimension of the model, d Indicates the number of heads of multi-head attention, represents the dimension of each attention head, For the The query matrix of each head, For the The transpose of the key matrix of the head, For the The value matrix of the head.
8. The tool wear prediction method based on the Transformer physical information neural network according to claim 1 is characterized in that: The method of using the validation set to perform validation and select the optimal prediction model includes: Use Adam as the optimizer and perform normalization via Batch Norm. Set the activation function to Relu and perform specific model training. During the training process, set the training rounds, early stopping rounds, and batch processing parameters, and save the optimal prediction model during the training process.
Citation Information
Patent Citations
Deep learning and physical model guide fused dense medium sorting density decision-making method
CN119049595A
Distribution network single-phase earth fault prediction method based on LSTM-Transform neural network
CN120178095A