Cutter wear prediction method and system
By constructing a tool wear prediction model including KAN network, LSTM network and self-attention mechanism, the problem of difficulty in capturing nonlinear and timing characteristics in milling data in the prior art is solved, and higher prediction accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510279106.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art is difficult to simultaneously capture complex nonlinear and timing features present in milling data, resulting in low tool prediction accuracy.
The tool wear prediction model is constructed using KAN network, three-layer long and short-term memory network, two-layer fully connected layer and random inactivation layer, and a self-attention mechanism and weight penalty term are introduced into the model to improve the accuracy and efficiency of prediction.
Effectively capture complex dynamic features and timing information during tool wear, improve the accuracy and efficiency of prediction, reduce overfitting, and enhance the generalization ability of the model.
Smart Images

Figure CN119910504A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of tool wear prediction, and in particular to a tool wear prediction method and system. Background Art
[0002] In the field of industrial manufacturing, the efficient operation of CNC machine tools is crucial to the entire industrial system. As the key executive component of the machine tool, the performance of the tool directly affects the processing efficiency and workpiece quality. As the machining progresses, the friction between the tool and the workpiece causes material loss and wear, which not only reduces the machining efficiency and prolongs the cycle, but also affects the surface quality of the workpiece and increases the cost. Therefore, accurate prediction of tool wear is of great significance for improving machining efficiency, ensuring workpiece quality and reducing costs. Traditional tool wear monitoring methods include cutting force measurement, motor current monitoring, sound feature analysis and vibration signal detection. Although these methods can reflect the tool wear state, they have limitations under complex cutting conditions. In recent years, research models based on machine learning and deep learning have gradually become the mainstream methods for tool wear prediction due to their advantages in processing complex data relationships. Domestic and foreign scholars have conducted a lot of research on tool wear prediction and proposed a variety of algorithms, such as least squares support vector machine, residual variational autoencoder, reinforcement learning algorithm, etc. These methods have improved the accuracy of prediction to a certain extent. Deep learning models, especially those that combine convolutional neural networks, long short-term memory networks, and particle swarm optimization, are better at automatically extracting features and processing complex functions, thereby improving the accuracy and efficiency of predictions. Although research has made significant progress, there are still challenges in capturing complex long-term dependencies in time series, computing efficiency in processing large-scale data, and computing resource requirements of models. Summary of the invention
[0003] In view of the defects in the prior art, the present invention provides a tool wear prediction method and system, which solves the problem of low tool prediction accuracy caused by the difficulty in simultaneously capturing the complex nonlinear characteristics and timing characteristics existing in milling data in the prior art.
[0004] In order to achieve the above-mentioned purpose, one aspect of the present invention provides a tool wear prediction method, which includes: obtaining milling data related to tool milling, and preprocessing the milling data to obtain sample data; constructing a tool wear prediction model using a KAN network, a three-layer long short-term memory network, two fully connected layers and a random inactivation layer; setting a weight penalty term related to the weight of the tool wear prediction model, and adding the weight penalty term to the loss function of the tool wear prediction model to obtain a weight-optimized tool wear prediction model; introducing a self-attention mechanism into the weight-optimized tool wear prediction model to obtain an optimized tool wear prediction model; using the sample data to train and evaluate the optimized tool wear prediction model, and based on the evaluation results, using the optimized tool wear prediction model to predict tool wear.
[0005] The KAN network and three-layer long short-term memory network in the tool wear prediction model of the present invention can effectively capture the complex dynamic characteristics and time series information in the tool wear process, set a weight penalty term for the tool wear prediction model loss function, and prevent the tool wear prediction model from overfitting during training. The introduction of the self-attention mechanism enables the model to automatically focus on the key features that have a greater impact on tool wear, further improving the accuracy and efficiency of the prediction.
[0006] Optionally, the preprocessing of the milling data to obtain sample data includes: truncating the milling data using the third quartile method to obtain first data; performing an outlier analysis on the first data, and adjusting the first data according to the result of the outlier analysis to obtain second data; performing denoising on the second data using a wavelet threshold denoising method to obtain third data; performing time domain and frequency domain analysis on the third data to obtain multiple feature data; performing a phase relationship analysis on the feature data using the Spearman correlation coefficient to obtain the most relevant feature data; and normalizing the most relevant feature data to obtain the sample data.
[0007] The present invention truncates the original data by the third quartile method, removes the interference of extreme values, and ensures the centralization and stability of the data. By analyzing and adjusting the data by outliers, erroneous or unreasonable data points are further eliminated to ensure the accuracy of the data. The wavelet threshold denoising method is used to denoise the data, which reduces the interference of noise on subsequent analysis and makes the data clearer. By performing time domain and frequency domain analysis on the data, multiple feature data are extracted, and then the Spearman correlation coefficient is used to screen out the most relevant feature data, avoiding the interference of redundant information and improving the representativeness of the data. The most relevant feature data is normalized to unify the data scale, which is convenient for model training and prediction. This series of preprocessing steps significantly improves the quality of sample data and lays a solid foundation for efficient training and accurate prediction of tool wear prediction models.
[0008] Optionally, performing an outlier analysis on the first data and adjusting the first data according to the result of the outlier analysis to obtain the second data includes: introducing a sliding window mechanism and using a window to slide and sample on the first data to obtain sampled data, wherein the sampled data includes multiple data points; calculating the median of the sampled data and calculating the absolute deviation of the data point to the median through the median; comparing the data point with the absolute deviation and adjusting the data point according to the result of the comparison to obtain the second data.
[0009] The present invention slides and samples on the first data through a sliding window to form multiple local samples, and calculates the median of each sample and the absolute deviation of the data point to the median. By comparing the median with the absolute deviation, the abnormal data points are accurately identified and adjusted, avoiding local abnormalities that may be missed by the global analysis. This method can dynamically adapt to the local characteristics of the data, flexibly handle outliers in different regions, and retain the overall trend and key information of the data, thereby improving the stability and authenticity of the data.
[0010] Optionally, the use of the sample data to train the optimized tool wear prediction model includes: according to the self-attention mechanism, assigning attention weights to the feature data in the sample data to form weighted sample data; using the KAN network to extract the nonlinear change features in the weighted sample data to obtain first feature data; using the three-layer long short-term memory network to extract the time series features of the first feature data to obtain second feature data; using the second feature data to adjust the weights of the optimized tool wear prediction model until the training is completed.
[0011] The present invention dynamically allocates attention weights according to the importance of feature data in sample data through the self-attention mechanism to form weighted sample data, so that the model can focus on features that have a greater impact on tool wear prediction, thereby improving the representativeness of features. The KAN network extracts nonlinear change features in weighted sample data, further mines complex nonlinear relationships in the data, and enhances the adaptability of the model to complex working conditions. The three-layer long short-term memory network (LSTM) deeply mines the time series features of the extracted first feature data, effectively capturing the dynamic information that changes over time during tool wear, making up for the defect of insufficient extraction of time series features by traditional methods, and improving the training effect of optimizing tool wear prediction model.
[0012] Optionally, according to the self-attention mechanism, attention weights are assigned to the feature data in the sample data to form weighted sample data, including: performing a linear transformation on the feature data to obtain a query matrix, a key matrix and a value matrix; calculating an attention score of the feature data based on the query matrix and the key matrix; introducing a Softmax function, and using the Softmax function to convert the attention score into the attention weight; obtaining weighted feature data based on the attention weight and the value matrix, and forming the weighted sample data based on the weighted feature data.
[0013] The present invention performs linear transformation on feature data to generate query matrix, key matrix and value matrix, which provides a basis for subsequent attention calculation. The attention score of feature data is calculated by query matrix and key matrix, the correlation between features is quantified, and the Softmax function is introduced to convert the attention score into normalized attention weight to ensure the rationality and uniqueness of weight distribution. Finally, weighted feature data is obtained according to the attention weight and value matrix, and weighted sample data is formed. The more influential features in the data on tool wear prediction are highlighted, while the interference of unimportant features is suppressed, thus improving the quality of training samples for optimizing the tool wear prediction model.
[0014] Optionally, the weighted feature data satisfies the following formula: is the weighted feature data, is the query matrix, is the bond matrix, is the value matrix, is the permutation of the bond matrix, is the dimension of the key matrix.
[0015] This formula has a simple structure and is easy to calculate.
[0016] Optionally, using the second feature data to adjust the weight of the optimized tool wear prediction model includes: according to the second feature data, using the optimized tool wear prediction model to make a prediction to obtain a prediction result; according to the prediction result, calculating the loss value of the loss function; according to the loss value, calculating the gradient of the loss function relative to the weight of the optimized tool wear prediction model; respectively calculating the weighted moving average of the gradient and the weighted moving average of the square of the gradient to obtain a first estimate and a second estimate; adjusting the weight of the optimized tool wear prediction model according to the first estimate and the second estimate.
[0017] The formula of the present invention realizes weighting of feature data through the interaction of the query matrix, the key matrix and the value matrix, has a simple structure and is easy to calculate.
[0018] Optionally, adjusting the weight of the optimized tool wear prediction model according to the first estimation and the second estimation includes: obtaining an optimized weight according to the first estimation and the second estimation; and adjusting the weight of the optimized tool wear prediction model using the optimized weight; The optimization weight satisfies the following formula: For the optimization weight, is the current weight of the optimized tool wear prediction model, is the current learning rate of the optimized tool wear prediction model, For the first estimate, For the second estimate, To prevent extremely small numbers with a denominator of 0.
[0019] The present invention dynamically adjusts and optimizes the weight of the tool wear prediction model by combining the first estimation and the second estimation, thereby improving the adaptability and prediction accuracy of the model. This method based on dynamically adjusting weights enables the model to better capture the complex changes in the tool wear process.
[0020] Optionally, the first characteristic data satisfies the following formula: is the approximate function of the weighted sample data, For KAN network The number of nodes in the layer, To connect to the KAN network Layer neurons and Layer The activation function of a neuron is For KAN network The number of nodes in the layer, is the number of nodes in the second layer of the KAN network, To connect to the second layer of the KAN network neurons and the third layer The activation function of a neuron is is the number of nodes in the first layer of the KAN network, To connect to the first layer of the KAN network neurons and the second layer The activation function of a neuron is is the number of nodes in the 0th layer of the KAN network, To connect to the 0th layer of the KAN network neurons and the first layer The activation function of a neuron is is the weighted sample data A quantity.
[0021] It should be noted that the number of nodes in the 0th layer of the KAN network refers to the number of nodes in the input layer of the KAN network.
[0022] The activation function of each layer of the formula of the present invention is responsible for converting the output of the previous layer into the input of the current layer. In this way, the model can capture complex patterns and high-order features in the data.
[0023] Another aspect of the present invention also provides a tool wear prediction system, the system comprising: a processor, an input device, an output device and a memory, the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute a tool wear prediction method as described in any one of the previous aspects of the present invention.
[0024] A tool wear prediction system of the present invention has a compact structure, stable performance, high integration and simple composition, and can stably execute a tool wear prediction method provided in the previous aspect of the present invention, further improving the overall applicability and practical application capability of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flow chart of a tool wear prediction method according to an embodiment of the present invention; Figure 2 The figure is a schematic diagram of the structure of a tool wear prediction system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are only for illustration and are not intended to limit the present invention. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that these specific details do not need to be adopted to implement the present invention. In other examples, in order to avoid confusing the present invention, known circuits, software or methods are not specifically described.
[0027] Throughout the specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment of the present invention. Therefore, the phrases "in one embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily all refer to the same embodiment or example. In addition, particular features, structures, or characteristics may be combined in one or more embodiments or examples in any suitable combination and / or subcombination. In addition, it should be understood by those of ordinary skill in the art that the figures provided herein are for illustrative purposes and that the figures are not necessarily drawn to scale.
[0028] Figure 1 2 is a flow chart of a tool wear prediction method according to an embodiment of the present invention. In order to solve the problem of low tool prediction accuracy caused by the difficulty in simultaneously capturing the complex nonlinear characteristics and time series characteristics existing in milling data in the prior art, the method shown in 2 includes the following steps: Step S1, acquiring milling data related to tool milling, and preprocessing the milling data to obtain sample data.
[0029] In this embodiment, the milling data comes from the public data set of the tool remaining life prediction competition held by the American Prognostics and Health Management Society (PHM) in 2010. The experimental conditions of the data set are shown in the following table: Under the above conditions, a full life cycle test was conducted on stainless steel HRC52 material, with a total of 315 milling times. The milling data included: cutting force signals in the X, Y, and Z directions, acceleration vibration signals, and acoustic emission signals during the machining process.
[0030] The preprocessing of the milling data to obtain sample data specifically includes the following sub-steps: Step S101, truncating the milling data using the third quartile method to obtain first data.
[0031] In this embodiment, in order to remove unstable signals during the tool cutting-in and cutting-out stages, the third quartile method is used for truncation. The third quartile method is a statistical method that is often used to identify and process outliers or unstable data in data processing. Compared with other truncation methods, the third quartile method is suitable for processing such dynamic signals: it has strong robustness and can effectively filter out outliers in the signal. Secondly, this method does not rely on the distribution characteristics of the data, so it performs well when processing non-normally distributed force signals and vibration signals.
[0032] The first data refers to data obtained by truncating the milling data using the third quartile method.
[0033] Step S102: performing an outlier analysis on the first data, and adjusting the first data according to a result of the outlier analysis to obtain second data.
[0034] The step of performing an outlier analysis on the first data and adjusting the first data according to the result of the outlier analysis to obtain the second data specifically includes the following sub-steps: Step S10201, introduce a sliding window mechanism, use a window to slide and sample on the first data to obtain sampled data, wherein the sampled data includes multiple data points.
[0035] The sliding window mechanism is a widely used technical means in data processing and analysis. It is like a "window" moving on the data sequence, which is used to observe and process the data locally. First, a fixed-size window is set. The window size represents the number of data points contained in each sampling. For example, the window size is set to 100 data points. The window will start from the starting position of the first data and "frame" the first 100 data points. These 100 data points constitute a set of sampled data. After that, the window will slide forward according to the pre-set step size. The step size refers to the number of data points spanned by the window each time it moves. Assuming the step size is 10, the window will move forward by 10 data points and frame 100 new data points again to form a new set of sampled data. This process is repeated until the window slides to the end of the data sequence.
[0036] Step S10202, calculating the median of the sampled data, and calculating the absolute deviation of the data point to the median through the median.
[0037] In this embodiment, the median is the value in the middle position after the data is arranged in ascending or descending order. If the number of data is an even number, the average of the two middle numbers is taken. After calculating the median, the absolute deviation of each data point to the median can be calculated. The absolute deviation is the absolute value of the difference between the data point and the median, which can reflect the degree to which each data point deviates from the center position. By analyzing these absolute deviations, the discreteness of the sampled data can be understood, providing a basis for subsequent more accurate data analysis and processing.
[0038] Step S10203: compare the data point with the absolute deviation, and adjust the data point according to the comparison result to obtain the second data.
[0039] In this embodiment, the data point is compared with the absolute deviation. If the data point is more than three times the absolute deviation, it is considered an outlier and replaced with the median in the window. This method can effectively filter out outliers and is robust to non-normally distributed data.
[0040] The second data refers to data obtained by replacing the median of the first data.
[0041] Step S103: performing denoising processing on the second data using a wavelet threshold denoising method to obtain third data.
[0042] In this embodiment, the second data is subjected to denoising by using a wavelet threshold denoising method. First, a multi-scale wavelet decomposition operation is performed on the second data, and the second data is decomposed into several levels of approximate signals and detail signals. The approximate signal represents the main features and overall trend of the data, while the detail signal contains local changes in the data and possible noise information.
[0043] Next, threshold processing is applied to each layer of detail signal. Here, soft thresholding is used, which can effectively weaken or remove noise components smaller than the soft threshold while retaining signal characteristics. When a value in the detail signal is smaller than the soft threshold, the value will be adjusted or directly set to zero, thereby reducing the impact of noise on the data.
[0044] Finally, the processed approximate signal and detail signal are reconstructed through inverse wavelet transform. After this series of operations, the noise in the second data is effectively suppressed, providing a more reliable basis for subsequent data analysis and processing.
[0045] Step S104: analyzing the third data in the time domain and the frequency domain to obtain a plurality of feature data.
[0046] In this embodiment, when the third data obtained after the noise reduction process is deeply analyzed, its characteristic information is comprehensively mined from three dimensions: time domain, frequency domain and time-frequency domain.
[0047] In the time domain, we focus on extracting 12 characteristic signals, including maximum value, peak value, variance, etc. The maximum value reflects the maximum amplitude of the data in the time series, the peak value reflects the instantaneous fluctuation intensity of the data, and the variance measures the degree of dispersion of the data relative to the mean. These time domain characteristic signals can clearly show the change law and fluctuation characteristics of the data in the time dimension.
[0048] Frequency domain analysis focuses on four characteristic signals, such as spectral power. Spectral power reveals the energy proportion of different frequency components in the signal. By analyzing the frequency domain characteristics, we can understand the frequency distribution of the signal and identify the main frequency components in the signal.
[0049] The time-frequency domain analysis uses three layers of wavelet decomposition to extract nine characteristic signals. Wavelet decomposition can perform localized analysis of signals in both time and frequency. These time-frequency domain characteristic signals help capture the details of signal changes at different time and frequency scales.
[0050] Therefore, 175 feature data can be extracted for each pass.
[0051] Step S105, using the Spearman correlation coefficient to perform a correlation analysis on the characteristic data to obtain the most relevant characteristic data.
[0052] In this embodiment, in order to screen out the most valuable information for analysis, Spearman correlation coefficient is used for correlation analysis. Spearman correlation coefficient is a non-parametric statistical method that does not depend on the distribution form of data and mainly measures the monotonic relationship between two variables.
[0053] In specific operations, the Spearman correlation coefficient between any two variables in the feature data is calculated. The value range of the coefficient is between -1 and 1. A value close to 1 indicates that the two variables are positively correlated, a value close to -1 indicates a negative correlation, and a value close to 0 indicates a weak correlation. By comparing the sizes of the various correlation coefficients, the features with the strongest correlation with the target variable can be found. Finally, these features with the strongest correlation are determined as the most relevant feature data, which can more accurately reflect the inherent laws of the data and provide strong support for subsequent analysis.
[0054] The Spearman correlation coefficient satisfies the following formula: in is the Spearman correlation coefficient, is the rank difference of each pair of observations, n is the number of observations Step S106: normalize the most relevant feature data to obtain the sample data.
[0055] In this embodiment, normalization is a data preprocessing technology, which converts data according to certain rules so that the data falls within a specific interval.
[0056] Normalizing the most relevant feature data is a key step in obtaining high-quality sample data. As a data preprocessing technique, normalization can convert data into a specific range according to specific rules. Due to the different dimensions and units of different feature data, some features will dominate in value and affect model training. Normalization can eliminate the dimension effect and make each feature equally important. The convergence speed of machine learning algorithms is affected by the scale of data. Normalization makes the data scale similar and can speed up model convergence. At the same time, it can also weaken the impact of outliers on the model and improve the stability and generalization ability of the model. After normalization, sample data that is more suitable for subsequent analysis and modeling can be obtained.
[0057] Step S2, construct a tool wear prediction model using a KAN network, a three-layer long short-term memory network, two fully connected layers, and a random dropout layer.
[0058] In this embodiment, the KAN network is Kolmogorov–Arnold Networks, which is a new type of neural network architecture. The KAN network architecture indicates that the activation function is no longer a fixed Sigmoid or ReLU; instead, it is learnable and applied to connections rather than nodes. Compared with the traditional MLP algorithm, the advantage of the KAN algorithm is that it achieves higher accuracy with fewer parameters and can perform curve learning instead of traditional straight line learning.
[0059] The KAN algorithm states that any continuous multivariable function can be composed of a series of simple univariate functions, which are formally presented as a nested structure. The mathematical expression of the Kolmogorov-Arnold theorem explains that each node can be regarded as a highly nonlinear mapping. The network effectively approximates complex functional relationships through a combination of these mappings. Its key advantages lie in its efficient expression ability and compact model structure. KAN relies on supervised learning and aims to approximate a function f so that the input data point x can be mapped to the corresponding output y. With the help of the Kolmogorov-Arnold theorem, any multivariate function can be decomposed into several univariate functions and their sum. In addition, for each input dimension , there exists a univariate function , which can aggregate the outputs of multiple univariate functions. The flexibility of this structure gives KAN a significant advantage in capturing complex data patterns, especially when combined with the long short-term memory (LSTM) network, which can enhance the model's ability to process time series data, thereby improving the accuracy and stability of predictions.
[0060] The core of KAN lies in its unique structural design, which is significantly different from the architecture of traditional neural networks. KAN learns the activation function as part of the network, so that each connection has not only weight parameters but also a trainable activation function. This feature enables the network to adaptively discover the optimal nonlinear transformation, providing greater flexibility for handling complex nonlinear problems, especially in scenarios with complex data distribution or highly nonlinear relationships.
[0061] In addition, the design of KAN allows switching between coarse-grained and fine-grained grids, using parameterized B-spline activation functions. This flexibility allows the network to effectively simplify calculations and improve training efficiency while capturing subtle features. In this way, KAN not only enhances the expressiveness of the model, but also demonstrates better generalization capabilities in practical applications, especially in areas such as complex time series analysis such as tool wear value prediction.
[0062] The long short-term memory (LSTM) network is a special type of recurrent neural network that aims to solve the gradient vanishing and exploding problems that traditional recurrent neural networks may encounter when processing long sequences. The force signal, acceleration vibration signal, and acoustic emission signal of tool wear used in the dataset are a kind of time-varying sequence data, which reflects the change of tool wear status with processing time. Therefore, it is suitable to use LSTM network to predict tool wear.
[0063] The LSTM network satisfies the following formula: Where: They are the weights of the forget gate, input gate, coupled forget gate, and output gate respectively; They are the biases of the forget gate, input gate, coupled forget gate and output gate respectively.
[0064] Forget Gate: used to determine the current moment , keep the last moment The amount of information in the current input And the output of the previous moment , decide which information needs to be forgotten and which needs to be retained, and store the retained information in the cell state middle.
[0065] Input Gate: Controls the writing of new information. Input Gate Determines the new information that the neuron will store, combined with the current input And the output of the previous moment , generate candidate memory cells , used to update the current cell state.
[0066] 3. Cell status update: cell status , is the result of the combined action of the forget gate and the input gate. The forget gate determines which old information needs to be discarded, and the input gate determines which new information needs to be added.
[0067] Output Gate: Output Gate By current cell state The final output value is determined by the input . Cell State After being processed by a nonlinear function, the output gate Combined to generate output .
[0068] The tool wear prediction model uses the KAN network to extract complex nonlinear features in the data, and the three-layer LSTM is used as the core time series feature extraction unit. LSTM can effectively solve the gradient vanishing problem of traditional recurrent neural networks (RNN) in long time series modeling, and through its memory and forget gate mechanism, it can efficiently model the potential long-term and short-term dependencies in tool wear data. In order to improve the nonlinear expression ability of the model, the rectified linear unit (ReLU) activation function is introduced after the LSTM layer to enhance the model's ability to capture complex wear patterns. In order to prevent overfitting, a random inactivation layer (Dropout) is added to the model to enhance the generalization ability of the model by randomly inactivating some neurons.
[0069] In addition, the original one fully connected layer is expanded to two fully connected layers, which further enhances the model's feature expression ability and ensures the stability and accuracy of the output results. In terms of model hyperparameter optimization, combined with the powerful function of the Kolmogorov-Arnold representation theorem, it can map complex multidimensional nonlinear systems into linear combinations, and overcome the limitations of traditional optimization algorithms that converge quickly but easily fall into local optimality by modeling the implicit laws within the system, enhancing the algorithm's global search ability and the diversity of the exploration space, ensuring that the optimal hyperparameter configuration is found during the training process.
[0070] Step S3, setting a weight penalty term related to the weight of the tool wear prediction model, and adding the weight penalty term to the loss function of the tool wear prediction model to obtain a weight-optimized tool wear prediction model; In this embodiment, by introducing a weight penalty term into the loss function of the tool wear prediction model, the tool wear prediction model is prevented from overfitting. In the process of model training, if the weight value is too large, the tool wear prediction model will easily grasp the noise and outlier features in the training data for learning. The introduced weight penalty term is the product of the sum of the squares of the weights and a coefficient. This will prompt the optimization process to proceed in the direction of making the weights smaller. Smaller weights mean that the model will not rely too much on a certain feature, but will consider the information of each feature more evenly, making the tool wear prediction model perform more smoothly. In this way, the overfitting of the tool wear prediction model to the training data can be effectively reduced, and its generalization ability on unknown data can be enhanced, making the tool wear prediction model more stable and reliable.
[0071] The weight penalty term satisfies the following formula: in, is the weight penalty term, is the regularization coefficient, is the vector dimension of the weight of the tool wear prediction model, is the vector of weights for the tool wear prediction model.
[0072] Step S4, introducing a self-attention mechanism into the weight-optimized tool wear prediction model to obtain an optimized tool wear prediction model; In this embodiment, the self-attention mechanism is introduced into the weight-optimized tool wear prediction model, which can significantly improve the performance and prediction accuracy of the model, thereby obtaining a more optimized tool wear prediction model.
[0073] Although the tool wear prediction model can estimate the tool wear to a certain extent, it has limited ability to capture the long-distance dependencies between data when processing complex and changeable processing data. The introduction of the self-attention mechanism breaks this limitation, which allows the model to perform correlation analysis on data at different positions when processing tool wear-related data sequences.
[0074] Through the query, key and value calculation and attention score allocation of the self-attention mechanism, the model can adaptively focus on the larger key data features in the input data, assign higher weights to these features, and relatively weaken the influence of secondary features. In this way, not only can the complex patterns and potential laws in tool wear data be captured more accurately, but also the ability to predict tool wear status can be effectively improved, so that the optimized tool wear prediction model can more reliably warn tool wear in advance in actual production applications, providing strong support for the reasonable arrangement of tool replacement time, improving production efficiency and reducing production costs.
[0075] Step S5: training and evaluating the optimized tool wear prediction model using the sample data, and predicting tool wear using the optimized tool wear prediction model based on the evaluation result.
[0076] The step of training the optimized tool wear prediction model using the sample data specifically includes the following sub-steps: Step S501: assign attention weights to the feature data in the sample data according to the self-attention mechanism to form weighted sample data.
[0077] According to the self-attention mechanism, allocating attention weights to the feature data in the sample data to form weighted sample data includes: Step S50101, performing linear transformation on the feature data to obtain a query matrix, a key matrix and a value matrix.
[0078] In this embodiment, the feature data is a multi-dimensional information set reflecting the tool wear condition. In order to enable the model to more effectively mine the potential relationship between these data, a linear transformation is required, which is achieved with the help of a learnable weight matrix.
[0079] For the feature data matrix composed of feature data, three different weight matrices are used to multiply it to obtain the query matrix, key matrix and value matrix. These three matrices have different functions. The query matrix is used to explore the relationship between data in subsequent steps; the key matrix cooperates with the query matrix to determine the importance of different data elements by calculating the similarity; the value matrix is the carrier that actually contains the feature information and is used to aggregate information according to the degree of association. Through such linear transformation, the foundation is laid for the subsequent calculation of the self-attention mechanism, so that the model can more accurately capture the long-distance dependency in the tool wear feature data and improve the accuracy of tool wear prediction.
[0080] Step S50102, calculating the attention score of the feature data according to the query matrix and the key matrix.
[0081] In this embodiment, in the tool wear prediction model based on the self-attention mechanism, calculating the attention score of the feature data based on the query matrix and the key matrix is an extremely critical link. First, each query vector in the query matrix is dot-producted with all the key vectors in the key matrix, because the dot product can measure the similarity between vectors. For the i-th query vector in the query matrix and the j-th key vector in the key matrix, their dot product results reflect the degree of attention of the data at the i-th position to the data at the j-th position. All these dot product results constitute the initial attention score matrix. It can also be scaled and processed according to needs in the future, and finally an attention score that can accurately reflect the degree of correlation between the feature data is obtained, providing a basis for the subsequent accurate prediction of tool wear.
[0082] The attention score satisfies the following formula: Score for attention, is the query matrix, is the bond matrix, is the permutation of the bond matrix, is the dimension of the key matrix.
[0083] Step S50103, introduce the Softmax function, and use the Softmax function to convert the attention score into the attention weight.
[0084] In this embodiment, after the attention scores are calculated based on the query matrix and the key matrix, the Softmax function is introduced to convert these scores into meaningful attention weights. The Softmax function has a unique advantage that it can map any real number attention score to the (0,1) interval, and the sum of all attention weights is 1.
[0085] Specifically, for the attention score corresponding to each query vector, the weight corresponding to each key vector is calculated by the Softmax function. The larger the weight value, the more important the feature information represented by the corresponding key vector is under the query. This transformation enables the model to weight the feature data according to importance, so as to focus more on key information in subsequent processing and improve the accuracy and efficiency of the model in tool wear prediction. Step S50104, obtain weighted feature data according to the attention weight and the value matrix, and form the weighted sample data according to the weighted feature data.
[0086] The weighted feature data satisfies the following formula: is the weighted feature data, is the query matrix, is the bond matrix, is the value matrix, is the permutation of the bond matrix, is the dimension of the key matrix.
[0087] Step S502: Using the KAN network, extract the nonlinear change features in the weighted sample data to obtain first feature data.
[0088] In this embodiment, the KAN network is constructed based on the Kolmogorov-Arnold theorem and has a strong nonlinear mapping capability, which can decompose the complex multidimensional nonlinear relationship in the weighted sample data into a series of simple single-dimensional functions. When the weighted sample data is input into the KAN network, it can process the data in a unique way. The network can adaptively fit the complex patterns in the data through a series of basis function combinations.
[0089] When processing weighted sample data, the neurons of the KAN network are activated and adjusted according to the dynamic changes of the data. It can identify nonlinear features in the data at different scales and dimensions.
[0090] After the KAN network carefully analyzes and extracts features from the weighted sample data, the first feature data is finally obtained. These data contain key nonlinear information in the tool wear process and are a more in-depth and accurate description of the tool wear state. The first feature data provides the core basis for the subsequent establishment of a more accurate tool wear prediction model, which helps to manage tools more scientifically in actual production and improve production efficiency and product quality.
[0091] Step S503: Use the three-layer long short-term memory network to extract the temporal features of the first feature data to obtain second feature data.
[0092] In this embodiment, after the first feature data is acquired, its time series features are extracted with the help of a three-layer long short-term memory network to obtain more valuable second feature data.
[0093] The three-layer LSTM network processes the first feature data in a progressive manner. Each layer of LSTM units has an input gate, a forget gate, and an output gate. These gating mechanisms can control the inflow, retention, and outflow of information. When processing the first feature data, the first layer of LSTM first captures the shallower temporal patterns in the data; the second layer further explores more complex and long-term dependencies on this basis, and the third layer integrates and deepens the features extracted by the previous two layers.
[0094] After being processed by the three-layer LSTM network, the time series features closely related to the change of tool wear over time can be accurately extracted from the first feature data, and the second feature data can be finally obtained. These data reflect the dynamic process of tool wear more comprehensively and deeply, providing strong support for the accurate prediction of subsequent tool wear status, and helping to take measures such as tool replacement in time in actual production to ensure the smooth progress of production.
[0095] Step S504: using the second feature data to adjust the weight of the optimized tool wear prediction model until the training is completed.
[0096] The step of adjusting the weight of the optimized tool wear prediction model by using the second characteristic data specifically includes the following sub-steps: Step S50401: perform prediction based on the second feature data using the optimized tool wear prediction model to obtain a prediction result.
[0097] In this embodiment, the second characteristic data is input into the optimized tool wear prediction model, and the optimized tool wear prediction model outputs a prediction result.
[0098] Step S50402, calculating the loss value of the loss function according to the prediction result.
[0099] In this embodiment, the loss value of the loss function is a quantitative indicator to measure the degree of difference between the model prediction result and the true label. Specifically, the prediction result and the true value are substituted into the mathematical expression of the loss function to calculate the loss value of the loss function.
[0100] Step S50403, calculating the gradient of the loss function relative to the weight of the optimized tool wear prediction model according to the loss value.
[0101] In this embodiment, the loss value is obtained by comparing the tool wear predicted by the model with the actual tool wear through the loss function, which can intuitively reflect the degree of deviation between the current model prediction and the actual situation. The essence of calculating the gradient of the loss function relative to the model weight is to explore how the loss function changes with the model weight. The error information between the prediction and the actual situation contained in the loss value is the key basis for calculating this change. In the process of calculating the gradient, the error size reflected by the loss value will affect the calculation result, determine the direction and amplitude of the gradient, and thus derive the gradient of the loss function we need relative to the weight of the optimized tool wear prediction model.
[0102] Specifically, starting from the loss value, the partial derivatives of the loss function with respect to the weights of the optimized tool wear prediction model are calculated layer by layer using the chain rule to obtain the gradient of the weights of the optimized tool wear prediction model.
[0103] In this embodiment, starting from the loss value, the gradient of the loss function with respect to the model weight is calculated layer by layer using the chain rule. Specifically, for the weight, the partial derivative of the loss function with respect to the weight is calculated.
[0104] Step S50404, respectively calculating the weighted moving average of the gradient and the weighted moving average of the square of the gradient to obtain a first estimate and a second estimate.
[0105] In this embodiment, the weighted moving average is a statistical method that gives a higher weight to recent data and a lower weight to distant data. In the gradient optimization scenario, the weighted moving average can be used to smooth the fluctuation of the gradient, making the model update more stable. First, the gradient estimate is initialized to 0.
[0106] The first estimate satisfies the following formula: in, For the first estimate, is the attenuation rate, usually close to , is the weighted moving average of the gradient of the previous time step, is the gradient.
[0107] The second estimate satisfies the following formula: in, For the second estimate, is the attenuation rate, which is usually close to , is the weighted moving average of the square of the gradient at the previous time step, is the gradient.
[0108] Step S50405: adjusting the weight of the optimized tool wear prediction model according to the first estimation and the second estimation.
[0109] The step of adjusting the weight of the optimized tool wear prediction model according to the first estimation and the second estimation specifically includes the following sub-steps: Step S5040501, obtaining an optimization weight according to the first estimation and the second estimation.
[0110] In this embodiment, since the gradient is estimated to be 0 during initialization, the first estimate and the second estimate obtained in the early stage will also be biased towards 0, so the first estimate and the second estimate need to be corrected to satisfy the following formula: in, is the first estimate after bias correction, is the first estimate without bias correction, is the first estimated decay rate of Power, is the time step.
[0111] is the second estimate after bias correction, is the second estimate without bias correction, is the second estimated decay rate of Power, is the time step.
[0112] The optimization weight satisfies the following formula: For the optimization weight, is the current weight of the optimized tool wear prediction model, is the current learning rate of the optimized tool wear prediction model, For the first estimate, For the second estimate, To prevent extremely small numbers with a denominator of 0.
[0113] Step S5040502: Use the optimization weight to adjust the weight of the optimized tool wear prediction model.
[0114] In this embodiment, the weight of the optimized tool wear prediction model is replaced by the optimization weight to adjust the weight and gradually optimize the performance of the optimized tool wear prediction model.
[0115] In an optional embodiment, the root mean square error RMSE, the mean absolute error MAE and the mean absolute percentage error MAPE are used as the criteria for determining the accuracy of the model.
[0116] The root mean square error satisfies the following formula: The mean absolute error satisfies the following formula: The mean absolute percentage error satisfies the following formula: In the above formula, n is the number of samples in the validation set. is the predicted tool wear value, is the actual tool wear value.
[0117] In an optional embodiment, the sample data is divided into 6 parts, namely C1, C2, C3, C4, C5 and C6, and C1, C4 and C6 are divided into training set and test set at a ratio of 8:2.
[0118] The following table is obtained through the above evaluation and evaluation of other models: As can be seen from the table, the LSTM model performs well in the mean absolute error MAPE of tool wear prediction only on the c1 dataset and the mean absolute percentage error MAPE of tool wear prediction on the c4 dataset, and the variance is the largest in other cases. This shows that the LSTM model has poor stability and accuracy, and is prone to problems such as gradient disappearance and underfitting during training. In addition, it can be seen from the figure that the LSTM model will have large differences in some tool passes when predicting tool wear conditions, and the maximum errors are 17.31 microns, 16.74 microns, and 18.92 microns on the c1, c4, and c6 datasets, respectively.
[0119] Compared with the LSTM model, the KAN-LSTM model has the ability to extract all features. Table 6 shows that the KAN-LSTM model performs well in the root mean square error RMSE and mean absolute error MAE of tool wear prediction on data sets c1, c4 and c6, but performs poorly in the mean absolute percentage error MAPE. This shows that the KAN-LSTM model can capture local feature patterns, but has limitations in predicting the numerical scale and range of tool wear. It cannot predict targets with a large range of changes well, and may not perform as well as other tool wear prediction models that focus more on time feature optimization in terms of long time series dependence, and may lose some important time information.
[0120] As can be seen from the figure, the LSTM model has large differences in predicting tool wear conditions in some tool passes, with the largest errors reaching 10.42 μm, 11.64 μm, and 12.37 μm on datasets c1, c4, and c6, respectively.
[0121] The accuracy and fluctuation error of the tool flank loss prediction value of the optimized tool wear prediction model are better than those of the LSTM model and KAN-LSTM model, and it has a better overall prediction effect.
[0122] In summary, the present invention combines the complex function decomposition capability of the KAN network, the long-term dependency information capture capability of the LSTM, and the weighting capability of the attention mechanism for key features, and provides an efficient prediction model for nonlinear, multivariate, and time series features in the tool wear process. Under this framework, the KAN network uses the Kolmogorov-Arnold representation theorem to decompose complex multidimensional nonlinear relationships into a series of simple one-dimensional functions, thereby effectively reducing the complexity of modeling. Then, the self-attention mechanism enhances the model's attention to the key moments of tool wear and improves the prediction performance by weighting the importance of the features. The LSTM retains the long-term dependency information in the sequence data through its gating mechanism, which is suitable for describing the process of gradual accumulation of tool wear over time.
[0123] By comparing performance indicators such as MSE, RMSE, MAPE and recognition accuracy, it can be found that the proposed KAN-AM-LSTM neural network model has more advantages in prediction accuracy than the LSTM model and KAN-LSTM model. The KAN-LSTM model shows lower values in the two error metrics of MSE and RMSE, indicating that it is superior to the traditional LSTM and KAN-LSTM architectures in terms of error convergence speed and overall error level. In terms of MAPE (mean absolute percentage error), the error ratio of the optimized tool wear prediction model is significantly reduced, indicating that it is more robust to data of different scales and helps to predict tool wear more stably.
[0124] The optimized tool wear prediction model reduces the complexity of nonlinear modeling through KAN function decomposition, LSTM efficiently processes time series features, and the attention mechanism focuses on key features to reduce unnecessary computational overhead, thereby significantly shortening the model training time.
[0125] like Figure 2 As shown, the present invention also provides a tool wear prediction system, comprising: a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute relevant steps of a relevant embodiment of a tool wear prediction method in the previous aspect of the present invention.
[0126] In the tool wear prediction system provided by the present invention, each functional component can be integrated into one processing component, or each component can exist physically separately, or two or more components can be integrated into one component. The above-mentioned integrated components can be implemented in the form of hardware or in the form of software functions, further improving the overall applicability and practical application ability of the present invention.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
Claims
1. A tool wear prediction method, characterized in that: The method comprises: Acquire milling data related to tool milling, and pre-process the milling data to obtain sample data; The tool wear prediction model was constructed using KAN network, three-layer long short-term memory network, two fully connected layers and random inactivation layer; Setting a weight penalty term related to the weight of the tool wear prediction model, and adding the weight penalty term to the loss function of the tool wear prediction model to obtain a weight-optimized tool wear prediction model; Introducing a self-attention mechanism into the weight-optimized tool wear prediction model to obtain an optimized tool wear prediction model; The optimized tool wear prediction model is trained and evaluated using the sample data, and according to the evaluation result, the optimized tool wear prediction model is used to predict tool wear.
2. A tool wear prediction method according to claim 1, characterized in that: The preprocessing of the milling data to obtain sample data comprises: Using the third quartile method to truncate the milling data to obtain first data; Performing an outlier analysis on the first data, and adjusting the first data according to a result of the outlier analysis to obtain second data; Performing noise reduction processing on the second data by using a wavelet threshold noise reduction method to obtain third data; Performing time domain and frequency domain analysis on the third data to obtain a plurality of feature data; Using the Spearman correlation coefficient to perform a phase relationship analysis on the characteristic data to obtain the most relevant characteristic data; The most relevant feature data is normalized to obtain the sample data.
3. A tool wear prediction method according to claim 2, characterized in that: The performing outlier analysis on the first data and adjusting the first data according to the result of the outlier analysis to obtain the second data comprises: A sliding window mechanism is introduced to slide and sample the first data using a window to obtain sampled data, wherein the sampled data includes a plurality of data points; Calculate the median of the sampled data, and calculate the absolute deviation of the data point to the median through the median; The data point is compared with the absolute deviation, and the data point is adjusted according to a result of the comparison, so as to obtain the second data.
4. A tool wear prediction method according to claim 1, characterized in that: The sample data includes a plurality of feature data, and the training of the optimized tool wear prediction model using the sample data includes: According to the self-attention mechanism, allocating attention weights to the feature data in the sample data to form weighted sample data; Using the KAN network to extract nonlinear change features in the weighted sample data to obtain first feature data; Extracting the time series features of the first feature data using the three-layer long short-term memory network to obtain second feature data; The weight of the optimized tool wear prediction model is adjusted using the second feature data until the training is completed.
5. A tool wear prediction method according to claim 4, characterized in that: The allocating attention weights to the feature data in the sample data according to the self-attention mechanism to form weighted sample data includes: Performing linear transformation on the feature data to obtain a query matrix, a key matrix and a value matrix; Calculating an attention score of the feature data according to the query matrix and the key matrix; Introducing a Softmax function, and using the Softmax function to convert the attention score into the attention weight; Weighted feature data is obtained according to the attention weight and the value matrix, and the weighted sample data is formed according to the weighted feature data.
6. A tool wear prediction method according to claim 5, characterized in that: The weighted feature data satisfies the following formula: is the weighted feature data, is the query matrix, is the bond matrix, is the value matrix, is the permutation of the bond matrix, is the dimension of the key matrix.
7. A tool wear prediction method according to claim 4, characterized in that: The step of adjusting the weight of the optimized tool wear prediction model by using the second feature data includes: According to the second characteristic data, using the optimized tool wear prediction model to perform prediction to obtain a prediction result; According to the prediction result, calculating the loss value of the loss function; Calculating the gradient of the loss function relative to the weight of the optimized tool wear prediction model according to the loss value; respectively calculating a weighted moving average of the gradient and a weighted moving average of the square of the gradient to obtain a first estimate and a second estimate; The weights of the optimized tool wear prediction model are adjusted according to the first estimate and the second estimate.
8. A tool wear prediction method according to claim 7, characterized in that: The adjusting the weight of the optimized tool wear prediction model according to the first estimation and the second estimation includes: Obtaining an optimization weight according to the first estimation and the second estimation; Using the optimization weight to adjust the weight of the optimized tool wear prediction model; The optimization weight satisfies the following formula: For the optimization weight, is the current weight of the optimized tool wear prediction model, is the current learning rate of the optimized tool wear prediction model, For the first estimate, For the second estimate, To prevent extremely small numbers with a denominator of 0.
9. A tool wear prediction method according to claim 4, characterized in that: The first characteristic data satisfies the following formula: is the approximate function of the weighted sample data, For KAN network The number of nodes in the layer, To connect to the KAN network Layer neurons and Layer The activation function of a neuron is For KAN network The number of nodes in the layer, is the number of nodes in the second layer of the KAN network, To connect to the second layer of the KAN network neurons and the third layer The activation function of a neuron is is the number of nodes in the first layer of the KAN network, To connect to the first layer of the KAN network neurons and the second layer The activation function of a neuron is is the number of nodes in the 0th layer of the KAN network, To connect to the 0th layer of the KAN network neurons and the first layer The activation function of a neuron is is the weighted sample data A quantity.
10. A tool wear prediction system, characterized in that: include: A processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute a tool wear prediction method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Tool wear amount prediction method based on self-attention mechanism and depth learning
CN110355608A
Cutter wear and life prediction method and device
CN113043073A
Milling tool wear monitoring method based on wavelet noise reduction and attention mechanism fused GRU network
CN114619292A
Cutter wear prediction method based on improved subtraction optimizer combined with improved bidirectional long-short term memory neural network
CN117549139A
Intelligent tool wear state monitoring method and system based on improved dung beetle algorithm
CN118536391A