Full-period cutting load and tool wear coupling evolution prediction method based on deep learning

By using an improved CNN-LSTM neural network and an Encoder-Decoder framework, combined with cutting load signals and tool wear values, a load-wear mapping model is established. This solves the problem of insufficient prediction accuracy in existing technologies, achieves accurate prediction of tool wear, and improves production efficiency and product quality.

CN120911034APending Publication Date: 2025-11-07NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511109730.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing deep learning-based tool wear prediction methods fail to effectively consider the coupling relationship between cutting load and tool wear, resulting in insufficient prediction accuracy and a lack of interpretability, making it difficult to meet the requirements for accuracy and understandability in production.

Method used

An improved CNN-LSTM neural network based on a multi-head self-attention mechanism is adopted, combined with an Encoder-Decoder framework, to establish a load prediction model and a load-wear mapping model. By acquiring cutting load signals and corresponding tool wear values, signal feature extraction and correlation analysis are performed to achieve accurate prediction of tool wear.

Benefits of technology

It achieves dynamic correlation between tool wear and cutting load, improves the accuracy and interpretability of prediction, and can accurately predict tool wear throughout the entire cycle, thereby improving production efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911034A_ABST
    Figure CN120911034A_ABST
Patent Text Reader

Abstract

The invention discloses a full-period cutting load and tool wear coupling evolution prediction method based on deep learning, and the method comprises the steps: firstly obtaining a multi-element load signal and wear data in a cutting process, and obtaining time domain and frequency domain features strongly related to a wear behavior; secondly, spatial-temporal features in the load and wear data are extracted through a convolutional neural network and a long-short term memory neural network, a memory mechanism is introduced to pay attention to global important information, mapping of signal features and tool wear is achieved, and therefore a load-wear mapping model is established; and meanwhile, through matching of the wear value and the cutting force data, a load prediction model based on a coding-decoding architecture is trained, load prediction is realized, and wear is further predicted by using a load-wear mapping model. According to the method, the coupling evolution characteristics of the cutting load and the tool wear in the CFRP milling process are fully considered, and accurate prediction of the cutting load and the tool wear progress in the machining period is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of tool wear prediction, and particularly relates to a full-cycle cutting load and tool wear coupling evolution prediction method based on deep learning. BACKGROUND

[0002] Under the environment of intelligent manufacturing factory, data monitoring and automatic service are widely applied in the production process, real-time monitoring of the production link is realized, and the production line yield is effectively improved and the labor cost is reduced. However, in the actual production scene, the tool wear problem becomes a key factor affecting the production efficiency and product quality. In the production process, the tool frequently causes poor machining surface quality due to exceeding the wear threshold, which seriously hinders the on-time delivery of products. In order to avoid the temporary shutdown of the production line caused by tool failure, the front-line operators have to pay attention to the state of the tool in use at all times, and replace the tool when it approaches the wear threshold according to their own experience. This method not only depends on the experience of the operators, but also is difficult to accurately judge, and cannot fundamentally solve the problem. If the real-time monitoring and accurate prediction of the tool wear amount in the machining process can be realized, the tool can be replaced in time at the appropriate time, thereby significantly improving the production efficiency. Taking milling as an example, the rake and flank faces of the tool will gradually wear during the machining process. With the continuous accumulation of wear, the cutting force increases significantly, the cutting temperature rises, and then the vibration phenomenon is caused, which directly leads to the reduction of workpiece machining precision and the deterioration of machining surface quality.

[0003] In recent years, with the rapid development of sensor technology and artificial intelligence technology, significant progress has been made in the field of tool wear prediction, and a complete process from machining process information acquisition to wear prediction has been established. At present, the tool wear prediction method based on deep learning network has become a research hotspot, but most of the existing methods focus on the tool wear value itself, ignoring the coupling relationship between the wear amount and the load of the tool in the full-cycle cutting process. This limitation affects the accuracy of the prediction results, and the prediction results obtained based on the deep learning method lack a certain interpretability, which is difficult to meet the dual demands of accuracy and understandability of tool wear prediction in actual production. Therefore, it is of important practical significance to study the prediction method considering the coupling evolution of cutting load and tool wear. SUMMARY

[0004] In view of the problem that the existing technology only depends on the wear value and leads to insufficient prediction accuracy, the application provides a full-cycle cutting load and tool wear coupling evolution prediction method based on deep learning, which considers the influence law of the coupling of the cutting load and the wear of the tool in the whole machining cycle, and realizes the accurate prediction of the tool wear in the machining process by establishing a load prediction model and a load-wear mapping model.

[0005] In order to achieve the above technical purposes, the following technical solutions are specifically adopted in the present application:

[0006] In one aspect of the present application, a deep learning-based full-cycle cutting load and tool wear coupling evolution prediction method is provided, comprising the following steps:

[0007] S1: Obtain the cutting load signal and the corresponding tool wear value, and perform noise reduction preprocessing on the cutting load signal;

[0008] S2: Extract time domain features and frequency domain features from the preprocessed cutting load signal, and filter the signal features that are strongly correlated with tool wear through correlation analysis;

[0009] S3: Construct an improved CNN-LSTM neural network based on a multi-head self-attention mechanism, and train a load-wear mapping model using the filtered signal feature data and the corresponding tool wear value;

[0010] S4. Construct a load prediction model based on an Encoder-Decoder framework, obtain predicted cutting load features, map the predicted cutting load features to tool wear values through the load-wear mapping model, and output tool wear prediction results.

[0011] In one embodiment, the preprocessing in step S1 uses a Butterworth filter to perform noise reduction processing on the cutting load signal.

[0012] In one embodiment, the correlation analysis in step S2 includes:

[0013] S21: Calculate the Pearson correlation coefficient of the time domain / frequency domain features and the tool wear value, and filter the features with an absolute correlation coefficient greater than a preset threshold:

[0014]

[0015] where X represents time domain / frequency domain feature data; Y represents tool wear history data; x i represents the time domain / frequency domain feature data at the i-th time; y i represents the tool wear history data at the i-th time; n represents the sample size; cov(X,Y) represents the covariance of X and Y; σ X σ Y represents the standard deviation of X and Y;

[0016] S22: Perform maximum and minimum normalization on the filtered time domain / frequency domain features.

[0017] In one embodiment, the convolution calculation formula in the improved CNN-LSTM neural network in step S3 is:

[0018] c i = T([x i :x i+f-1 ]*k i + b i )

[0019] C = [c1, c2, …, c i+f-1 ]

[0020] wherein k i represents a convolution kernel; I represents a width of the convolution kernel; f represents a convolution scale; T represents an activation function; b i represents a bias vector; x[x i :x i+f-1 ] represents the i-th to the (i+f-1)-th matrix vector in x, and C represents a feature map output by the convolution layer.

[0021] In one embodiment, the input gate i t , the forget gate f t , and the output gate o t of the LSTM layer in the improved CNN-LSTM neural network in step S3 are as follows:

[0022] i t = s(W i [h t-1 , x t ]+ b i )

[0023] f t = s(W f [h t-1 , x t ]+ b f )

[0024] o t = s(W o [h t-1 , x t ]+ b o )

[0025] wherein W i , W f , and W o represent weight matrices, respectively; h t-1 represents a hidden state at the previous time step; x t represents an input at the current time step; b i , b f , and b o represent bias matrices, respectively; and s represents a Sigmoid activation function.

[0026] In one embodiment, the multi-head attention calculation in step S3 is as follows:

[0027]

[0028] head i = Attention(CW i q ,CW i k ,CW i ν ),1≤i≤h

[0029] H = Concat(head1, head2,..., head h )W

[0030] where Q represents a query matrix; K represents a key matrix; V represents a value matrix; respectively represent a query projection matrix, a key projection matrix, and a value projection matrix; d k represents the dimension of the query array and the key array; h represents the number of attention heads; W represents a linear transformation matrix; and C represents a feature map of convolution output.

[0031] In one embodiment, the step S4 of acquiring the predicted load feature comprises:

[0032] S41: constructing a data set using the continuous tool wear values within a time series and the cutting load signals at the corresponding time;

[0033] S42: sequentially reading the constructed data set through the LSTM unit of the load prediction model, and updating the hidden state according to the following formula:

[0034] h (t) = f(h (t-1) ,x t )

[0035] where h t-1 represents the hidden state at the previous time; and x t represents the input at the current time step;

[0036] When reading to the end of the sequence, a summary context vector c of the input sequence is generated;

[0037] S43: based on the context vector c and the output y t-1 at the previous time, updating the hidden state and predicting the wear value sequence of the future m steps according to the following formula:

[0038] h (t) = f(h (t-1) ,y t-1 ,c).

[0039] In one embodiment, the mapping of the predicted cutting load feature to the tool wear value in step S4 is realized by the load-wear mapping model, which is represented as:

[0040]

[0041] where n represents the prediction time step; f t represents the load at time t; y t represents the wear at time t; M1 represents the load prediction model, and M2 represents the load-wear mapping model.

[0042] The beneficial effects of the present application are:

[0043] The present application simultaneously mines the space-time features in load and wear through convolutional neural networks and long short-term memory neural networks, introduces a multi-head attention mechanism to focus on important information in the continuous cutting process, establishes a mapping relationship between signal features and tool wear, realizes the dynamic correlation between wear and load, and establishes a prediction model based on an Encoder-Decoder framework, fuses the wear features and load coupling factors in the cutting process, constructs a multi-dimensional feature information matrix, provides data support for accurate prediction of tool load and wear values, breaks through the limitations of traditional methods that only focus on predicting tool wear itself, and thus achieves more accurate prediction results. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 is a flowchart of the cutting load and tool wear coupling evolution prediction method of the embodiments of the present application;

[0045] Figure 2 is a structural diagram of the CNN-LSTM of the embodiments of the present application;

[0046] Figure 3 is an effect diagram of the coupling evolution prediction method of the embodiments of the present application. DETAILED DESCRIPTION

[0047] The technical solutions of the present application will be described in detail below with specific embodiments, but those skilled in the art will understand that the following described embodiments are part of the embodiments of the present application, not all the embodiments, and are only used to illustrate the present application, and should not be regarded as limiting the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0048] In one embodiment of the present application, a deep learning-based full-cycle cutting load and tool wear coupling evolution prediction method is provided, which specifically includes the following steps:

[0049] S1: Obtain a cutting load signal and a corresponding tool wear value, and perform noise reduction preprocessing on the load signal.

[0050] Specifically, a cutting load signal collected by a force sensor and a corresponding tool wear value are obtained, and the cutting load signal is subjected to noise reduction processing. In this regard, Dewesoft_X3 force analysis software is used to filter the cutting load signal, and a Butterworth filter is selected.

[0051] In some embodiments, the order N of the Butterworth filter is set to 3, and the cutoff frequency is 500 Hz. The Butterworth filter has a maximum flat amplitude characteristic in the passband, and is characterized by the order N and the cutoff frequency at -3 dB.

[0052] S2: Extract time domain and frequency domain features from the preprocessed cutting load signal, and screen signal features strongly correlated with tool wear through correlation analysis.

[0053] In some embodiments, the noise-reduced cutting load signal is subjected to noise-reduced signal feature extraction from multiple angles such as time domain features and frequency domain features. The time domain features include at least one of the maximum value, the mean value, and the root mean square value; and the frequency domain features include at least one of the center of gravity frequency and the frequency variance.

[0054] In some embodiments, Pearson correlation coefficients are used to associate factors of tool wear values and the time domain / frequency domain features, and X represents time domain and frequency domain feature data, Y represents tool wear values, to obtain signal features strongly correlated with wear behavior. The formula of the Pearson correlation coefficient used is as follows:

[0055]

[0056] wherein x i represents the time domain and frequency domain feature data at the i-th moment; y i represents the tool wear history data at the i-th moment; n represents the sample size; cov(X, Y) represents the covariance of X and Y; σ X σ Y represents the standard deviation of X and Y.

[0057] In some embodiments, the screened time domain / frequency domain features are further subjected to maximum and minimum normalization processing. The calculation formula of the maximum and minimum normalization processing is as follows:

[0058]

[0059] wherein x represents the screened signal features; x min represents the minimum sample value; Xmax This represents the largest sample value; This represents the sample value after normalization.

[0060] S3: Construct an improved CNN-LSTM neural network based on a multi-head self-attention mechanism, and train it using the selected signal features and corresponding tool wear values ​​to obtain a load-wear mapping model. Specifically, this includes the following steps:

[0061] S31: Using the selected signal feature data as the input sequence, local features are extracted through convolutional layers to capture patterns and trends in the input data. Through convolution and pooling operations, the dimensionality of the data is reduced, the computational complexity is lowered, and important features are preserved.

[0062] In a convolutional neural network, the size and number of convolutional kernels are determined based on the data dimension. The convolution calculation formula is:

[0063] c i =T([x i :x i+f-1 ]*κ i +b i )

[0064] C = [c1, c2, ..., c i+f-1 ]

[0065] Among them, κ i Represents the convolution kernel; I represents the width of the convolution kernel; f represents the convolution scale; T represents the activation function; b i Represents the bias vector; x[x i :x i+f-1 ] represents the i-th to (i+f-1)-th matrix vector in x, and C represents the feature map output by the convolutional layer.

[0066] S32: The temporal features of the feature map are learned using a Long Short-Term Memory (LSTM) network, and global information is extracted; the calculation expression of the dual LSTM neural network is as follows:

[0067] Input gate: i t =σ(W i [h t-1 ,x t ]+b i )

[0068] Forgotten Gate: f t =σ(W f [h t-1 ,x t ]+b f )

[0069] Output gate: o t =σ(W o[h t-1 ,x t ]+b o )

[0070] wherein, W i , W f , W o respectively represent weight matrix; h t-1 represents the hidden state of the previous moment; x t represents the input of the current time step; b i , b f , b o respectively represent bias matrix; and σ represents Sigmoid activation function.

[0071] S33: add attention mechanism layer, which is used for dynamically focusing on different parts of time sequence features, helping the model to better capture important time step information.

[0072] Specifically, the multi-head attention mechanism expression is as follows:

[0073]

[0074] wherein, d k represents the dimension of the query array and the key array; Q represents the query matrix; K represents the key matrix; and V represents the value matrix.

[0075] Given the number of attention heads h, the dimension of the output H is obtained: H = h x d v ; wherein, d v represents the dimension of the array.

[0076] The multi-head attention calculation process is as follows:

[0077] head i = Attention(CW i q ,CW i k ,CW i ν ), 1≤i≤h

[0078] H = Concat (head1, head2,..., head h ) W

[0079] wherein, respectively represent the query projection matrix, the key projection matrix, and the value projection matrix, W∈R k / h ; h represents the number of attention heads; W represents the linear transformation matrix; and C represents the feature map of the convolution output.

[0080] In some embodiments, the features output by the hidden layer simultaneously include multi-dimensional features and sequence features of the original time-series data, and a regression layer is used to map these features to tool wear values. The output results are then... With actual wear and tear labels k To evaluate the performance of the improved CNN-LSTM neural network, we compared its root mean square error (RMSE) and mean absolute error (MAE). The calculation formulas are as follows:

[0081]

[0082] Where l represents the number of samples.

[0083] S4: Construct a load prediction model based on the Encoder-Decoder framework, obtain the predicted cutting load features, and realize the mapping from the predicted cutting load features to the tool wear value through the load-wear mapping model, outputting the tool wear prediction result. Specifically, it includes the following steps:

[0084] S41: Construct a training sample dataset using tool wear values ​​and cutting load signals. This dataset contains continuous tool wear values ​​and corresponding cutting load signals over a time series.

[0085] S42: Establish a load prediction model based on the Encoder-Decoder framework.

[0086] In the Encoder phase, an LSTM unit sequentially reads the state of the dataset from step S41 at each time step. During the reading, the hidden state h of the LSTM unit... (t) The following changes occurred:

[0087] h (t) =f(h (t-1) ,x t );

[0088] After reading the end of the sequence (marked by the sequence end symbol), the hidden state is the context summary vector c of the entire input sequence.

[0089] In the Decoder phase, the LSTM unit is trained to predict a given hidden state h. (t) The next symbol y t Generate the output sequence. Unlike the computation in the Encoder stage, the hidden state h at time t is:

[0090] h (t) =f(h (t-1) ,y t-1 ,c).

[0091] S43: considering the coupling evolution process of cutting load and tool wear, inputting the load-wear mapping model data into the load prediction model, and the process is represented as:

[0092]

[0093] wherein n represents a prediction time step; f t represents a load at time t; y t represents wear at time t; M1 represents a load prediction model, and M2 represents a load-wear mapping model.

[0094] Exemplary embodiments

[0095] As shown in the accompanying Figure 1 A deep learning-based CFRP full-cycle cutting load and tool wear coupling evolution prediction method includes the following steps:

[0096] S1. Obtain the tool historical cutting load signal and the corresponding wear value.

[0097] A milling experiment is performed on the tool, and the test data in the processing process is collected through a force sensor installed on the machine tool.

[0098] S2. Preprocess the collected raw data, and the specific implementation steps are as follows:

[0099] S21. Wear signal processing and feature extraction

[0100] The signal collected by the force sensor has a lot of noise signals, and the original signal needs to be denoised. Dewesoft_X3 force analysis software is used to filter the original signal. In order to balance efficiency and signal quality, a Butterworth filter is selected. The Butterworth filter has the maximum flatness characteristic in the passband, and only two parameters are needed to characterize: the order N of the filter and the cutoff frequency ΩC at-3dB. Among them, the order N is 3, and the cutoff frequency ΩC is 500Hz.

[0101] Further, the load signal in the data set is extracted. Because the sampling frequency is high, the amount of signal recorded by the sensor in the cutting experiment is very large, and directly inputting the deep neural network for learning has high requirements on the performance of the computer and the calculation time, which will also lead to too long training time and reduced calculation accuracy. Therefore, the signal is first extracted, and the noise-reduced signal features are extracted from multiple angles such as time domain and frequency domain.

[0102] S22. Correlation analysis

[0103] In order to evaluate the effect of extracting signal features, the relationship between the extracted signal features and wear needs to be analyzed. If the correlation between the extracted signal features and tool wear is weak, the requirement for model training effect is higher when inputting into the deep learning model for learning, and directly inputting the screened strong correlation features will greatly improve the efficiency and accuracy of the model. Highly correlated features are selected by calculating the Pearson coefficient, and the extracted signal features are input into the deep learning model to establish the mapping relationship between the sensor signal features and the tool wear, and to realize the monitoring of the wear value. The Pearson correlation coefficient shows the linear correlation between different features and tool wear, and the formula of the Pearson correlation coefficient is:

[0104]

[0105] In the formula, cov(X, Y) represents the covariance of X and Y; σ X σ Y represents the standard deviation of X and Y; X, Y represents two signal features compared.

[0106] The Pearson correlation coefficient of two compared signal features is positive, which is positive correlation, and is negative, which is negative correlation, and the greater the value is, the higher the correlation is. Therefore, when considering the relationship between the extracted signal features, only the features with large absolute value of the coefficient between the wear and the wear value need to be selected. The features with Pearson correlation coefficient above 0.8 are defined as highly correlated with the wear value.

[0107] Further, the calculation formula of the maximum and minimum normalization processing is as follows:

[0108]

[0109] In the formula, x represents the screened signal feature; x min represents the minimum sample value; X max represents the maximum sample value; represents the normalized sample value.

[0110] S3. After the milling experiment, the tool wear data is obtained as the data label, and then the highly correlated features such as average value, variance, root mean square, maximum value are selected through the above correlation analysis, and the training data set is constructed by matching with the data label, which supports the subsequent model training process.

[0111] S4. An improved CNN-LSTM neural network based on multi-head self-attention mechanism (Multi-Head Self-Attention, MHSA) is established. The structure of CNN-LSTM is shown in the accompanying Figure 2 The features in the feature data set are input into the CNN-LSTM network, and the load-wear mapping model is trained.

[0112] S41. The constructed feature dataset is used as an input sequence, local features are extracted through a convolutional layer to capture patterns and trends in the input data, and the dimensionality of the data is reduced through convolution and pooling operations to reduce computational complexity while preserving important features.

[0113] S42. Then through the LSTM layer, the long-time dependence of the signal in the cutting process is captured.

[0114] The long short-term memory network (LSTM) successfully solves the defect of gradient disappearance during training of the original recurrent neural network, and adds three gate structures to the recurrent neural network, namely the forgetting gate, the output gate and the input gate. The calculation formulas of the three gate coefficients are as follows:

[0115] i t =σ(W i [h t-1 ,x t ]+b i )

[0116] f t =σ(W f [h t-1 ,x t ]+b f )

[0117] o t =σ(W o [h t-1 ,x t ]+b o )

[0118] In the formula, W i , W f , W o represent weight matrices respectively; h t-1 represents the hidden state of the previous moment; x t represents the input of the current time step; b i , b f , b o represent bias matrices respectively; and σ represents a Sigmoid activation function.

[0119] The candidate state value of the current neuron A is determined by the output h t-1 of the previous moment and the input x t of the current moment, then

[0120]

[0121] where W s represents the weight matrix of the current neuron, and b s represents the bias matrix of the current neuron.

[0122] the state value S of the previous time t-1 and the candidate state value of the current time The proportion of the new state value S t is calculated by the forgetting gate f t and the input gate i t As shown in the following formula:

[0123]

[0124] Finally, the output value y t of the current time is calculated:

[0125] y t = o t · tanh(S t ).

[0126] S43. In which the attention mechanism layer is added, which is used to dynamically focus on different parts of the time sequence features, helping the model to better capture important time step information.

[0127] Multi-head Self-attention can capture the information of input sequence in different subspaces at the same time by running multiple Self-Attention layers in parallel and integrating their results, thereby enhancing the expression ability of the model. For each attention head, first calculate its scaled dot-product attention. The input consists of a set of query arrays, key arrays and value arrays, and the scaled dot-product attention is calculated as follows:

[0128]

[0129] In the formula, d k represents the dimension of the query array and the key array.

[0130] Given the number of attention heads h, the dimension of the output H is H = h x d v ; d v represents the dimension of the array.

[0131] The multi-head attention calculation process is as follows:

[0132] head i = Attention(CW i q ,CW i k ,CW i ν ),1≤i≤h

[0133] H = Concat (head1, head2,..., head h ) W

[0134] wherein, W ∈ R k / h , W represents a linear transformation matrix, and C represents a feature map output by a convolutional layer.

[0135] S44. All data are divided into a training data set and a test data set, 80% of the total amount of data is randomly selected as a training set for model training to obtain optimal model parameters, and the remaining data is used as a test data set to evaluate the performance of the model. The model training uses the Adam algorithm to minimize the loss function, and the mean square error (Mean Square Error) is used as the loss function, which is defined as follows:

[0136]

[0137] wherein, n represents the number of samples, y k represents the true wear value, represents the wear value output by the model.

[0138] S45. The finally output features contain multi-dimensional features and sequence features of the original time series data, and the mapping from the features to the tool wear value is realized through a regression layer. The output result is compared with the true wear label y k , and the root mean square error (RMSE) and the absolute mean error (MAE) are calculated to evaluate the effect of the model. The calculation formula is as follows:

[0139]

[0140] S5. A prediction model considering the coupling evolution of cutting load and tool wear is established.

[0141] S51. A load prediction model based on an Encoder-Decoder framework is established. The encoder and decoder kernels are both long short-term memory network (LSTM) units described in the above steps. The Encoder-Decoder framework encodes the global information of the input sequence and decodes to generate the target sequence, and its training process depends on the end-to-end gradient optimization. The input sequence is input into the Encoder one by one, and the hidden state is updated step by step to finally generate the context vector. The initial hidden state of the Decoder is initialized by the context vector of the Encoder, and the output is generated step by step.

[0142] In the Encoder stage, each time of the input sequence is read in order by an LSTM unit, and when reading, the hidden state h (t) of the LSTM unit will change as follows:

[0143] h (t) = f(h (t-1) , xt );

[0144] After reading the end of the sequence (marked with an end-of-sequence token), the hidden state is the summary c of the entire input sequence.

[0145] The Decoder stage LSTM units are trained to generate an output sequence by predicting the next symbol yt given the hidden state h (t) at time t. Unlike the computation of the Encoder stage, the hidden state h at time t is:

[0146] h (t) = f(h (t-1) , y t-1 , c).

[0147] This model is trained using the load wear dataset, and for each time step, the predicted value is calculated. The true label is the time-step cross-entropy loss (Token-Level Cross-Entropy), and the model parameters are updated using the Adam algorithm. The formula for the time-step cross-entropy loss is as follows:

[0148]

[0149] In the formula, T represents the target sequence length; b t represents the true label at the t-th time step; P(b t ) represents the probability distribution at the t-th position.

[0150] S52. The continuous multiple load and wear data are input into the encoder module, and the predicted m-step load sequence is output in the decoder module, and then the predicted m-step wear sequence is obtained through the load-wear mapping model, and the process is represented as:

[0151]

[0152] In the formula, m represents the prediction time step; f t represents the load at time t; y t represents the wear at time t; M1 represents the load prediction model operation in S51; and M2 represents the load-wear mapping model operation in S4.

[0153] In order to fully prove the effectiveness of the monitoring model and the prediction model, the competition data of the PHM Association data competition in 2010 is used for experiment and evaluation analysis. The experimental machine tool is Roders-Tech RFM760 high-speed numerical control milling machine, the experimental tool is three-edge tungsten carbide ball end mill, and the cutting material is stainless steel HRC-52. The force signal, vibration signal and acoustic emission signal original time domain signal in the machining process are collected through the force sensor, acceleration sensor and acoustic emission sensor, the signal sampling frequency is 50KHz, each time the tool is cut along the X direction to cut 108mm, which is recorded as a cutting stroke, each tool cuts 315 strokes, and the wear of the tool is recorded after each cutting stroke. The force signal and the corresponding wear value are extracted from the data set as the experimental data set. The experimental data set is used to obtain the prediction wear of the next ten cutting through the prediction model, and the result is shown in Figure 3 .

[0154] Although the embodiments of the present application are described above in combination with the drawings, the present application is not limited to the above-described specific embodiments and application fields, and the above-described specific embodiments are only illustrative and guiding, but not limiting. Those skilled in the art can make many forms under the guidance of the present application and without departing from the scope protected by the claims of the present application, which all belong to the protection of the present application.

Claims

1. A deep learning-based full-cycle cutting load and tool wear coupling evolution prediction method, characterized in that, The method comprises the following steps: S1: acquiring a cutting load signal and a corresponding tool wear value, and performing noise reduction preprocessing on the cutting load signal; S2: extracting time domain features and frequency domain features from the preprocessed cutting load signal, and screening signal features that are strongly correlated with tool wear through correlation analysis; S3: constructing an improved CNN-LSTM neural network based on a multi-head self-attention mechanism, training a load-wear mapping model using the screened signal feature data and the corresponding tool wear value; S4: constructing a load prediction model based on an Encoder-Decoder framework, acquiring predicted cutting load features, mapping the predicted cutting load features to tool wear values through the load-wear mapping model, and outputting tool wear prediction results.

2. The full-cycle cutting load and tool wear coupling evolution prediction method according to claim 1, characterized in that, The preprocessing in step S1 uses a Butterworth filter to perform noise reduction processing on the cutting load signal.

3. The full-cycle coupled evolution prediction method of cutting load and tool wear according to claim 1, characterized in that, The correlation analysis in step S2 comprises: S21: calculating the Pearson correlation coefficient of the time domain / frequency domain features and the tool wear value, and screening the time domain / frequency domain features with an absolute value of the correlation coefficient greater than a preset threshold value; wherein X represents time and frequency domain feature data; Y represents tool wear history data; x i represents time and frequency domain feature data at the i-th time; y i represents tool wear history data at the i-th time; N represents sample size; cov(X, Y) represents covariance of X and Y; σ X σ Y represents standard deviation of X and Y; S22: performing maximum and minimum normalization processing on the screened time domain / frequency domain features.

4. The full-cycle coupled evolution prediction method of cutting load and tool wear according to claim 1, characterized in that, The convolution calculation formula in the improved CNN-LSTM neural network in step S3 is: c i = T([x i : x i+f-1 ]*k i + b i ) C = [c1, c2,..., c i+f-1 ] where, κ i represents a convolution kernel; I represents the width of the convolution kernel; f represents the convolution scale; T represents an activation function; b i represents a bias vector; x[x i :x i+f-1 ] represents the i-th to (i+f-1)-th matrix vectors in the input sequence x, and C represents the feature map output by the convolution layer.

5. The full-cycle coupled evolution prediction method of cutting load and tool wear according to claim 1, characterized in that, Step S3 input gate i of LSTM layer in improved CNN-LSTM neural network t forgetting gate f t output gate o t is: i t = σ(W i [ h t-1 , x t ] + b i ) f t = σ(W f [ h t-1 , x t ] + b f ) o t = σ(W o [ h t-1 , x t ] + b o ) wherein W i , W f , W o represent weight matrices, respectively; h t-1 represents the hidden state at the previous time step; x t represents the input at the current time step; b i , b f , b o represent bias matrices, respectively; and σ represents a Sigmoid activation function.

6. The full-cycle coupled cutting load and tool wear evolution prediction method of claim 1, wherein, The multi-head attention calculation in step S3 is: head i = Attention(CW i q ,CW i k ,CW i ν ), 1≤i≤h H = Concat(headl, head2,..., head h )W wherein Q represents a query matrix; K represents a key matrix; V represents a value matrix; respectively represent a query projection matrix, a key projection matrix, and a value projection matrix; d k represents the dimensions of the query array and the key array; h represents the number of attention heads; W represents a linear transformation matrix; and C represents a feature map output by a convolutional layer.

7. The full-cycle coupled cutting load and tool wear evolution prediction method of claim 1, wherein, The acquisition of the predicted load features in step S4 comprises: S41: using consecutive tool wear values and corresponding cutting load signals at the same time to construct a data set; S42: sequentially reading the constructed data set through the LSTM unit of the load prediction model, and updating the hidden state according to the following formula: h (t) = f(h (t-1) , x t ) where h t-1 denotes the hidden state at the previous time step; x t denotes the input at the current time step; When reading to the end of the sequence, generate a summary context vector c for the input data set; S43: update the hidden state based on the context vector c and the previous time output y t-1 The hidden state is updated and the future m-step wear value sequence is predicted according to the following formula: h (t) = f(h (t-1) ,y t-1 ,c).

8. The full-cycle coupled cutting load and tool wear evolution prediction method of claim 1, wherein, The mapping of the predicted cutting load features to tool wear values through the load-wear mapping model in step S4 is represented as: where n denotes the prediction time step; f t denotes the load at time t; y t denotes the wear at time t; M1 denotes a load prediction model, and M2 denotes a load-wear mapping model.