Cutter wear state determination method, device and equipment based on multi-wavelet decomposition
Through the deep learning network model of multi-wavelet decomposition, the traditional method has solved the problem of low interpretability and application limitation in tool wear status monitoring in noise environments, and achieved tool wear status monitoring with high accuracy and reliability in noise environments.
Patent Information
- Application Number
- CN202510408249.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional deep learning methods are limited in tool wear status monitoring in low interpretability and noise environments, so they cannot effectively monitor tool wear status.
The deep learning network model based on multi-wavelet decomposition is adopted, and the multi-scale feature extraction module, global average pooling layer and fully connected layer are combined with linearly gated soft threshold noise reduction function and multi-resolution hybrid layer to extract tool wear state features to improve the interpretability and noise robustness of the model.
Improve the accuracy and reliability of tool wear status monitoring, and can effectively monitor tool wear status in a noisy environment.
Smart Images

Figure CN120257094A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of simulation analysis, and in particular to a method, device and equipment for determining tool wear status based on multi-wavelet decomposition. Background Art
[0002] As the core executive component of CNC machine tools, the tool is subjected to severe mechanical loads during the milling process, and its wear directly affects the processing accuracy, surface quality and equipment safety. When the tool wear reaches a critical value, the dynamic characteristics of the cutting system will cause problems such as deterioration of processing accuracy and excessive surface roughness, and even lead to equipment failure. Therefore, real-time monitoring of tool status and implementation of predictive maintenance can effectively reduce the damage to processing parts caused by tool damage.
[0003] In the field of tool wear status monitoring, traditional deep learning methods are often opaque from input to output, which means that users cannot directly know why the model reaches a certain prediction result. At the same time, the tool processing environment is often accompanied by a lot of noise, which will drown out the key features that reflect tool wear information. Therefore, the model needs a certain degree of noise robustness to cope with the noisy environment in which the tool is milling. At present, there are few interpretable deep models used in tool wear monitoring tasks, which limits the application of deep learning methods in tool wear status monitoring. Summary of the invention
[0004] The present application provides a method, device and equipment for determining tool wear status based on multi-wavelet decomposition, so as to solve the problem that low interpretability and noisy environment limit the application of deep learning methods in tool wear status monitoring.
[0005] In a first aspect, the present application provides a method for determining tool wear state based on multiwavelet decomposition, the method comprising:
[0006] Obtain historical data and corresponding true labels of X-, Y-, and Z-axis force signals and vibration signals of the tool, perform standardization on the historical data and the true labels, and use the standardized historical data and the true labels to construct a training set, a validation set, and a test set;
[0007] Establishing a deep learning network model based on multi-wavelet joint decomposition, and setting hyperparameters for the deep learning network model, wherein the deep learning network model comprises: a plurality of stackable multi-scale feature extraction modules, a global average pooling layer and a fully connected layer, wherein the multi-scale feature extraction module comprises: a wavelet packet encoder, and the wavelet packet encoder comprises: a deep wavelet packet decomposition layer, a multi-resolution feature mixing layer and a channel mixing layer;
[0008] Taking the data of the training set, validation set, and test set as input signals, performing time-frequency feature extraction on the input signals through the multi-scale feature extraction module to obtain a feature vector for tool wear state determination;
[0009] Analyzing and processing the feature vector for tool wear state determination through the global average pooling layer and the fully connected layer, and outputting the tool wear state determination result.
[0010] Optionally, the step of taking the data of the training set, validation set, and test set as input signals, performing time-frequency feature extraction on the input signals through the multi-scale feature extraction module to obtain a feature vector for tool wear state determination includes:
[0011] Using a convolutional-based feature block embedding layer to perform convolutional-based feature block embedding on the input signal;
[0012] Performing feature extraction on the input signal after convolutional-based feature block embedding through the deep wavelet packet decomposition layer to obtain a corresponding wavelet time-frequency diagram;
[0013] Using a linear gated soft threshold denoising function to perform soft threshold denoising processing with an adaptive threshold on the high-frequency sub-bands in the wavelet time-frequency diagram;
[0014] Through the multi-resolution mixing layer, fusing the features in the wavelet time-frequency diagram between different spatial positions and different resolution levels to obtain a first feature vector;
[0015] Performing weighted fusion on the features of each channel of the first feature vector through the channel mixing layer to obtain a second feature vector.
[0016] Optionally, the construction process of the deep wavelet packet decomposition layer includes:
[0017] Initializing two sets of learnable three-dimensional tensors with a shape of (D, 1, K) according to the uniform distribution on (-1, 1) as custom wavelet filters, where D is the number of input signal channels and K is the wavelet kernel size, and dividing the tensors by for scaling;
[0018] Creating two globally shared one-dimensional convolutional neural networks as wavelet decomposers, with the convolutional kernels being the two defined sets of filters and and setting the convolutional stride to 2;
[0019] Setting the number of convolutional kernel groups to the number of input signal channels, and applying a depthwise separable one-dimensional convolutional neural network to independently decompose the data of each channel.
[0020] Optionally, the feature extraction of the input signal after embedding the convolutional-based feature block by the deep wavelet packet decomposition layer to obtain the corresponding wavelet time-frequency diagram includes:
[0021] Zero-padding the features input to the wavelet decomposer, and the wavelet decomposer recursively decomposes the padded data to obtain low-frequency coefficients and high-frequency coefficients, and applies a linear gated soft threshold function to the decomposed high-frequency coefficients for noise reduction processing until the maximum decomposition depth is reached;
[0022] Stack the decomposed low-frequency coefficients and the high-frequency coefficients after noise reduction processing according to the wavelet packet tree structure to obtain the corresponding wavelet time-frequency diagram.
[0023] Optionally, the linear gated soft threshold noise reduction function is defined as:
[0024]
[0025] where σ is the standard deviation calculation operation, is the element-wise multiplication operation, s is the sub-band signal of the input function, N is the length of the input signal, τ is the adaptive noise threshold, μ is the gating coefficient, and λ is the adaptive threshold weight coefficient, and β is the bias.
[0026] Optionally, the multi-resolution mixing layer includes: a first mixing layer and a second mixing layer. By the multi-resolution mixing layer, the features in the wavelet time-frequency diagram are fused between different spatial positions and different resolution levels to obtain a first feature vector, including:
[0027] Concatenate the wavelet time-frequency diagram matrix with the input signal, and add position encoding to obtain a concatenated feature map matrix;
[0028] Input the concatenated feature map matrix into the first mixing layer to obtain a first mixed time-frequency diagram matrix, where the first mixing layer is used to mix the information between different resolution levels of the wavelet time-frequency diagram;
[0029] Transpose the first mixed time-frequency diagram matrix and then input it into the second mixing layer for mixing, and then transpose the mixed matrix again to obtain a second time-frequency diagram matrix, where the second mixing layer is used to mix the wavelet coefficients within each resolution level of the wavelet time-frequency diagram;
[0030] Construct a residual connection between the wavelet time-frequency diagram matrix and the second mixed time-frequency diagram matrix to obtain a residual feature matrix;
[0031] Take the mean of all the features of the residual feature matrix and perform normalization processing to obtain a first feature vector.
[0032] Optionally, the method further includes: constructing an additional training loss function according to the biorthogonal wavelet kernel construction condition and the vanishing moment condition to impose an additional constraint on the learning and optimization process of the wavelet kernel.
[0033] In a second aspect, the present application provides a tool wear state determination device based on multi-wavelet decomposition, and the device includes:
[0034] An acquisition module, configured to acquire historical data and corresponding true labels of the tool X, Y, and Z axis force signals and vibration signals, perform normalization processing on the historical data and the true labels, and construct a training set, a validation set, and a test set by using the normalized historical data and the true labels;
[0035] A construction module, configured to establish a deep learning network model based on multi-wavelet joint decomposition, and perform hyperparameter setting on the deep learning network model. The deep learning network model includes: multiple stackable multi-scale feature extraction modules, a global average pooling layer, and a fully connected layer. The multi-scale feature extraction module includes: a wavelet packet encoder, and the wavelet packet encoder includes: a deep wavelet packet decomposition layer, a multi-resolution feature mixing layer, and a channel mixing layer;
[0036] A processing module, configured to use the data of the training set, the validation set, and the test set as input signals, perform time-frequency feature extraction on the input signals through the multi-scale feature extraction module, and obtain a feature vector for tool wear state determination;
[0037] The processing module is further configured to perform analysis and processing on the feature vector for tool wear state determination through the global average pooling layer and the fully connected layer, and output a tool wear state determination result.
[0038] In a third aspect, the present application provides a tool wear state determination device based on multi-wavelet decomposition, including:
[0039] A memory;
[0040] A processor;
[0041] Wherein, the memory stores computer execution instructions;
[0042] The processor executes the computer execution instructions stored in the memory to implement the tool wear state determination method based on multi-wavelet decomposition as described in the first aspect and various possible implementation manners of the first aspect.
[0043] In a fourth aspect, the present application provides a computer storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the tool wear state determination method based on multi-wavelet decomposition as described in the first aspect and various possible implementation manners of the first aspect.
[0044] The present application provides a method, device, and equipment for determining the tool wear state based on multi-wavelet decomposition. The method obtains the historical data and corresponding true labels of the tool X-axis, Y-axis, and Z-axis force signals and vibration signals, performs standardized processing on the historical data and true labels, and constructs a training set, a validation set, and a test set using the standardized historical data and true labels; establishes a deep learning network model based on multi-wavelet joint decomposition, and sets hyperparameters for the deep learning network model. The deep learning network model includes: multiple stackable multi-scale feature extraction modules, a global average pooling layer, and a fully connected layer. The multi-scale feature extraction module includes: a wavelet packet encoder, and the wavelet packet encoder includes: a deep wavelet packet decomposition layer, a multi-resolution feature mixing layer, and a channel mixing layer; uses the data of the training set, validation set, and test set as input signals, extracts time-frequency features of the input signals through the multi-scale feature extraction module to obtain feature vectors for determining the tool wear state; analyzes and processes the feature vectors for determining the tool wear state through the global average pooling layer and the fully connected layer, and outputs the result of determining the tool wear state, improving the accuracy and reliability of the output result. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0046] Figure 1 Schematic flow of the method for determining the tool wear state based on multi-wavelet decomposition provided by an embodiment of the present application Figure 1 ;
[0047] Figure 2 Structural diagram of the deep learning network model for multi-wavelet joint decomposition provided by an embodiment of the present application;
[0048] Figure 3 Architectural diagram of the multi-resolution feature mixing module provided by an embodiment of the present application;
[0049] Figure 4 Architectural diagram of the multi-resolution mixing layer provided by an embodiment of the application;
[0050] Figure 5 Architectural diagram of the multi-head attention layer provided by an embodiment of the present application;
[0051] Figure 6 Structural schematic diagram of the device for the method for determining the tool wear state based on multi-wavelet decomposition provided by an embodiment of the present application;
[0052] Figure 7 Structural schematic diagram of the equipment for the method for determining the tool wear state based on multi-wavelet decomposition provided by an embodiment of the present application.
[0053] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Invention
[0054] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0055] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein.
[0056] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0057] In the tool state monitoring task, the characteristics of the tool wear signal have cross-scale distribution characteristics. For example, the chatter characteristics, the impact energy of chipping, and the wear accumulation trend characteristics are mainly distributed in the frequency band regions at different scales. In order not to miss the key characteristics related to tool wear, the model needs to consider multiple scale information simultaneously.
[0058] However, using too many wavelet decomposition layer coefficients will lead to redundant model inputs, increasing the computational burden and the risk of overfitting. Using too few layer coefficients will not achieve effective time-frequency analysis results. Even within the same decomposition layer, not every segment of coefficients can equally reflect the wear characteristics. The cutting process of the tool is carried out under the coupling of multiple parts of the machine tool, and the monitoring signal is disturbed by many factors. For example, during the tool machining process, the vibration signal collected by the acceleration sensor is vulnerable to multi-source interference factors: on the one hand, the vibration of the coupled mechanical components and the characteristics of the tool structure itself will introduce signal disturbances; on the other hand, workpiece impacts, chips, cutting fluid flow, and non-stationary rotation of the spindle during the machining process will also introduce disturbances. These interferences cause a large number of frequency domain components unrelated to the tool wear state to be mixed in the monitoring signal.
[0059] To address the above problems, this application proposes a method for determining the tool wear state based on multi-wavelet decomposition. This method addresses the above problems by embedding wavelet analysis and multi-head self-attention in the model, and connecting a multi-perspective wavelet decomposition layer, a multi-resolution hybrid layer, and a channel feature hybrid layer that apply a linear gated soft-threshold function in series into a multi-scale hybrid module for extracting deeper features, ensuring the interpretability and noise robustness of the model.
[0060] The following uses specific embodiments to detail the technical solutions of this application and how the technical solutions of this application solve the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.
[0061] Figure 1 Flow schematic of a method for determining the tool wear state based on multi-wavelet decomposition provided by the embodiments of this application Figure 1 As Figure 1 shown, the method for determining the tool wear state based on multi-wavelet decomposition provided in this embodiment includes:
[0062] S1: Obtain the historical data of the tool X, Y, and Z axis force signals and vibration signals and their corresponding true labels, perform normalization processing on the historical data and true labels, and construct a training set, a validation set, and a test set using the normalized historical data and true labels.
[0063] S2: Establish a deep learning network model based on multi-wavelet joint decomposition, and perform hyperparameter settings on the deep learning network model. The deep learning network model includes: multiple stackable multi-scale feature extraction modules, a global average pooling layer, and a fully connected layer. The multi-scale feature extraction module includes: a wavelet packet encoder, and the wavelet packet encoder includes: a deep wavelet packet decomposition layer, a multi-resolution feature hybrid layer, and a channel hybrid layer.
[0064] Figure 2 This is the structural diagram of the deep learning network model for multi-wavelet joint decomposition provided by the embodiments of the present application. Figure 3 This is the architecture diagram of the multi-scale feature mixing module provided by the embodiments of the present application.
[0065] As Figure 2 and Figure 3 shown, the Depth Wavelet Decomposition layer (DepthWPConv) is mainly composed of two deep one-dimensional convolutional neural networks. When performing wavelet packet layer-by-layer decomposition, the globally shared filters h[n] and g[n] are used at each resolution level. Among them, 1x1Conv refers to a one-dimensional convolutional layer with a convolutional kernel size of 1, Gated Softshrink refers to the linear gating soft threshold function proposed by the present invention, GELU refers to the GELU activation function, GLU refers to Gated Linear Units, and Norm refers to the normalization operation through the BatchNorm1D function. A multi-scale feature mixing module (MixResBlock) provided in this embodiment extracts multi-scale features by a Depth Wavelet Decomposition layer (DepthWPConv) and a linear gating threshold function (GatedSoftshrink), mixes the extracted time-frequency map matrices by a multi-resolution mixing layer (MixresLayer), and uses a GLU layer to mix the feature information between different channels.
[0066] Specifically, the depth wavelet packet decomposition layer, the multi-resolution feature mixing layer, and the channel mixing layer are connected in series to form a wavelet packet encoder WaveletEncoder for extracting the time-frequency features of the input . At the same time, a multi-scale feature mixing module composed of a wavelet packet encoder, a normalization layer, and a residual connection is constructed. Stack several layers of multi-scale feature mixing modules, and the original tool wear sensor signal (where C represents the number of channels of the original input signal C = 6, that is, the force signals and vibration signals in the X, Y, and Z directions, and L represents the length of the original input signal, taking L = 4096), after being embedded into non-overlapping feature blocks with the number of input channels C, the number of output channels D, and the convolutional kernel and convolutional step size both P based on the one-dimensional convolutional neural network 1DCNN, obtain (where N = L / P, taking P = 64, D = 128), and feed into the MixResBlock mainly composed of WaveletEncoder to obtain the tensor for tool wear state prediction, and predict the result through the global average pooling layer and the fully connected layer.
[0067] It is understandable that the model hyperparameters include the learning rate, learning rate decay factor, maximum number of iterations, wavelet kernel size, depth of wavelet decomposition layer, regularization term coefficient, etc. The stackable multi-scale feature extraction module is used to extract features, and this multi-scale feature extraction module is mainly composed of a wavelet packet encoder, a normalization operation, and a residual connection.
[0068] It is understandable that the feature vector after the input is embedded by the feature block, after extracting features through the multi-scale feature extraction module and then normalizing, a feature matrix for predicting the tool wear state is obtained. At the same time, a residual connection is used between the stacked multi-layer multi-scale feature extraction modules to ensure stable gradient propagation.
[0069] S3: Use the data of the training set, validation set, and test set as input signals, and perform time-frequency feature extraction on the input signals through the multi-scale feature extraction module to obtain a feature vector for determining the tool wear state.
[0070] S31: Use the convolutional-based feature block embedding layer to perform convolutional-based feature block embedding on the input signal.
[0071] S32: Perform feature extraction on the input signal after convolutional-based feature block embedding through the deep wavelet packet decomposition layer to obtain the corresponding wavelet time-frequency diagram.
[0072] Specifically, for a deep wavelet packet decomposition layer of an embedding model, its construction process specifically includes the following steps:
[0073] Initialize two sets of learnable three-dimensional tensors with the shape of (D, 1, K) according to the uniform distribution on (-1, 1) as custom wavelet filters, where D is the number of input signal channels and K is the wavelet kernel size, and divide the tensors by for scaling to prevent unstable training caused by too large gradients;
[0074] Create two globally shared one-dimensional convolutional neural networks as wavelet decomposers, and the convolutional kernels are the two sets of defined filters and To fit the wavelet decomposition process, set the convolutional stride to 2;
[0075] To process the input signal with D channels Set the number of convolutional kernel groups to D, and apply a depthwise separable one-dimensional convolutional neural network to independently decompose the data of each channel. Here, the input signal refers to the signal after feature block embedding.
[0076] Furthermore, for this deep wavelet packet decomposition layer, its decomposition process is based on the Mallat algorithm and includes the following steps:
[0077] Before wavelet decomposition, zero-padding of length K - 2 is performed on the right side of the feature X input to the wavelet decomposer to obtain the padded data X padded ;
[0078] After padding, the wavelet decomposer decomposes the padded data X recursively according to and to obtain the low-frequency coefficients padded and the high-frequency coefficients and and apply the linear gating soft threshold function to the decomposed b j,n for noise reduction to obtain b thresholded j,n ;
[0079] Repeat the above steps until the maximum decomposition depth
[0080] Stack all the low-frequency coefficients a j,n and the noise-reduced high-frequency coefficients b thresholded j,n according to the wavelet packet tree structure to obtain a two-dimensional real-valued matrix of length N and width J constituting the time-frequency feature map.
[0081] S33: Use the linear gating soft threshold denoising function to perform soft threshold denoising with an adaptive threshold on the high-frequency subbands in the wavelet time-frequency map.
[0082] Specifically, use an adaptive soft threshold to denoise the input and control whether to apply the function to denoise the input through a learnable routing mechanism. Its function is defined as:
[0083]
[0084] σ is the standard deviation calculation operation, is the element-wise multiplication operation, s is the subband signal of the input function, N is the length of the input signal, τ is the adaptive noise threshold, μ is the gating coefficient, where are the adaptive threshold weight coefficient and bias respectively;
[0085] The method controls the threshold size using the adaptive threshold weighting coefficient λ and bias β on the basis of the threshold . Among them, λ and β are learnable parameters and are initialized according to the uniform distribution in the interval (-0.5, 0.5).
[0086] The method uses μ to control the noise reduction process. When μ is 0, the threshold fails, and the function will perform an identity mapping on the input. When μ is 1, the function will normally perform noise reduction on the input. Among them, μ is calculated by a globally shared 1x3 convolutional network, a Sigmoid activation function, and an adaptive pooling layer according to the high-frequency coefficient b of the input. j,n Calculated.
[0087] S34: Through the multi-resolution mixing layer, the features in the wavelet time-frequency diagram are fused between different spatial positions and different resolution levels to obtain the first feature vector.
[0088] Figure 4 This is the architecture diagram of the multi-resolution mixing layer provided by the embodiment of the present application. Figure 5 This is the architecture diagram of the multi-head attention layer provided by the embodiment of the present application.
[0089] As Figure 4 shown, the multi-resolution mixing layer (MixresLayer) includes two multi-head attention (Multi-Head-Attention, MHA) layers facing the resolution dimension and the wavelet coefficient dimension, a pooling layer (AvePool) that averages all channel information, and also includes a normalization layer (Norm) and a residual connection used to ensure gradient stability.
[0090] As Figure 5 shown, the multi-head attention layer includes a multi-head attention mechanism (Multi-HeadAttention), a feed-forward neural network (Feed-Forward), and layer normalization (LayerNorm).
[0091] Specifically, two MHA layers applicable to different dimensions, one for mixing information between different resolution levels of the wavelet time-frequency diagram, called the res-mixing layer, and one for mixing wavelet coefficients within each resolution level of the wavelet time-frequency diagram, called the coef-mixing layer;
[0092] The wavelet time-frequency feature map matrix is concatenated with the input and after adding positional encoding, we get where the definition of Position Embedding is to encode the position information for each position in the input sequence so that the model can recognize the order of the input elements;
[0093] The is first input into the res-mixing layer and mixed according to the correlation relationship of each resolution level to obtain the mixed time-frequency diagram matrix F′;
[0094] The mixed time-frequency diagram matrix F′ obtained by transposing (Transpose) Send it to the coef-mixing layer, mix it according to the correlation between the wavelet coefficients of the same resolution, and transpose the mixed matrix to obtain the mixed time-frequency map matrix
[0095] In order to ensure the stability of the gradient, a residual connection is constructed between the input time-frequency matrix F and the mixed time-frequency matrix F″ to obtain F res =F stack +F″;
[0096] right All features of the J+1 dimensions representing the resolution level are averaged and normalized using the BatchNorm1D function to obtain the vector It can be used as a matrix that can represent res The global features of all resolution level information are used for subsequent prediction tasks, where
[0097] Furthermore, when performing wavelet decomposition on the input signal x, the method uses a convolution-based patch embedding operation to perform non-overlapping local block cutting on the input signal x and linearly map it into an embedding vector, mapping the input to a space that is more conducive to subsequent learning while reducing the sequence length and the computational overhead of the self-attention mechanism, so that the multi-resolution hybrid layer can process high-dimensional signals more efficiently.
[0098] S35: Perform weighted fusion on the features of each channel of the first feature vector through a channel mixing layer to obtain a second feature vector.
[0099] It can be understood that in order to mix the features of different channels according to their importance, this embodiment uses a gated linear unit (GLU) to perform adaptive fusion of cross-channel features. Specifically, first, the wavelet time-frequency domain features extracted from each sensor channel are tensor stacked to form a three-dimensional feature matrix; then, the feature channels are linearly projected through a learnable weight matrix to generate two sets of feature maps with the same dimension; finally, a nonlinear transformation is applied to one set of maps using a sigmoid activation function, and then the two sets of maps are element-by-element multiplied to achieve adaptive weighting of the feature channels, and a second feature vector after weighted fusion is obtained.
[0100] Compared with the conventional weighted average or simple series method, the GLU mechanism achieves dual advantages through the differentiable gating structure: first, the gating unit can autonomously suppress invalid feature channels; second, it can retain the strength of feature responses that are sensitive to wear, thereby constructing a fused feature vector in the feature space with strong noise robustness and outstanding wear characterization capabilities.
[0101] S4: Analyze and process the feature vector for tool wear state determination through a global average pooling layer and a fully connected layer, and output the result of tool wear state determination.
[0102] It can be understood that the global average pooling layer performs a global average operation on each channel of the feature vector, thereby generating a one-dimensional vector with the same number of channels as the number of channels. The global average pooling layer can reduce the number of parameters and the computational complexity of the model, which helps to accelerate the training and inference speed of the model.
[0103] Furthermore, the fully connected layer maps the feature vector output by the global average pooling layer to the classification space of the tool wear state. By learning the weight parameters, the fully connected layer converts the information in the feature vector into the classification probability of the tool wear state. At the output layer, the fully connected layer converts the classification probability into a specific tool wear state label through the softmax function or other classifiers, and then the model outputs the result of the tool wear state determination according to the input feature vector.
[0104] It can be understood that only when the wavelet structure matches the structure of the input signal can the wavelet decomposition efficiently extract effective features. When using a single and fixed wavelet coefficient, the wavelet structure cannot be dynamically adjusted according to the input signal characteristics and the prediction task, and the perspective of feature extraction is single, which is likely to form a bottleneck in the representation of tool wear feature information. In an optional embodiment, a method for constructing an adaptive wavelet filter bank is proposed. The core idea is to construct a wavelet filter bank with learnable coefficients, and dynamically optimize and adjust the wavelet filter bank coefficients by combining the signal characteristics and the task prediction target, so as to break through the time-frequency characteristic constraints of the traditional fixed wavelet basis.
[0105] It can also be understood that two groups of unconstrained filters cannot specifically decompose the input signal x into low-frequency components and high-frequency components. At the same time, the learned filter bank may not necessarily form a wavelet filter. To solve this contradiction, an additional training loss function is added during training to constrain the learned filter coefficients. Orthogonal wavelets are widely used in time-frequency feature extraction for various tasks because of their orthogonality. During the signal decomposition process, there is no redundancy between sub-bands, and they have the characteristics of high computational efficiency and good energy preservation. However, on the one hand, the strict orthogonality of orthogonal wavelets leads to low design flexibility, resulting in a small optimization space for learnable wavelet filters. On the other hand, except for the Haar wavelet, orthogonal wavelets do not have symmetry, which will cause time-domain misalignment of sub-band signals, resulting in linear phase distortion, and the boundary effect will increase with the increase of the decomposition scale, causing decomposition errors, directly affecting the reliability of subsequent processing (such as feature extraction, compression, denoising), and affecting the quality of feature extraction. To achieve better training results, considering symmetry, orthogonality, and design flexibility, this embodiment proposes to use biorthogonal wavelet constraints instead of orthogonal wavelet constraints.
[0106] Specifically, the condition for the biorthogonal filter bank to achieve exact reconstruction of the input signal is:
[0107]
[0108] where is used to eliminate the aliasing distortion caused by downsampling in the wavelet decomposition process. is used to ensure energy conservation after the signal is decomposed and reconstructed.
[0109] where represents the frequency-domain coefficients of the low-pass filter for wavelet decomposition, h[n] represents the time-domain coefficients of the low-pass filter for wavelet decomposition, and the low-pass filter for wavelet decomposition is used to extract the approximation coefficients in the signal decomposition stage, H * (ω) represents taking the complex conjugate of H(ω), and the complex conjugate operation corresponds to time reversal in the time domain. Specifically,
[0110] Furthermore, represents the frequency-domain coefficients of the low-pass filter for wavelet synthesis, represents the time-domain coefficients of the low-pass filter for wavelet synthesis, and the low-pass filter for wavelet synthesis uses the coefficients obtained by wavelet decomposition to reconstruct the signal.
[0111] Similarly, represents the frequency-domain coefficients of the high-pass filter for wavelet decomposition, g[n] represents the time-domain coefficients of the high-pass filter for wavelet decomposition, and the high-pass filter for wavelet decomposition is used to extract the detail coefficients in the signal decomposition stage, G * (ω) represents taking the complex conjugate of G(ω),
[0112] Even further, represents the frequency-domain coefficients of the high-pass filter for wavelet synthesis, represents the time-domain coefficients of the high-pass filter for wavelet synthesis, and the high-pass filter for wavelet synthesis uses the coefficients obtained by wavelet decomposition to reconstruct the signal.
[0113] In addition, H(ω + π) represents the frequency response of the filter H(ω) at the phase of π, and the rest are similar. The appearing in the filter formula is used to normalize the energy of the filter to ensure energy conservation.
[0114] It can be understood that if g[n], h[n] and constitute a filter bank that satisfies biorthogonality and can be completely reconstructed, the following conditions need to be met:
[0115] where g[n] and need to satisfy high-pass characteristics, h[n] and It is necessary to satisfy the low-pass property and also
[0116] We can obtain:
[0117]
[0118] Combining the biorthogonality condition, the perfect reconstruction condition, the low-pass and high-pass conditions, and the vanishing moment order condition, when choosing to use a biorthogonal wavelet filter bank, it is necessary to use an additional loss function composed of the following constraints:
[0119]
[0120] Among them, is used to ensure that the DC gain of the low-pass filter is to ensure energy conservation during signal decomposition and reconstruction;
[0121] is used to ensure that the high-pass filter filters out the DC component, and the sum being zero indicates that it has no response to low frequencies;
[0122] is for the wavelet decomposition filter and the wavelet synthesis filter to be orthogonal under even displacements, ensuring no redundant subbands. < > represents the inner product operation. For example, <A, B> represents the inner product of A and B;
[0123] δ n represents the Kronecker function
[0124] has a zero response at the Nyquist frequency (ω = π), which is used to ensure the low-pass property;
[0125] is used to represent the vanishing moment order of the wavelet filter, ensuring that higher-order polynomial components are filtered out, where H (n) (ω) represents the nth derivative of the frequency response, and n represents the vanishing moment order here;
[0126] h[1 - n] represents shifting one position to the right after inverting the filter time-domain coefficients with respect to n = 0.
[0127] It can be understood that the first and second item formulas of the above constraints are used to constrain the low-pass property of the filters h and and the high-pass property of g and In this embodiment, and are used to represent this part of the loss;
[0128] The third item formula constrains the biorthogonal relationship between h and In this embodiment, Denote this part of the loss, where δ(n) is the Kronecker function. It should be noted that due to the conversion relationship shown in the 6th and 7th formulas of the above constraints, it can be represented by g in this embodiment.
[0129] The 4th formula of the above constraints constrains the relationship of the filter H(ω) at the π phase, which is represented by in this embodiment to represent this part of the loss;
[0130] The 5th formula of the above constraints is used to constrain the vanishing moment order o of the filter, where K is the length of the wavelet filter, which is represented by l vanish = l h-vanish (h, o) + l g-vanish (g, o) to represent this part of the loss, where
[0131] In this embodiment, it is defined that l wavelet = l highpass + l lowpass + l π-phase + l bior , and the overall loss is defined as
[0132] A method for determining the tool wear state based on multi-wavelet joint decomposition provided by an embodiment of the present application. This method constructs a multi-perspective wavelet packet decomposition layer based on the Mallat algorithm. This layer uses a learnable multi-channel wavelet filter and a depth one-dimensional convolutional neural network with a step size of 2 to simulate the wavelet packet decomposition operation under the multi-wavelet perspective, and uses the biorthogonal wavelet construction condition and the vanishing moment condition to constrain the learning process of the filter. By constructing a linear gated soft threshold function, the gating mechanism and the adaptive threshold are used to denoise the key high-frequency signals. By constructing a multi-resolution hybrid layer based on the MHA mechanism, the time-frequency feature map is weighted and mixed in the wavelet coefficient dimension and the resolution dimension according to the context correlation. The GLU is used to construct a channel feature hybrid layer to fuse the information of each channel. The multi-perspective wavelet decomposition layer, the multi-resolution hybrid layer, and the channel feature hybrid layer applying the linear gated soft threshold function are connected in series into a multi-scale hybrid module, and the model can stack this module multiple times to extract deeper features. This method embeds wavelet analysis and multi-head self-attention in the model, and stacks the above modules, pursuing the design concept of improving performance by stacking layers in deep learning while ensuring the interpretability and noise robustness of the model.
[0133] Figure 6 It is a structural schematic diagram of a device for determining the tool wear state based on multi-wavelet decomposition provided by an embodiment of the present application. As Figure 6As shown in the figure, the tool wear state determination device 600 based on multi-wavelet decomposition provided in this embodiment includes:
[0134] An acquisition module 601, configured to acquire historical data of the tool X, Y, and Z axis force signals and vibration signals and corresponding true labels, perform normalization processing on the historical data and the true labels, and construct a training set, a validation set, and a test set by using the normalized historical data and true labels;
[0135] A construction module 602, configured to establish a deep learning network model based on multi-wavelet joint decomposition, and perform hyperparameter setting on the deep learning network model. The deep learning network model includes: a plurality of stackable multi-scale feature extraction modules, a global average pooling layer, and a fully connected layer. The multi-scale feature extraction module includes: a wavelet packet encoder, and the wavelet packet encoder includes: a deep wavelet packet decomposition layer, a multi-resolution feature mixing layer, and a channel mixing layer;
[0136] A processing module 603, configured to use the data of the training set, the validation set, and the test set as input signals, perform time-frequency feature extraction on the input signals through the multi-scale feature extraction module, and obtain a feature vector for tool wear state determination;
[0137] The processing module 603 is further configured to perform analysis and processing on the feature vector for tool wear state determination through the global average pooling layer and the fully connected layer, and output a tool wear state determination result.
[0138] The tool wear state determination device based on multi-wavelet decomposition provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0139] Figure 7 It is a structural schematic diagram of a tool wear state determination device based on multi-wavelet decomposition provided in an embodiment of the present application. As Figure 7 shown, the tool wear state determination device based on multi-wavelet decomposition provided in the present application, the tool wear state determination device 700 based on multi-wavelet decomposition includes: a receiver 701, a transmitter 702, a processor 703, and a memory 704.
[0140] The receiver 701 is configured to receive instructions and data;
[0141] The transmitter 702 is configured to send instructions and data;
[0142] The memory 704 is configured to store computer execution instructions;
[0143] A processor 703 is configured to execute computer-executable instructions stored in a memory 704 to implement each step performed by the method for determining a tool wear state based on multi-wavelet decomposition in the above embodiments. For specific details, reference may be made to the relevant descriptions in the embodiments of the method for determining a tool wear state based on multi-wavelet decomposition described above.
[0144] Optionally, the above-mentioned memory 704 may be either independent or integrated with the processor 703.
[0145] When the memory 704 is independently provided, the electronic device further includes a bus for connecting the memory 704 and the processor 703.
[0146] This application also provides a computer storage medium storing computer-executable instructions, which, when executed by a processor, implement the method for determining a tool wear state based on multi-wavelet decomposition performed by the device for determining a tool wear state based on multi-wavelet decomposition as described above.
[0147] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be executed by several physical components in cooperation. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium.
[0148] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0149] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for determining the tool wear state based on multi-wavelet decomposition, characterized in that, The method includes: Obtaining historical data of the X, Y, and Z axis force signals and vibration signals of the tool and corresponding true labels, performing normalization processing on the historical data and the true labels, and constructing a training set, a validation set, and a test set by using the normalized historical data and the true labels; Establishing a deep learning network model based on multi-wavelet joint decomposition, and performing hyperparameter setting on the deep learning network model. The deep learning network model includes: multiple stackable multi-scale feature extraction modules, a global average pooling layer, and a fully connected layer. The multi-scale feature extraction module includes: a wavelet packet encoder, and the wavelet packet encoder includes: a deep wavelet packet decomposition layer, a multi-resolution feature mixing layer, and a channel mixing layer; Using the data of the training set, the validation set, and the test set as input signals, performing time-frequency feature extraction on the input signals through the multi-scale feature extraction module to obtain a feature vector for tool wear state determination; Performing analysis and processing on the feature vector for tool wear state determination through the global average pooling layer and the fully connected layer, and outputting a tool wear state determination result.
2. The method according to claim 1, characterized in that The step of using the data of the training set, the validation set, and the test set as input signals, performing time-frequency feature extraction on the input signals through the multi-scale feature extraction module to obtain a feature vector for tool wear state determination includes: Using a convolutional-based feature block embedding layer to perform convolutional-based feature block embedding on the input signals; Performing feature extraction on the input signals after convolutional-based feature block embedding through the deep wavelet packet decomposition layer to obtain corresponding wavelet time-frequency diagrams; Using a linear gated soft threshold denoising function to perform soft threshold denoising processing with an adaptive threshold on the high-frequency subbands in the wavelet time-frequency diagrams; Through the multi-resolution mixing layer, fusing the features in the wavelet time-frequency diagrams between different spatial positions and different resolution levels to obtain a first feature vector; Performing weighted fusion on the features of each channel of the first feature vector through the channel mixing layer to obtain a second feature vector.
3. The method according to claim 2, characterized in that, The construction process of the deep wavelet packet decomposition layer includes: Initialize two learnable three-dimensional tensors of shape (D, 1, K) as custom wavelet filters according to the uniform distribution on (-1, 1), where D is the number of input signal channels and K is the wavelet kernel size, and divide the tensors by for scaling; Create two globally shared one-dimensional convolutional neural networks as wavelet decomposers, with the convolutional kernels being two defined sets of filters and Set the convolutional stride to 2; Setting the number of groups of convolutional kernels to the number of channels of the input signal, and applying a depthwise separable one-dimensional convolutional neural network to independently decompose the data of each channel.
4. The method according to claim 2, wherein The step of performing feature extraction on the input signals after convolutional-based feature block embedding through the deep wavelet packet decomposition layer to obtain corresponding wavelet time-frequency diagrams includes: Performing zero-padding on the features input to the wavelet decomposer, and the wavelet decomposer recursively decomposing the padded data to obtain low-frequency coefficients and high-frequency coefficients, and applying a linear gated soft threshold function to the decomposed high-frequency coefficients for denoising processing until the maximum decomposition depth is reached; Stacking the decomposed low-frequency coefficients and the denoised high-frequency coefficients according to the wavelet packet tree structure to obtain corresponding wavelet time-frequency diagrams.
5. The method according to claim 2, wherein The linear gated soft threshold denoising function is defined as: where σ is the standard deviation calculation operation, is the element-wise multiplication operation, s is the subband signal of the input function, N is the length of the input signal, τ is the adaptive noise threshold, μ is the gating coefficient, and λ is the adaptive threshold weight coefficient, and β is the bias.
6. The method according to claim 2, wherein The multi-resolution hybrid layer includes: a first hybrid layer and a second hybrid layer. Through the multi-resolution hybrid layer, the features in the wavelet time-frequency diagram are fused between different spatial positions and different resolution levels to obtain a first feature vector, including: Concatenate the wavelet time-frequency diagram matrix with the input signal, and add position encoding to obtain a concatenated feature map matrix; Input the concatenated feature map matrix into the first hybrid layer to obtain a first mixed time-frequency diagram matrix after mixing, where the first hybrid layer is used to mix information between different resolution levels of the wavelet time-frequency diagram; Transpose the first mixed time-frequency diagram matrix after mixing and input it into the second hybrid layer for mixing, and then transpose the mixed matrix to obtain a second time-frequency diagram matrix, where the second hybrid layer is used to mix the wavelet coefficients within each resolution level of the wavelet time-frequency diagram; Construct a residual connection between the wavelet time-frequency diagram matrix and the second mixed time-frequency diagram matrix after mixing to obtain a residual feature matrix; Calculate the mean of all features of the residual feature matrix and perform normalization processing to obtain a first feature vector.
7. The method according to claim 1, wherein The method further includes: constructing an additional training loss function according to the biorthogonal wavelet kernel construction condition and the vanishing moment condition to impose an additional constraint on the learning and optimization process of the wavelet kernel.
8. A tool wear state determination device based on multi-wavelet joint decomposition, characterized in that The device includes: An acquisition module, configured to acquire historical data and corresponding true labels of the tool X, Y, and Z axis force signals and vibration signals, perform normalization processing on the historical data and the true labels, and construct a training set, a validation set, and a test set by using the normalized historical data and true labels; A construction module, configured to establish a deep learning network model based on multi-wavelet joint decomposition, and perform hyperparameter setting on the deep learning network model. The deep learning network model includes: multiple stackable multi-scale feature extraction modules, a global average pooling layer, and a fully connected layer. The multi-scale feature extraction module includes: a wavelet packet encoder, and the wavelet packet encoder includes: a deep wavelet packet decomposition layer, a multi-resolution feature hybrid layer, and a channel hybrid layer; A processing module, configured to use the data of the training set, the validation set, and the test set as input signals, perform time-frequency feature extraction on the input signals through the multi-scale feature extraction module to obtain a feature vector for determining the tool wear state; The processing module is further configured to analyze and process the feature vector for determining the tool wear state through the global average pooling layer and the fully connected layer, and output a result for determining the tool wear state.
9. An equipment for determining the tool wear state based on multi-wavelet joint decomposition, characterized in that The device includes: A memory; A processor; Wherein, the memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement a method for determining tool wear state based on multi-wavelet joint decomposition according to any one of claims 1-7.
10. A computer storage medium, characterized in that, The computer storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement a method for determining tool wear state based on multi-wavelet joint decomposition according to any one of claims 1-7.