Tool wear degree identification method and device based on Mel-ResNet-Transformer

By converting sensor timing data into Mel spectra and using the improved ResNet-Transformer model for feature extraction and fusion, the problem of low fusion efficiency of multimodal data is solved, and efficient and accurate identification of tool wear state is achieved.

CN120431338APending Publication Date: 2025-08-05SHENYANG UNIVERSITY OF TECHNOLOGY +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510497669.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the industrial Internet of Things, multimodal data fusion methods are inefficient, making it difficult for industrial equipment to accurately identify tool wear conditions.

Method used

The Mel-ResNet-Transformer-based method is adopted to convert the sensor timing data into Mel spectra, and feature extraction is performed through the improved ResNet18 model, and feature fusion and classification are performed by combining two-layer Transformer encoders.

Benefits of technology

The accuracy of tool wear recognition is improved, and through the efficient fusion of multimodal data, the ability to identify tool wear status is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431338A_ABST
    Figure CN120431338A_ABST
Patent Text Reader

Abstract

The invention relates to industrial tool wear degree monitoring, in particular to a tool wear degree identification method and device based on Mel-ResNet-Transform. The method is applied to industrial cutter wear state recognition, and the accuracy of cutter wear recognition is improved. Comprising the following steps: S1, converting time sequence data of a sensor into a Mel spectrogram, and realizing modal unification of a one-dimensional signal and a picture; s2, parallel feature extraction: inputting the converted Mel spectrogram and the tool operation image into a feature extraction model for feature extraction, and generating 512-dimensional abstract feature representation; and S3, mixed feature fusion: establishing two layers of Transform encoders, carrying out fusion encoding on the extracted features according to importance, and selecting the feature with the most discriminative property for classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to industrial tool wear degree monitoring, and in particular to a tool wear degree identification method and device based on Mel-ResNet-Transformer. Background Art

[0002] The Industrial Internet of Things (IIoT) provides real-time, detailed process and status data for smart manufacturing through the widespread deployment of sensors. Monitoring systems then return decision-making information to relevant equipment, thereby improving production efficiency and reducing costs. Monitoring systems model and analyze multimodal data to identify possible equipment states. However, complex industrial processes result in data with high-dimensional heterogeneity and complex interactions, posing challenges to multimodal data analysis. To address this challenge, multimodal data fusion integrates data from different distributions, sources, and types into a global space, forming a unified representation. Therefore, identifying tool wear status in industrial equipment can be transformed into a classification task based on multimodal data fusion.

[0003] Traditional multimodal data fusion strategies are categorized into three types: data-level fusion, feature-level fusion, and decision-level fusion. Data-level fusion combines the raw or preprocessed data from each modality before sending it to the model, enabling joint representation learning. This fusion strategy typically directly connects unimodal modalities to exploit cross-modal correlations. However, due to the redundancy and similarity of the raw data, it struggles to deeply extract cross-modal correlations in complex, high-dimensional data. Feature-level fusion combines features extracted from different modalities and sends them to the model for decision-making. Much of the work in feature-level fusion focuses on feature engineering, optimizing the raw data of each modality through mathematical or statistical methods to form a concise and efficient representation. Feature fusion is performed through multiplication, concatenation, cross product, or element-wise addition to extract cross-modal correlations and interactions at different levels. Decision-level fusion combines the individual decisions obtained from each modality to form a final prediction. For example, data from different modalities are fed into their respective sub-neural networks, and the final classification result is determined by a majority voting system. However, this strategy focuses on the competition between models and does not fully utilize the complementary relationship between multimodal data, which makes the equipment unable to accurately identify tool wear.

[0004] In order to make up for the problem that the equipment cannot identify tool wear effectively due to the inadequacy of the fusion method, the present invention introduces deep learning to improve the existing fusion method, so as to solve the problem that the feature fusion method is inefficient and the industrial equipment cannot identify tool wear effectively in an industrial environment. Summary of the Invention

[0005] This invention addresses the issues of data fusion's poor ability to identify tool wear in industrial IoT applications, as well as the low accuracy of single-modality tool wear identification. By providing a tool wear identification method and device based on Mel-ResNet-Transformer, this method improves the accuracy of tool wear identification by applying it to industrial tool wear status identification.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a tool wear degree identification method based on Mel-ResNet-Transformer, comprising the following steps:

[0007] S1. Convert sensor time series data into Mel spectrogram to achieve modal unification of one-dimensional signal and image;

[0008] S2, parallel feature extraction, input the converted Mel spectrum and tool operation image into the feature extraction model (improved ResNet18 model) for feature extraction, and generate a 512-dimensional abstract feature representation;

[0009] S3: Hybrid feature fusion, establish a two-layer Transformer encoder, fuse and encode the extracted features according to their importance, and select the most discriminative features for classification.

[0010] Furthermore, the step S1 includes:

[0011] S1.1. Obtaining time series signals x[n] collected by a vibration sensor, an acoustic wave sensor, and a cutting force sensor, wherein the time series signals represent dynamic change data of vibration, acoustic wave, and cutting force during tool machining;

[0012] S1.2. Perform signal framing processing; split the time series signal x[n] into short time series of length N along the time axis according to the fixed sliding window method, and set overlapping windows for adjacent series;

[0013] S1.3. Apply the Hann window function w[n] to each short time series. The calculation formula is as follows:

[0014]

[0015] S1.4. Perform time-frequency analysis;

[0016] Perform short-time Fourier transform (STFT) on the windowed signal to obtain the time-frequency spectrum representation. The calculation formula is as follows:

[0017]

[0018] Where ω represents the normalized angular frequency, w[n-τ] is the window function, and τ is the center of the window function, which represents the time position of the current analysis.

[0019] S1.5, perform frequency scale conversion;

[0020] The linear Hz frequency f in the extracted frequency domain information Hz Mapping to the nonlinear Mel frequency scale f Mel , where f Hz The calculation formula is:

[0021]

[0022] f Mel The calculation formula is as follows:

[0023]

[0024] S1.6, constructing a Mel filter bank;

[0025] Distribute M triangular filters equally spaced on the Mel frequency scale and calculate the value H of the mth filter m (k); the calculation formula is as follows:

[0026]

[0027] Where k is the frequency, f(m) is the center frequency of the triangular filter, and ∑ m H m (k) = 1;

[0028] S1.7, perform Mel feature extraction;

[0029] Calculate the convolution of the signal power spectrum and the Mel filter bank to obtain the Mel frequency domain feature s(m). The calculation formula is:

[0030]

[0031] The final output is a Mel spectrogram containing time-frequency features.

[0032] Furthermore, the step S2 includes:

[0033] S2.1. Data preprocessing: scaling the Mel spectrum and tool operation image to a uniform size of 224×224 pixels and performing normalization processing;

[0034] S2.2, perform feature extraction:

[0035] The ResNet18 model of IMAGENET1K_V1 pre-trained on the ImageNet dataset is used to extract features from the pre-processed Mel spectrogram and tool operation image. The feature extraction model (improved ResNet18 model) includes:

[0036] The initial convolution layer uses a 7×7 convolution kernel with a stride of 2 and outputs a 112×112×64 dimensional feature map.

[0037] The maximum pooling layer uses a 3×3 pooling window with a step size of 2 and outputs a 56×56×64-dimensional feature map;

[0038] There are four residual block groups, namely Conv2_x, Conv3_x, Conv4_x and Conv5_x. Each residual block group contains at least two residual units, where:

[0039] Conv2_x outputs a 56×56×64 dimensional feature map;

[0040] Conv3_x outputs a 28×28×128-dimensional feature map;

[0041] Conv4_x outputs a 14×14×256 dimensional feature map;

[0042] Conv5_x outputs a 7×7×512-dimensional feature map;

[0043] The global average pooling layer reduces the 7×7×512-dimensional feature map to a 512-dimensional feature vector;

[0044] S2.3. Input the 512-dimensional feature vector into a shared dimensionality reduction layer, reduce the dimensionality to a d-dimensional feature vector through a fully connected layer, where d < 512, and process it through batch normalization and ReLU activation function;

[0045] S2.4. Output the d-dimensional feature vector as input of S3.

[0046] Furthermore, the step S3 includes:

[0047] S3.1. Concatenate n d-dimensional feature vectors along the direction of the number of features. The shape of the concatenated feature matrix X is n×d, where n is the number of features and d is the feature dimension.

[0048] S3.2. Model the feature interaction of the feature matrix X through the multi-head attention mechanism, specifically including:

[0049] a) Define three learnable weight matrices: query weight matrix (Queries) W Q , key weight matrix (Keys) W K , Value weight matrix (Values) W V ,in,

[0050] b) Project the feature matrix X onto the weight matrix to obtain the query matrix Q = XW Q , key matrix K = XWK Sum matrix V = XW V ;

[0051] c) Divide the query matrix, key matrix, and value matrix into h attention heads respectively, Then the dimension of each attention head is n×d k ;

[0052] d) Calculate the scaled dot product attention for each attention head i:

[0053]

[0054] Among them, Softmax(·) is used to convert a real number vector into a probability distribution, reflecting the relative importance of each input to the output; that is, given z = [z1, z2, ... z n ],i∈n

[0055]

[0056] S3.3. Perform feature fusion, specifically including:

[0057] a) Splice the output features of each attention head; specifically: i Splicing is performed to obtain the result of multi-head attention transformation, whose shape is still n×d;

[0058] b) Perform layer normalization on the concatenated features;

[0059] c) Perform residual connection on the normalized features and the input feature matrix X to obtain the intermediate features X1;

[0060] S3.4 performs feedforward enhancement, specifically including:

[0061] a) Project the intermediate feature X1 into 4D space through the fully connected layer;

[0062] b) Restore the 4D-dimensional features to dimension d through the fully connected layer;

[0063] c) Perform layer normalization on the restored features and perform residual connection with the intermediate feature X1 to obtain the final output. The shape of the final output matrix is still n×d;

[0064] S3.5 performs classification output.

[0065] Furthermore, the S3.5 specifically includes:

[0066] a) Input the feature-fused n×d-dimensional feature matrix into the second-layer Transformer encoder and repeat steps S3.2 to S3.4 to further enhance the feature expression capability;

[0067] b) Select the most discriminative eigenvector from the final output feature matrix;

[0068] c) passing the selected feature vectors through the fully connected layer and the softmax classification layer in sequence, and outputting the classification results of the tool wear degree;

[0069] Among them, the structure of the second-layer Transformer encoder is the same as that of the first layer, which is used to deepen the feature fusion effect.

[0070] A tool wear degree recognition device based on Mel-ResNet-Transformer includes: an image capture unit for collecting visual data of the tool wear edge and outputting the tool operation image to a data processing unit;

[0071] The sensor data capture unit is used to collect vibration, sound wave and cutting force signals during the tool processing process and output the conditioned time series data to the data processing unit;

[0072] The data processing unit receives the tool operation image and sensor time series data, and realizes the tool wear status identification.

[0073] Furthermore, the image capture unit includes an image acquisition module, an illumination module, a synchronization control module and a mechanical stabilization module;

[0074] The image acquisition module uses a camera with a resolution of 3.2 microns per pixel and a field of view of 17.8 mm × 14.3 mm to capture microscopic deformation data of the tool wear edge;

[0075] The lighting module is used to ensure that the camera can obtain a clear image;

[0076] The synchronization control module is used to receive the spindle position signal of the CNC machine tool and trigger the image capture action to ensure that the image capture is synchronized with the milling action to avoid motion blur caused by tool rotation;

[0077] The mechanical stabilization module is used to maintain the rigid connection between the camera box and the tool cutting workbench, eliminating image distortion caused by vibration or displacement.

[0078] Furthermore, the sensor data capture unit includes a sensor module, a signal conditioning module, a data conversion module, a data storage module, a processing and transmission module, a synchronization control module, and a power management module; wherein,

[0079] The sensor data capture unit includes a sensor module, a signal conditioning module, a data conversion module, a data storage module, a processing and transmission module, a synchronization control module, and a power management module;

[0080] The sensor module directly senses vibration, sound waves and cutting force through sensitive elements and converts them into electrical signals;

[0081] The signal conditioning module is used to amplify, remove noise and perform impedance matching on the weak electrical signal output by the sensor module to improve the signal quality;

[0082] The data conversion module is used to convert the conditioned analog signal into a digital signal for subsequent processing and storage;

[0083] The data storage module is used to store the data obtained by sampling at different sampling frequencies locally on the client, and the stored data needs to be cleaned to filter out low-quality data for data fusion analysis;

[0084] The processing and transmission module is used to perform data encapsulation and format conversion, and send the data to the host computer or the cloud via wired / wireless means;

[0085] The synchronization control module is used to receive the machine tool spindle position signal to ensure that the sensor data acquisition is synchronized with the processing action and eliminate timing errors;

[0086] The power management module is used to provide stable power supply for each module, optimize energy consumption and prevent overvoltage / overcurrent from damaging hardware.

[0087] Furthermore, the data processing unit includes a signal image conversion module, a feature extraction module, and a feature fusion module;

[0088] The signal image conversion module is used to convert the sensor time series data into Mel spectrogram to achieve modal unification of one-dimensional signal and image;

[0089] The feature extraction module is used to input the converted Mel spectrum and tool operation image into the improved ResNet18 model for feature extraction;

[0090] The feature fusion module is used to fuse and encode the extracted features according to their importance, and the number and dimension of each feature remain unchanged after encoding; finally, the tool wear status is classified based on the fusion result.

[0091] Compared with the prior art, the present invention has beneficial effects.

[0092] 1. This paper proposes a method for unifying multimodal data based on Mel spectrum transformation. This method converts time series signals from linear Hertz frequencies to nonlinear Mel frequencies and maps the Mel spectrum into an image. By leveraging the Mel spectrum's sensitivity to low-frequency signals and its insensitivity to high-frequency signals, this method captures subtle features of device signals within a specific frequency range. This data-driven method can losslessly transform time series information into images, preserving the long-term dependencies of the sequence.

[0093] 2. This paper designs a multimodal feature extraction model based on a transfer learning deep residual network to achieve joint feature extraction from different images. By introducing pre-trained parameters, the model does not need to be trained from scratch, reducing computational overhead and enhancing the ability of features to represent images. It also solves the feature extraction problem caused by sample sparsity in industrial processing tasks with limited data.

[0094] 3. This paper designs a Transformer-based feature fusion model that identifies the impact of different modal features on a task and fully considers the contribution of each feature to the task during fusion. While reducing noise interference, it performs self-attention fusion on global features, achieving efficient fusion of multimodal data and knowledge sharing between tasks, improving the utilization of multimodal data and enabling efficient and accurate identification of tool wear status in industrial equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] The present invention is further described below with reference to the accompanying drawings and specific embodiments. The scope of protection of the present invention is not limited to the following description.

[0096] Figure 1 This is a diagram summarizing the system structure of the present invention;

[0097] Figure 2 Schematic diagram of the system structure of the present invention;

[0098] Figure 3 A schematic diagram of a simulated industrial scenario of the present invention;

[0099] Figure 4 Schematic diagram of the method framework of the present invention;

[0100] Figure 5 Schematic diagram of the identity mapping module of the present invention. DETAILED DESCRIPTION

[0101] In order to make the purpose, technical solutions and beneficial effects of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0102] Example 1: A tool wear degree recognition device based on Mel-ResNet-Transformer includes an image capturing part, a sensor data capturing part and a data processing part.

[0103] The image capture unit includes an image acquisition module, an illumination module, a synchronization control module, and a mechanical stabilization module. The image acquisition module captures microscopic deformation data of the tool's worn edge using high resolution (3.2 microns per pixel) and a 17.8mm x 14.3mm field of view. The illumination module provides uniform illumination in all directions, optimizing the incident angle to suppress reflection interference and enhance the contrast of features in the worn area. The synchronization control module receives the CNC machine tool spindle position signal and precisely triggers the image capture action, ensuring strict synchronization of the timing with the milling action to avoid motion blur. The mechanical stabilization module maintains a rigid connection between the camera box and the workbench, eliminating image distortion caused by vibration or displacement.

[0104] The sensor data capture unit includes a sensor module, a signal conditioning module, a data conversion module, a processing and transmission module, a synchronization control module, and a power management module. The sensor module directly senses physical quantities such as vibration, sound waves, and cutting force through sensitive elements and converts them into electrical signals. The signal conditioning module is used to amplify, remove noise, and impedance match the weak electrical signals output by the sensor to improve signal quality. The data conversion module is used to convert the conditioned analog signals into digital signals for subsequent processing and storage. The data storage module is used to store data obtained by sampling at different sampling frequencies locally on the client. The stored data needs to be cleaned and low-quality data is filtered out for data fusion analysis. The processing and transmission module is used to perform data encapsulation and format conversion, and send the data to the host computer or cloud via wired or wireless means. The synchronization control module is used to receive the machine tool spindle position signal to ensure that sensor data acquisition is strictly synchronized with the processing action and eliminate timing errors. The power management module is used to provide stable power supply for each module, optimize energy consumption, and prevent overvoltage / overcurrent from damaging the hardware.

[0105] The data processing unit includes a feature extraction module and a feature fusion module. The feature extraction module inputs the converted Mel spectrogram and tool operation image into a modified ResNet18 for feature extraction. The feature fusion module fuses and encodes the extracted features according to their importance, maintaining the number and dimension of each feature after encoding. Finally, the most discriminative features are selected to classify the tool wear status.

[0106] Example 2: The method for identifying the degree of tool wear specifically comprises the following steps:

[0107] S1: Data converted into Mel spectrum:

[0108] Convert sensor time series data into Mel spectrograms to achieve modal unification of one-dimensional signals and images to meet the model's feature extraction requirements. By amplifying low-frequency details in the signal, a Mel spectrogram with discriminative properties for different tool wear levels is generated, allowing the model to deeply extract spectrogram features. This includes:

[0109] Power spectrum calculation: Apply STFT to the original signal to mine the signal's time-frequency domain information;

[0110] Mel spectrum mapping: nonlinearly maps the frequency domain unit Hz to amplify low-frequency details in the signal;

[0111] S2: Parallel feature extraction:

[0112] The converted Mel spectrogram and tool operation image are input into ResNet18 for feature extraction. Each image data is extracted into a 512-dimensional abstract feature representation. A shared dimensionality reduction layer is established to integrate and reduce the common information of the features to avoid the increase in computing resources due to excessive dimensionality.

[0113] S3: Hybrid feature fusion:

[0114] A two-layer Transformer encoder is built to fuse and encode the extracted features according to their importance, while preserving the number and dimension of each feature. Finally, the most discriminative features are selected to classify the tool wear state.

[0115] Furthermore, the specific steps of S1 are as follows:

[0116] In the first step, the signal is divided into short time series of length N along the time axis according to the fixed sliding window method, and overlapping windows of adjacent sequences are set.

[0117] In the second step, STFT is applied to each short time series to obtain the spectrum representation of the series at different frequencies and times. The calculation formula is as follows;

[0118]

[0119] Where x[n] is the discrete time series signal to be transformed, ∈ represents the normalized angular frequency, w[n-τ] is the window function, and τ is the center of the window function, indicating the time position of the current analysis. The window function uses a zero-centered Hann window w[n], and the calculation formula for w[n] is as follows.

[0120]

[0121] The third step is to map the Hz frequency in the extracted frequency domain information to the Mel scale. The calculation formula is as follows:

[0122]

[0123] Where f Mel and f Hz represent the frequency in Mel scale and the frequency in Hz scale respectively.

[0124] The fourth step is to convert the Hz frequency to the Mel scale, distribute m triangular filters evenly spaced on the Mel scale, and then convert it to Hertz. The calculation formula is as follows:

[0125]

[0126] The fifth step is to calculate the value of the filter, where the value of the mth filter H m (k) is calculated as follows:

[0127]

[0128] Where k is the frequency, f(m) is the center frequency of the triangular filter, and ∑ m H m (k)=1.

[0129] The sixth step is to extract the signal features by convolving the power spectrum of the signal with the pre-designed Mel filter bank to obtain the bandpass filtering result s(m) in the Mel frequency domain. The calculation formula is as follows:

[0130]

[0131] Where M is the total number of Mel filters. Each point on the Mel spectrogram corresponds to the bandpass filtering result of a short time series at that moment in the frequency band. Its pixel value represents the signal energy intensity at a specific time and frequency. Since a short time series represents a period of time, the position of the moment is taken as the center point of the short time series.

[0132] Furthermore, the specific steps of S2 are as follows:

[0133] The feature extraction algorithm is based on the PyTorch deep learning framework, and the ResNet18 model uses the weight parameters of the first version IMAGENET1K_V1 pre-trained on the ImageNet dataset.

[0134] The first step is to scale the Mel spectrum and tool operation image to 224×224.

[0135] In the second step, the image passes through the first layer, Conv1, which uses a 7×7 convolution kernel, meaning that each convolution operation processes a 7×7 pixel block. This layer generates 64 output channels, and the final output size is 112×112×64, i.e., 64 112×112 feature maps.

[0136] Step 3: The feature map passes through a 3×3 max pooling layer with a stride of 2. This pooling layer is used to reduce the spatial size of the feature map and retain the salient features of each region. The final output size of the feature map is 56×56×64.

[0137] Step 4: The feature map passes through the second layer Conv2, which consists of the following submodules:

[0138] Submodule 1:

[0139] A 3×3 convolutional layer with 64 output channels, stride 1, and padding 1 is used for feature extraction.

[0140] A batch normalization layer that normalizes each channel to reduce internal covariate shift and speed up training.

[0141] An activation function ReLU, defined as f(x) = max(0, x), is used to introduce nonlinearity so that the model can learn more complex features.

[0142] A 3×3 convolutional layer with 64 output channels, stride 1, and padding 1 is used for feature extraction.

[0143] A batch normalization layer that normalizes each channel to reduce internal covariate shift and speed up training.

[0144] Submodule 2:

[0145] A 3×3 convolutional layer with 64 output channels, stride 1, and padding 1 is used for feature extraction.

[0146] A batch normalization layer that normalizes each channel to reduce internal covariate shift and speed up training.

[0147] An activation function ReLU, defined as f(x) = max(0, x), is used to introduce nonlinearity so that the model can learn more complex features.

[0148] A 3×3 convolutional layer with 64 output channels, stride 1, and padding 1 is used for feature extraction.

[0149] A batch normalization layer that normalizes each channel to reduce internal covariate shift and speed up training.

[0150] The feature map undergoes feature extraction in submodule 1, outputting a feature map of the same size as the one that did not pass through submodule 1. This feature map is residually connected to the feature map that did not pass through submodule 1 and serves as the input to submodule 2. This allows the network to learn the difference between input and output, promoting information flow. In Conv2, submodule 2 has the same network structure as submodule 1. The feature map undergoes feature extraction again in submodule 2, outputting a feature map of the same size as the input. This feature map is residually connected to the feature map that did not pass through submodule 2 and serves as the input to the third layer, Conv3. The final feature map output after Conv2 has a size of 56×56×64.

[0151] Step 4: The feature map passes through the third layer Conv3, which consists of the following submodules:

[0152] Submodule 1:

[0153] A 3×3 convolutional layer with 256 output channels, a stride of 2, and a padding of 1.

[0154] A batch normalization layer that performs per-channel normalization.

[0155] An activation function ReLU is used to introduce nonlinearity.

[0156] A 3×3 convolutional layer with 128 output channels, a stride of 1, and padding of 1.

[0157] A batch normalization layer that performs per-channel normalization.

[0158] An identity mapping module, including:

[0159] A 1×1 convolutional layer with 256 output channels, a stride of 2, and a padding of 1.

[0160] A batch normalization layer that performs per-channel normalization.

[0161] Submodule 2:

[0162] A 3×3 convolutional layer with 256 output channels, a stride of 1, and padding of 1.

[0163] A batch normalization layer that performs per-channel normalization.

[0164] An activation function ReLU is used to introduce nonlinearity.

[0165] A 3×3 convolutional layer with 256 output channels, a stride of 1, and padding of 1.

[0166] A batch normalization layer that performs per-channel normalization.

[0167] After entering submodule 1, the input feature map passes through a 3×3 convolutional layer with a stride of 2, a batch normalization layer, an activation function, a 3×3 convolutional layer with a stride of 1, and a batch normalization layer. After passing through the 3×3 convolutional layer with a stride of 2, the feature map size is halved, and the number of channels is doubled. Due to this change in feature map size, the output feature map cannot be residually connected to the input feature map. Therefore, the input feature map must enter the identity mapping module, where the output is the same size as the output feature map. This is then residually connected to the output feature map and serves as the input to submodule 2. Since submodule 2 does not have a stride of 2 convolutional layer, the feature map remains unchanged after passing through each layer. Therefore, the feature map output from submodule 2 is directly residually connected to the input feature map from submodule 2 and serves as the input to Conv4. The final feature map output after Conv3 has a size of 28×28×128.

[0168] Step 5: The feature map passes through the fourth layer Conv4, which consists of the following submodules:

[0169] Submodule 1:

[0170] A 3×3 convolutional layer with 128 output channels, a stride of 2, and a padding of 1. This is used for downsampling to improve the model's ability to recognize large-scale features.

[0171] A batch normalization layer that performs per-channel normalization.

[0172] An activation function ReLU is used to introduce nonlinearity.

[0173] A 3×3 convolutional layer with 128 output channels, a stride of 1, and padding of 1.

[0174] A batch normalization layer that performs per-channel normalization.

[0175] An identity mapping module, including:

[0176] A 1×1 convolutional layer with 128 output channels, a stride of 2, and a padding of 1. This is used to keep the output feature size consistent.

[0177] A batch normalization layer that performs per-channel normalization.

[0178] Submodule 2:

[0179] A 3×3 convolutional layer with 128 output channels, a stride of 1, and padding of 1.

[0180] A batch normalization layer that performs per-channel normalization.

[0181] An activation function ReLU is used to introduce nonlinearity.

[0182] A 3×3 convolutional layer with 128 output channels, a stride of 1, and padding of 1.

[0183] A batch normalization layer that performs per-channel normalization.

[0184] After entering submodule 1, the input feature map passes through a 3×3 convolutional layer with a stride of 2, a batch normalization layer, an activation function, a 3×3 convolutional layer with a stride of 1, and a batch normalization layer. After passing through the 3×3 convolutional layer with a stride of 2, the feature map size is halved, and the number of channels is doubled. Due to this change in feature map size, the output feature map cannot be residually connected to the input feature map. Therefore, the input feature map must enter the identity mapping module, where the output is the same size as the output feature map. This is then residually connected to the output feature map and serves as the input to submodule 2. Since submodule 2 does not have a stride of 2 convolutional layer, the feature map output remains unchanged after passing through each layer. Therefore, the feature map output from submodule 2 is directly residually connected to the input feature map from submodule 2 and serves as the input to Conv5. The final feature map output after Conv4 is 14×14×256 in size.

[0185] Step 6: The feature map passes through the fourth layer Conv5, which consists of the following submodules:

[0186] Submodule 1:

[0187] A 3×3 convolutional layer with 512 output channels, a stride of 2, and a padding of 1. This is used for downsampling to improve the model's ability to recognize large-scale features.

[0188] A batch normalization layer that performs per-channel normalization.

[0189] An activation function ReLU is used to introduce nonlinearity.

[0190] A 3×3 convolutional layer with 512 output channels, a stride of 1, and padding of 1.

[0191] A batch normalization layer that performs per-channel normalization.

[0192] An identity mapping module, including:

[0193] A 1×1 convolutional layer with 512 output channels, a stride of 2, and a padding of 1. This is used to keep the output feature size consistent.

[0194] A batch normalization layer that performs per-channel normalization.

[0195] Submodule 2:

[0196] A 3×3 convolutional layer with 256 output channels, a stride of 1, and padding of 1.

[0197] A batch normalization layer that performs per-channel normalization.

[0198] An activation function ReLU is used to introduce nonlinearity.

[0199] A 3×3 convolutional layer with 512 output channels, a stride of 1, and padding of 1.

[0200] A batch normalization layer that performs per-channel normalization.

[0201] After entering submodule 1, the input feature map passes through a 3×3 convolutional layer with a stride of 2, a batch normalization layer, an activation function, a 3×3 convolutional layer with a stride of 1, and a batch normalization layer. After passing through the 3×3 convolutional layer with a stride of 2, the feature map size is halved, and the number of channels is doubled. Due to this change in feature map size, the output feature map cannot be residually connected to the input feature map. Therefore, the input feature map must enter the identity mapping module, where the output is the same size as the output feature map. This is then residually connected to the output feature map as the input to submodule 2. Since submodule 2 does not have a stride of 2 convolutional layer, the feature map remains the same size after passing through each layer. Therefore, the feature map output from submodule 2 is directly residually connected to the input feature map from submodule 2. The final output feature map after Conv5 is 7×7×512 in size.

[0202] Step 7: Perform an average pooling operation on the 512 7×7 feature maps, that is, calculate the average value of all pixels in the 7×7 area of the map. The final output size is 1×1×512, that is, a 512-dimensional feature.

[0203] Step 8: Input the 512-dimensional features into the shared dimensionality reduction layer and output d-dimensional features, where d is less than 512.

[0204] The shared dimensionality reduction layer consists of the following layers:

[0205] A fully connected layer is used for feature dimensionality reduction and information integration. The weight parameters of this layer are shared, projecting different modal features into a common feature space to achieve feature sharing and alignment.

[0206] A batch normalization layer to speed up training and enhance model stability.

[0207] Furthermore, the steps of hybrid feature fusion in S3 are:

[0208] In the first step, the number of features n is concatenated along the direction of number. The shape of the concatenated matrix X is n×d, where n is the number of features and d is the feature dimension.

[0209] The second step is to define three learnable weight matrices, namely the query weight matrix (Queries) W Q ,Key weight matrix (Keys)W K ,Value weight matrix (Values)W V ,in,

[0210] In the third step, X is projected onto these weight matrices to obtain Q = XW Q ,K=XW K and V=XW V .

[0211] The fourth step is to divide Q, K, and V into h matrices respectively, where h is the number of attention heads. Then the dimension of each matrix is n×d k .

[0212] Step 5: For each attention head i, calculate its scaled dot product attention.

[0213]

[0214] Softmax(·) is used to convert a real number vector into a probability distribution, reflecting the relative importance of each input to the output. That is, given z = [z1, z2, ... z n ],i∈n

[0215]

[0216] Step 6: h Attention i After splicing, we get the result of multi-head attention conversion, and its shape is still n×d.

[0217] In the seventh step, the result of the multi-head attention conversion is normalized. The processed result is then residually connected to the input matrix X to serve as the input X1 of the fully connected layer.

[0218] In the eighth step, X1 passes through a fully connected layer to project the embedding dimension d of each feature to a higher dimension 4d.

[0219] In the ninth step, the feature group passes through a fully connected layer to restore the embedding dimension 4d of each feature back to the embedding dimension d.

[0220] In the tenth step, the feature group is layer-normalized and the processed result is residually connected with the input matrix X1 to obtain the final output. The shape of the final output matrix is still n×d.

[0221] In the 11th step, repeat steps 2 to 10. The feature group passes through an encoder layer to enhance the expression of the fused features.

[0222] In the twelfth step, the most discriminative features are selected and classified through the fully connected layer and the softmax layer, and the model is trained using cross entropy loss.

[0223] Figure 1 This is a system block diagram, consisting of the image capture, sensor data capture, and data processing components. The image capture component accurately captures visual data of the tool's worn edge, providing raw input for subsequent wear detection algorithms. The sensor data capture component collaboratively collects physical signals from multiple sensor types, providing multidimensional dynamic data support for tool wear detection. The data processing and analysis component extracts features and fuses and analyzes multimodal data for tool wear status identification.

[0224] Figure 2The figure shows a schematic diagram of the system structure of the present invention. Specifically, it includes an image capture part, a sensor data capture part and a data processing part; the image capture part mainly includes an image acquisition module, an illumination module, a synchronization control module and a mechanical stabilization module. The image acquisition module is used to capture the microscopic deformation data of the worn edge of the tool through high resolution (3.2 microns / pixel) and a field of view of 17.8mm×14.3mm; the illumination module is used to provide all-round uniform illumination, optimize the incident angle to suppress reflective interference, and enhance the feature contrast of the worn area; the synchronization control module is used to receive the spindle position signal of the CNC machine tool, accurately trigger the image capture action, ensure that the timing is strictly synchronized with the milling action, and avoid motion blur; the mechanical stabilization module is used to maintain a rigid connection between the camera box and the workbench, and eliminate image distortion caused by vibration or displacement. The illumination module, mechanical stabilization module and synchronization control module are used to assist the image acquisition module in image capture. The sensor data capture part mainly includes a sensor module, a signal conditioning module, a data conversion module, a processing and transmission module, a synchronization control module and a power management module. The synchronization control module receives the machine tool spindle position signal, ensuring strict synchronization between sensor data acquisition and machining operations, eliminating timing errors. The power management module provides stable power to each module, optimizing energy consumption and preventing hardware damage from overvoltage and overcurrent. The sensor module directly senses physical quantities such as vibration, sound waves, and cutting force through sensitive components, converting them into electrical signals and sending them to the signal conditioning module. The signal conditioning module amplifies, removes noise, and performs impedance matching on the weak electrical signals output by the sensors, improving signal quality and sending them to the data conversion module. The data conversion module converts the conditioned analog signals into digital signals for subsequent processing and storage. The data storage module stores data sampled at different sampling frequencies locally on the client. This stored data undergoes data cleaning to filter out low-quality data for data fusion analysis. The processing and transmission module performs data encapsulation and format conversion, and transmits the data to a host computer or cloud via wired or wireless means. The data processing section includes a feature extraction module and a feature fusion module. The feature extraction module feeds the converted Mel spectra and tool operation images into a modified ResNet18 for feature extraction. The module then sends these features to the feature fusion module, where the extracted features are fused and encoded according to their importance, maintaining the same number and dimension of the encoded features. Finally, the most discriminative features are selected to classify the tool wear state.

[0225] Figure 3 Shown Figure 2The schematic diagram simulates an industrial scenario, primarily consisting of a CNC machine tool and a camera box. The CNC machine tool performs milling operations and provides a platform for data capture. The camera box contains image capture and controller components, synchronizing image and sensor data capture. The worktable secures and supports the workpiece, ensuring stability during machining and measurement. The tool is mounted on the spindle and used to machine the workpiece. The tool insert is attached to the tool and directly participates in the cutting process. The spindle supports the tool and rotates to drive the tool during machining. Acoustic emission sensors detect acoustic emission signals generated during machining to monitor for internal damage or cracks. Accelerometers measure vibration and acceleration of the workpiece or tool during machining to analyze machining conditions and structural response. A force platform measures and records forces applied to the workpiece during machining to analyze machining forces and mechanical properties. When the tool is in the capture position, i.e., ready for inspection or measurement, it is synchronized with image capture or sensor data acquisition. The camera in the camera box is used to capture images of the workpiece and tool, the macro lens is used to magnify details so that the camera can capture high-resolution images, the LED light is used to provide lighting to ensure that the camera can obtain clear images, and the transparent protective cover is used to protect the internal components while allowing the camera to capture images.

[0226] Figure 4 Shown is a schematic diagram of the method framework of the present invention. Figure 4 (a) shows the conversion of sensor time series data into a Mel spectrogram. First, the power spectrum is calculated: The first step is to split the signal into short time series of length N along the time axis using a fixed sliding window method, and set overlapping windows for adjacent series, which corresponds to the framing operation. Then, STFT is applied to each short time series to obtain the frequency spectrum representation of the series at different frequencies and times. The calculation formula is as follows:

[0227]

[0228] Where x[n] is the discrete time series signal to be transformed, ω represents the normalized angular frequency, w[n-τ] is the window function, and τ is the center of the window function, indicating the time position of the current analysis. This corresponds to a windowing operation. The window function uses a zero-centered Hann window w[n], and the calculation formula for w[n] is as follows.

[0229]

[0230] Map the Hz frequency in the extracted frequency domain information to the Mel scale. The calculation formula is as follows:

[0231]

[0232] Where f Mel and f Hzrepresent the frequency in Mel scale and the frequency in Hz scale respectively.

[0233] After converting the Hz frequency to the Mel scale, m triangular filters are evenly spaced on the Mel scale and then converted to Hertz. The calculation formula is as follows:

[0234]

[0235] Calculate the value of the filter, where the value of the mth filter H m (k) is calculated as follows:

[0236]

[0237] Where k is the frequency, f(m) is the center frequency of the triangular filter, and ∑ m H m (k)=1.

[0238] The sixth step is to extract the signal features by convolving the power spectrum of the signal with the pre-designed Mel filter bank to obtain the bandpass filtering result s(m) in the Mel frequency domain. The calculation formula is as follows:

[0239]

[0240] Where M is the total number of mel filters. Each point on the mel spectrogram corresponds to the bandpass filtering result of a short time series at that moment in time within that frequency band. Its pixel value represents the signal energy intensity at that specific time and frequency. Since a short time series represents a period of time, the position of a moment is taken as the center point of the short time series. The mel spectrograms at different moments are then spliced together to obtain the final mel spectrogram.

[0241] Figure 4 (b) The converted Mel spectrogram and tool operation image are input into the improved ResNet18 for feature extraction. Each image data is extracted into a 512-dimensional abstract feature representation and a shared dimensionality reduction layer is established to integrate and reduce the common information of the features.

[0242] Figure 4 (c) shows the establishment of a two-layer Transformer encoder, which fuses and encodes the extracted features according to their importance. After encoding, the number and dimension of each feature remain unchanged. Finally, the most discriminative features are selected to classify the tool wear state. Specifically, the number of features n is spliced along the direction of number, and the shape of the spliced matrix X is n×d, where n is the number of features and d is the feature dimension. Three learnable weight matrices are defined, namely the query weight matrix (Queries) W Q ,Key weight matrix (Keys)W K,Value weight matrix (Values)W V ,in, X is projected onto these weight matrices to obtain Q = XW Q ,K=XW K and V=XW V . Divide Q, K, and V into h matrices respectively, where h is the number of attention heads, and let Then the dimension of each matrix is n×d k For each attention head i, calculate its scaled dot product attention.

[0243]

[0244] Softmax(·) is used to convert a real number vector into a probability distribution, reflecting the relative importance of each input to the output. That is, given z = [z1, z2, ... z n ],i∈n

[0245]

[0246] h Attention i The multi-head attention transformation results are concatenated to obtain the result, which still has the shape of n×d. Then, the multi-head attention transformation results are layer-normalized. The processed results are residually connected to the input matrix X and used as the input X1 of the fully connected layer. X1 passes through a fully connected layer to project the embedding dimension d of each feature to a higher dimension 4d. The feature group passes through a fully connected layer to restore the embedding dimension 4d of each feature back to the embedding dimension d. The feature group is layer-normalized and the processed results are residually connected to the input matrix X1 to obtain the final output. The shape of the final output matrix is still n×d. Repeat steps 2 to 10. The feature group passes through an encoder layer to enhance the expression of the fused features. Finally, the most discriminative features are selected for classification through the fully connected layer and softmax layer. The model is trained using cross-entropy loss.

[0247] Figure 5 The data fusion analysis diagram of the present invention is shown below. The specific process is as follows:

[0248] The data flow path of the identity mapping module is divided into two parts. The first part undergoes a convolution operation with a kernel size of 3×3 and a stride of 2, which halves the size of its feature map and doubles the number of channels. The next part undergoes a convolution operation with a kernel size of 3×3 and a stride of 1, which does not change the size of the feature map or the number of channels. Because residual connections are required with the feature map before entering the module, the same data needs to pass through the second part to ensure that its feature map size and channels are the same. The second part is a convolution operation with a kernel size of 1×1 and a stride of 2. After the data passes through this operation, the feature map size and number of channels are halved and the number of channels is doubled. After the data passes through the second part, the output feature map size and number of channels are the same as the data that passed through the first part, so residual connections can be performed. If there is no convolution operation with a stride of 2, the feature map size and number of channels do not change, and residual connections can be performed directly.

[0249] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "preferred embodiments," "specific implementations," or "preferred implementations" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0250] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. Therefore, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. Tool wear degree recognition method based on Mel-ResNet-Transformer, characterized by: Including steps: S1. Convert sensor time series data into Mel spectrogram to achieve modal unification of one-dimensional signal and image; S2, parallel feature extraction, input the converted Mel spectrum and tool operation image into the feature extraction model for feature extraction, and generate a 512-dimensional abstract feature representation; S3: Hybrid feature fusion, establish a two-layer Transformer encoder, fuse and encode the extracted features according to their importance, and select the most discriminative features for classification.

2. The method according to claim 1, wherein: The step S1 comprises: S1.

1. Obtaining time series signals x[n] collected by a vibration sensor, an acoustic wave sensor, and a cutting force sensor, wherein the time series signals represent dynamic change data of vibration, acoustic wave, and cutting force during tool machining; S1.2, perform signal framing processing; The time series signal x[n] is divided into short time series of length N along the time axis according to the fixed sliding window method, and overlapping windows of adjacent series are set; S1.

3. Apply the Hann window function w[n] to each short time series. The calculation formula is as follows: S1.

4. Perform time-frequency analysis; Perform short-time Fourier transform on the windowed signal to obtain the time-frequency spectrum representation. The calculation formula is as follows: Where ω represents the normalized angular frequency, w[n-τ] is the window function, and τ is the center of the window function, which represents the time position of the current analysis. S1.5, perform frequency scale conversion; The linear Hz frequency f in the extracted frequency domain information Hz Mapping to the nonlinear Mel frequency scale f Mel , where f Hz The calculation formula is: f Mel The calculation formula is as follows: S1.6, constructing a Mel filter bank; Distribute M triangular filters equally spaced on the Mel frequency scale and calculate the value H of the mth filter m (k); the calculation formula is as follows: Where k is the frequency, f(m) is the center frequency of the triangular filter, and ∑ m H m (k) = 1; S1.7, perform Mel feature extraction; Calculate the convolution of the signal power spectrum and the Mel filter bank to obtain the Mel frequency domain feature s(m). The calculation formula is: The final output is a Mel spectrogram containing time-frequency features.

3. The method according to claim 1, wherein: The step S2 comprises: S2.

1. Data preprocessing: scaling the Mel spectrum and tool operation image to a uniform size of 224×224 pixels and performing normalization processing; S2.2, perform feature extraction: The ResNet18 model of IMAGENET1K_V1, pre-trained on the ImageNet dataset, is used to extract features from the pre-processed Mel spectrogram and tool operation image. The feature extraction model includes: The initial convolution layer uses a 7×7 convolution kernel with a stride of 2 and outputs a 112×112×64 dimensional feature map. The maximum pooling layer uses a 3×3 pooling window with a step size of 2 and outputs a 56×56×64-dimensional feature map; There are four residual block groups, namely Conv2_x, Conv3_x, Conv4_x and Conv5_x. Each residual block group contains at least two residual units, where: Conv2_x outputs a 56×56×64 dimensional feature map; Conv3_x outputs a 28×28×128-dimensional feature map; Conv4_x outputs a 14×14×256 dimensional feature map; Conv5_x outputs a 7×7×512-dimensional feature map; The global average pooling layer reduces the 7×7×512-dimensional feature map to a 512-dimensional feature vector; S2.

3. Input the 512-dimensional feature vector into a shared dimensionality reduction layer, reduce the dimensionality to a d-dimensional feature vector through a fully connected layer, where d < 512, and process it through batch normalization and ReLU activation function; S2.

4. Output the d-dimensional feature vector as input of S3.

4. The method according to claim 1, wherein: The step S3 comprises: S3.

1. Concatenate n d-dimensional feature vectors along the direction of the number of features. The shape of the concatenated feature matrix X is n×d, where n is the number of features and d is the feature dimension. S3.

2. Model the feature interaction of the feature matrix X through the multi-head attention mechanism, specifically including: a) Define three learnable weight matrices: query weight matrix W Q , key weight matrix W K , value weight matrix W V ,in, b) Project the feature matrix X onto the weight matrix to obtain the query matrix Q = XW Q , key matrix K = XW K Sum matrix V = XW V ; c) Divide the query matrix, key matrix, and value matrix into h attention heads respectively, Then the dimension of each attention head is n×d k ; d) Calculate the scaled dot product attention for each attention head i: Among them, Softmax(·) is used to convert a real number vector into a probability distribution, reflecting the relative importance of each input to the output; that is, given z = [z1, z2, ... z n ],i∈n S3.

3. Perform feature fusion, specifically including: a) Splice the output features of each attention head; specifically: i Splicing is performed to obtain the result of multi-head attention transformation, whose shape is still n×d; b) Perform layer normalization on the concatenated features; c) Perform residual connection on the normalized features and the input feature matrix X to obtain the intermediate features X1; S3.4 performs feedforward enhancement, specifically including: a) Project the intermediate feature X1 into 4D space through the fully connected layer; b) Restore the 4D-dimensional features to dimension d through the fully connected layer; c) Perform layer normalization on the restored features and perform residual connection with the intermediate feature X1 to obtain the final output. The shape of the final output matrix is still n×d; S3.5 performs classification output.

5. The method according to claim 4, characterized in that: S3.5 specifically includes: a) Input the feature-fused n×d-dimensional feature matrix into the second-layer Transformer encoder and repeat steps S3.2 to S3.4 to further enhance the feature expression capability; b) Select the most discriminative eigenvector from the final output feature matrix; c) passing the selected feature vectors through the fully connected layer and the softmax classification layer in sequence, and outputting the classification results of the tool wear degree; Among them, the structure of the second-layer Transformer encoder is the same as that of the first layer, which is used to deepen the feature fusion effect.

6. Tool wear degree identification device based on Mel-ResNet-Transformer, characterized by: include: An image capture unit, used for collecting visual data of the tool wear edge and outputting the tool operation image to the data processing unit; The sensor data capture unit is used to collect vibration, sound wave and cutting force signals during tool processing and output the conditioned time series data to the data processing unit; The data processing unit receives the tool operation image and sensor time series data, and realizes the tool wear status identification.

7. The device according to claim 6, characterized in that: The image capture unit includes an image acquisition module, an illumination module, a synchronization control module and a mechanical stabilization module; The image acquisition module uses a camera with a resolution of 3.2 microns per pixel and a field of view of 17.8 mm × 14.3 mm to capture microscopic deformation data of the tool wear edge; The lighting module is used to ensure that the camera can obtain a clear image; The synchronization control module is used to receive the spindle position signal of the CNC machine tool and trigger the image capture action to ensure that the image capture is synchronized with the milling action to avoid motion blur caused by tool rotation; The mechanical stabilization module is used to maintain the rigid connection between the camera box and the tool cutting workbench, eliminating image distortion caused by vibration or displacement.

8. The device according to claim 6, characterized in that: The sensor data capture unit includes a sensor module, a signal conditioning module, a data conversion module, a data storage module, a processing and transmission module, a synchronization control module, and a power management module; wherein, The sensor data capture unit includes a sensor module, a signal conditioning module, a data conversion module, a data storage module, a processing and transmission module, a synchronization control module, and a power management module; The sensor module directly senses vibration, sound waves and cutting force through sensitive elements and converts them into electrical signals; The signal conditioning module is used to amplify, remove noise and perform impedance matching on the weak electrical signal output by the sensor module to improve the signal quality; The data conversion module is used to convert the conditioned analog signal into a digital signal for subsequent processing and storage; The data storage module is used to store the data obtained by sampling at different sampling frequencies locally on the client, and the stored data needs to be cleaned to filter out low-quality data for data fusion analysis; The processing and transmission module is used to perform data encapsulation and format conversion, and send the data to the host computer or the cloud via wired / wireless means; The synchronization control module is used to receive the machine tool spindle position signal to ensure that the sensor data acquisition is synchronized with the processing action and eliminate timing errors; The power management module is used to provide stable power supply for each module, optimize energy consumption and prevent overvoltage / overcurrent from damaging hardware.

9. The device according to claim 6, characterized in that: The data processing unit includes a signal image conversion module, a feature extraction module, and a feature fusion module; The signal image conversion module is used to convert the sensor time series data into Mel spectrogram to achieve modal unification of one-dimensional signal and image; The feature extraction module is used to input the converted Mel spectrum and tool operation image into the improved ResNet18 model for feature extraction; The feature fusion module is used to fuse and encode the extracted features according to their importance, and the number and dimension of each feature remain unchanged after encoding; finally, the tool wear status is classified based on the fusion result.

Citation Information

Cited By

  • Intelligent cutter wear monitoring method and related equipment

    CN121340033A