Ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism

By introducing multi-scale convolution and attention mechanisms in ball screw diagnosis, combined with S transform and convolution kernels of different sizes, the problem of fault type identification under complex operating conditions is solved, and higher diagnostic accuracy and stability are achieved.

CN120125840APending Publication Date: 2025-06-10LANZHOU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510208302.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art ball screw fault types are difficult to identify under complex working conditions and have low classification accuracy.

Method used

The ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism is adopted, and the one-dimensional vibration signal is converted into two-dimensional images through S transformation, and the convolution kernel and attention mechanism module of different sizes are combined to extract and fuse feature information.

Benefits of technology

It improves the accuracy and stability of fault diagnosis, enhances attention to important features, reduces the risk of overfitting, and significantly improves the diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125840A_ABST
    Figure CN120125840A_ABST
Patent Text Reader

Abstract

The invention discloses a ball screw intelligent diagnosis method based on multi-scale convolution and an attention mechanism, and belongs to the technical field of ball screw fault type identification. Comprising the following steps: collecting vibration signals of different fault types, converting a one-dimensional vibration signal into a two-dimensional image by using S transformation, setting a classification label, and dividing data into a training set and a test set; performing feature extraction on the two-dimensional image by adopting multi-scale convolution layers of convolution kernels with different sizes; initializing the structure of the network model, and setting hyper-parameters of the network model; importing the training set into a network model for training; inputting the test set into the network model, verifying the fault diagnosis effect and performance of the network model, and carrying out the fault recognition and classification of the to-be-tested ball screw of the numerical control machine tool based on the network model. Compared with a traditional method, the method is higher in diagnosis accuracy and better in stability, and intelligent diagnosis of the ball screw of the numerical control machine tool can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ball screw fault type identification, and more specifically, to an intelligent diagnosis method for ball screws based on multi-scale convolution and attention mechanism. Background Art

[0002] Currently, common time-frequency analysis methods include short-time Fourier transform, wavelet transform, and S transform, etc. Since the short-time Fourier transform uses a fixed short-time window function, it is essentially a single-resolution signal analysis method and is difficult to maintain good resolution in both the time domain and frequency domain of non-stationary signals. There are certain difficulties in selecting wavelet bases for wavelet transform, and at the same time, data redundancy is relatively serious, and the analysis results of different wavelet bases are also different. S transform is a new time-frequency analysis tool with adjustable time-frequency resolution, which can meet the time-frequency analysis requirements of different frequency signals. Due to its excellent anti-noise ability, S transform is particularly suitable for the analysis and processing of vibration signals.

[0003] Zheng Lei et al. used wavelet transform to extract fault information, and then used the feature vector after regularization and dimensionality reduction as the input of the neural network. The test results show that it can effectively improve the fault diagnosis speed and accuracy of different circuits. Gao Wei et al. directly input the original one-dimensional time series into a one-dimensional convolutional neural network (CNN). Although using a global pooling layer instead of a fully connected layer to reduce model parameters can effectively extract fault features, there is a problem of redundant feature quantity. Yu Yueqiang et al. combined continuous wavelet transform and CNN, used continuous wavelet transform to replace Fourier transform to process the transient signal of impulse frequency response analysis, and based on a one-dimensional neural network to diagnose the fault of the signal processed by CWT. The noise comparison test shows that the proposed method can improve the robustness and generalization ability of the model. Wang Huidong et al. used a multi-scale convolution module to establish a fault diagnosis model and performed data augmentation based on the generative adversarial network to solve the problem of unbalanced fault samples. Zheng Yizhen et al. completed the fault diagnosis of the cage of cylindrical roller bearings based on an adaptive convolutional neural network fault diagnosis model for "end-to-end" identification, solving the problems of instability of the fault signal of the rolling bearing cage and lack of impact characteristics. Since the receptive field of 1D-CNN is small and the network depth is insufficient, and overfitting is likely to occur during the training process, resulting in low diagnostic accuracy. Therefore, in contrast, the effect of 1D-CNN in processing time-frequency signals is inferior to that of 2D-CNN.

[0004] Zhang Jun et al. integrated the attention mechanism with a two-dimensional convolutional neural network. First, the original signal was processed by short-time Fourier transform. Then, the obtained two-dimensional image was input into the two-dimensional convolutional neural network to extract features and classify them. The result does not depend on the feature vector, meeting the requirements of lightweight and easy training. Wang Qingrong et al. used the S transform to convert the original data into a time-frequency diagram, then performed secondary feature extraction with a CNN, and finally used a classifier to classify the faults. Gong Jun et al. used the synchrosqueezed wavelet transform (SWT) to convert the original one-dimensional signal into a two-dimensional time-frequency image containing high-frequency information, improving the network's representation ability by increasing the attention to important features, and combined with a two-dimensional convolutional neural network for fault diagnosis. Although the two-dimensional convolutional neural network has strong image processing capabilities and plays a certain advantage in extracting fault features and improving the diagnostic accuracy, it may not pay more attention to some important features.

[0005] Although the above research shows good performance in fault recognition under single working conditions, for ball screws in complex working conditions, there are still problems such as difficult fault type recognition and low classification accuracy.

[0006] Therefore, how to provide an intelligent diagnosis method for ball screws based on multi-scale convolution and attention mechanism to improve the accuracy and stability of diagnosis is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides an intelligent diagnosis method for ball screws based on multi-scale convolution and attention mechanism. Aiming at the problems existing in the above-mentioned prior art, a two-dimensional convolutional neural network model combined with an attention mechanism is provided to improve the accuracy and stability of fault diagnosis.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] An intelligent diagnosis method for ball screws based on multi-scale convolution and attention mechanism, comprising:

[0010] S100: Collect vibration signals of different fault types, use the S transform to convert the one-dimensional vibration signal into a two-dimensional image, set classification labels at the same time, and divide the data into a training set and a test set;

[0011] S200: Use a multi-scale convolutional layer with different sizes of convolutional kernels to extract features from the two-dimensional image;

[0012] S300: Initialize the structure of the network model and set the hyperparameters of the network model;

[0013] S400: Import the training set into the network model for training;

[0014] S500: Input the test set into the network model to verify the fault diagnosis effect and performance of the network model;

[0015] S600: Based on the network model, perform fault identification and classification on the ball screw of the numerically controlled machine tool to be measured.

[0016] Furthermore, the expression for converting the one-dimensional vibration signal into a two-dimensional image using the S transform is:

[0017]

[0018] In the formula: τ is time, which controls the position of the window function on the time axis; h(t) is the analysis signal; f is the frequency; S(τ, f) is the time-frequency spectrum matrix obtained by the transformation.

[0019] Furthermore, different sizes of convolution kernels include:

[0020] 1*1 convolution kernel, 3*3 convolution kernel, 5*5 convolution kernel.

[0021] Furthermore, it also includes: fusing the outputs of the multi-scale convolution layers with different sizes of convolution kernels.

[0022] Furthermore, input the fused output image into the attention mechanism module.

[0023] Furthermore, the attention mechanism module includes: a channel attention module and a spatial attention module, and the expression is:

[0024]

[0025] Among them, F ∈ R C×W×H , M C ∈ R C×1×1 , M S ∈ R 1×H×W ;

[0026] In the formula, F is the fused output image, is the element-wise multiplication, C, W, and H respectively represent the number of channels, width, and height; F′ is the image after passing through the channel attention module, F″ is the image after passing through the spatial attention module, M S is the one-dimensional channel attention map, M C is the two-dimensional spatial attention map.

[0027] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an intelligent diagnosis method for ball screws based on multi-scale convolution and attention mechanism. The one-dimensional vibration signals collected are transformed into two-dimensional images through S transform as the input. The time-frequency diagram can present more abundant fault information, and the traditional convolution layer is improved, and a multi-scale convolution layer is designed, which can extract more subtle important features horizontally. The attention mechanism and two-dimensional convolutional neural network are introduced into the fault identification of ball screws in numerical control machine tools. The attention mechanism can enhance the attention to important features, thereby improving the accuracy and stability of fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0029] Figure 1 It is a schematic flow chart of the method of the present invention;

[0030] Figure 2 It is a schematic structural diagram of a fault diagnosis model based on multi-scale convolution and attention mechanism provided by an embodiment of the present invention;

[0031] Figure 3 It is a schematic diagram of an improved convolutional neural network (multi-scale convolutional neural network) provided by an embodiment of the present invention;

[0032] Figure 4 It is a schematic structural diagram of a convolutional neural network provided by an embodiment of the present invention;

[0033] Figure 5 It is a schematic structural diagram of an attention mechanism module provided by an embodiment of the present invention;

[0034] Figure 6 It is a time-domain waveform diagram of a fault signal provided by an embodiment of the present invention;

[0035] Figure 7 It is a schematic diagram of the time-frequency transformation results of six time-domain signals provided by an embodiment of the present invention;

[0036] Figure 8 It is a curve diagram of the change of the training accuracy of the model provided by an embodiment of the present invention;

[0037] Figure 9 It is a curve diagram of the change of loss provided by an embodiment of the present invention;

[0038] Figure 10Schematic diagram of the confusion matrix provided by the embodiments of the present invention;

[0039] Figure 11 Comparison chart of the accuracy rates of the model provided by the present invention and other different models. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] See Figure 1 , the embodiments of the present invention disclose an intelligent diagnosis method for ball screws based on multi-scale convolution and attention mechanism, including:

[0042] S100: Collect vibration signals of different fault types, use the S transform to convert the one-dimensional vibration signal into a two-dimensional image, set classification labels at the same time, and divide the data into a training set and a test set;

[0043] S200: Use multi-scale convolution layers with different sizes of convolution kernels to extract features from the two-dimensional image;

[0044] S300: Initialize the structure of the network model and set the hyperparameters of the network model;

[0045] S400: Import the training set into the network model for training;

[0046] S500: Input the test set into the network model to verify the fault diagnosis effect and performance of the network model;

[0047] S600: Based on the network model, perform fault identification and classification on the ball screws of the numerically controlled machine tools to be measured.

[0048] In a specific embodiment, the fault diagnosis model structure based on multi-scale convolution and attention mechanism is as Figure 2 shown. By using multi-scale convolution kernels to splice and fuse the features of the data at different time scales, deep features can be extracted. The attention mechanism can capture the important information in the horizontal features, enhance the influence factors of this part of the features, so as to improve the accuracy of the model and reduce the risk of overfitting.

[0049] The detailed steps are as follows:

[0050] 1) In the data preprocessing stage, use the S transform to convert the one-dimensional vibration signal into a two-dimensional image, set classification labels at the same time, and divide the data into a training set and a test set.

[0051] 2) Design multi-scale convolutional layers using convolutional kernels of different sizes 1*1, 3*3, and 5*5 to extract features from the image processed by the S transform.

[0052] 3) Initialize the network structure and set the hyperparameters of the network.

[0053] 4) Import the training set into the network model for training.

[0054] 5) Input the test set into the model to verify the fault diagnosis effect and performance of the model.

[0055] In a specific embodiment, the S transform is a reversible time-frequency analysis technique that combines the characteristics of the short-time Fourier transform and the wavelet transform. It solves the defect that the short-time Fourier transform cannot adjust the analysis window frequency, introduces the multi-resolution analysis of the wavelet transform, and at the same time maintains a direct connection with the Fourier spectrum. Its definition is:

[0056]

[0057] In the formula: τ is time, which controls the position of the window function on the time axis; h(t) is the analysis signal; f is the frequency; S(τ, f) is the obtained time-frequency spectrum matrix.

[0058] In a specific embodiment, CNN is a neural network specifically for image data processing, with the ability of feature learning, and can use its hierarchical structure to perform translation-invariant classification on the input information.

[0059] In a specific embodiment, local features of the input data are extracted through convolution operations. Each convolutional kernel in the convolutional layer can extract a specific feature, and multiple convolutional kernels can work in parallel to extract different types of features. Its mathematical model can be expressed as:

[0060]

[0061] In the formula, X is the input of the convolutional layer, M j is the set of output feature maps of the l-1 layer, ω is the weight matrix of the corresponding convolutional kernel, is the bias term, l is the convolutional layer number, i and j are two connected neurons, f is the activation function, which can improve the non-linear expression ability of the network. The commonly used activation function in CNN is RELU, which can be expressed as:

[0062] f(x) = max{0, log[1 + exp(x)]} (3)

[0063] In a specific embodiment, the pooling layer, also known as the downsampling layer, mainly serves to downsample (or reduce the dimension) the output of the convolutional layer, with the aim of reducing the number of parameters and improving the computational efficiency. Common pooling methods include max pooling and average pooling, which are specifically represented as follows:

[0064]

[0065] In the formula, y is the output of the pooling layer, X down is the downsampling function, x is the input, and is the bias term.

[0066] In a specific embodiment, before the output layer of the network is the fully connected layer. The fully connected layer is actually a fully connected neural network that integrates the features extracted by the previous layers for tasks such as classification or regression. Each neuron in the fully connected layer is connected to all neurons in the previous layer. Each neuron in the network is interconnected with other neurons at different levels, which makes the number of parameters in the entire network reach the maximum. Its mathematical model can be expressed as:

[0067] y = f(ω·x + b) (5)

[0068] In the formula, x is the input matrix, y is the output matrix, f is the activation function, and b is the bias of the fully connected layer.

[0069] See Figure 3 , in a specific embodiment, an improved version of the convolutional neural network (CNN), called the multi-scale convolutional neural network, is disclosed, which is specifically used for fault diagnosis of the ball screw of a numerically controlled machine tool. This method uses three different sizes of convolutional kernels to extract features from the image processed by the short-time Fourier transform. Different from the traditional longitudinal deepening method (i.e., convolution, pooling, and then convolution), this method can extract subtle and important features more comprehensively in the horizontal direction. Through multi-layer convolution, the network can gradually learn and extract more abstract and semantically rich features from the original image, thereby achieving more accurate and effective feature extraction and processing of the input data.

[0070] In a specific embodiment, it further includes: fusing the outputs of the multi-scale convolutional layers with different sizes of convolutional kernels.

[0071] In a specific embodiment, the working environment load of the ball screw of a numerically controlled machine tool fluctuates greatly, resulting in the time-varying and non-linear characteristics of its vibration signal. Therefore, under the same conditions, the signal features obtained at different times are different. Some features can effectively reflect the fault information, while others may bring interference, thereby affecting the generalization ability of the model. The attention mechanism adaptively assigns weights to the features of different signal segments for information screening, highlighting important fault features and suppressing invalid features at the same time. Its structure is asFigure 4 as shown

[0072] In a specific embodiment, the fused output image is input into the attention mechanism module.

[0073] Specifically, referring to Figure 5 , the attention mechanism was initially used in machine translation and usually adopted an auto-encoding method for sequence conversion. This mechanism originated from the research on human vision and has since been widely applied in fields such as natural language processing. The attention mechanism module CBAM consists of two sub-modules: the channel attention module and the spatial attention module, which focus on channels and space respectively. The input image F (height × width × channels) is processed by the max-pooling layer and the global average pooling layer to generate a feature map with a height and width of 1. Then, these feature maps are fed into two shared perceptron networks and output by adding them one by one. After passing through the activation function, the channel attention feature is generated. The expression is:

[0074]

[0075] where F ∈ R C×W×H , M C ∈ R C×1×1 , M S ∈ R 1×H×W ;

[0076] In the formula, F is the fused output image, is the element-wise multiplication, C, W, and H represent the number of channels, width, and height respectively; F' is the image after passing through the channel attention module, F'' is the image after passing through the spatial attention module, M S is the one-dimensional channel attention map, and M C is the two-dimensional spatial attention map.

[0077] In a specific embodiment, it also includes experimental analysis based on a ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism disclosed in the present invention.

[0078] Specifically, the experimental data includes:

[0079] Taking a numerically controlled machine tool as an example for fault diagnosis experiments, the GD4010 lead screw used in the experiments has specific parameters as shown in Table 1.

[0080] Table 1 Process parameters of GD4010 ball screw

[0081] Name Lead Screw Diameter Ball Diameter Contact Angle Helix Angle Unit <![CDATA[d 0 / mm]]> <![CDATA[d b / mm]]> α / ° λ / ° Value 40 5.953 45 4.55

[0082] Simulate normal state, ball screw raceway wear fault, rolling element wear fault, ball screw misalignment fault, ball screw bending fault, and ball screw pitting fault respectively. Use YMC121A100 type unidirectional IEPE acceleration sensor, YMC9216 type signal collector, and YMC9800 signal analysis software. At a rotational speed of 1772 r / min, collect vibration signals of different fault types. The sample length is selected as 1024, and the sampling interval is 128. For the specific division method of the training set and the test set, see Table 2.

[0083] Table 2 Fault types and labels

[0084]

[0085] Collect one-dimensional vibration signals. The time-domain waveforms of each fault signal are as Figure 6 shown. As can be seen from Figure 6 , the time-domain waveform can only reflect the fault characteristics in the time domain, and the features extracted from it cannot comprehensively express the fault characteristics of the ball screw. Therefore, the S transform is used to process the time-domain signal, and the obtained time-frequency diagram is as Figure 7 shown. It can be found that the two-dimensional feature matrix formed by the vibration signal processed by the S transform contains richer fault information, which provides sufficient basis for subsequent fault classification.

[0086] Specifically, the experiment is carried out on a computer with a Windows system, a processor of Inter(R) Xeon(R) silver 4210R CPU @ 2.40 GHz, 2394 MHz, and a memory of 64 GB. The specific network parameters are shown in Table 3.

[0087] Table 3 Model parameters

[0088]

[0089] Specifically, the performance verification includes:

[0090] For hyperparameters, the model cannot be adjusted through training and is generally set before model training. The optimization and adjustment of hyperparameters are an important part of fault diagnosis research. The number of input samples Batchsize is taken as 32, the learning rate is taken as 0.01, the commonly used Adam optimizer is selected, and the training samples and validation samples are input into the model for parameter initialization training. As the training progresses, the model performance gradually improves. The accuracy rate of model training is as Figure 8 shown.

[0091] Specifically, from Figure 9As can be seen, the network reached an accuracy of 90% on the training set at around 75 epochs, and the accuracy increased to 99% after 125 epochs. At the same time, as the number of iterations increased, the loss value of the network continued to decline, indicating that the network did not overfit. To show the recognition of fault diagnosis, a confusion matrix was used to visualize the model results. The recognition results of the network for each fault sample were shown in the form of a confusion matrix, as Figure 10 shown. The horizontal axis of the confusion matrix represents the predicted labels of the lead screw, and the vertical axis of the confusion matrix is the true labels of the lead screw. From Figure 10 it can be seen that except for the misjudgment of the recognition of lead screw pitting as other states, the recognition rates of the other 5 states are very high, the fault recognition rate reaches 100%, and the overall accuracy is 99.44%. Thus, it can be seen that using multi-feature extraction and fusing convolutional neural networks for fault recognition is very effective.

[0092] Specifically, to verify the effectiveness of the model of the present invention, the model proposed in the present invention, the S-transform - CNN model, the S-transform - CNN - SVM model, and the S-transform - CNN - BiGRU were compared. And the original one-dimensional data was used, which was transformed into a time-frequency diagram through the S-transform, and the number of iterations for all data was set to 1000.

[0093] From Figure 11 it can be known that the accuracy of network fault recognition based on multi-scale feature extraction and spatial attention mechanism is improved by 3.33% compared with the traditional convolutional neural network. Compared with the models CNN - SVM and CNN - BiGRU, the recognition accuracy of the model of the present invention is significantly higher, reflecting that the model of the present invention can effectively extract fault features and verifying that the model of the present invention has good fault diagnosis ability.

[0094] Specifically, the robustness analysis is as follows:

[0095] Since the machine tool may be interfered by noise when working under actual working conditions, it is necessary to consider the influence of noise on the diagnosis when diagnosing the faults of the ball screw of the numerically controlled machine tool. Noises with different signal-to-noise ratios were added to the original signal, and the diagnostic results are shown in Table 4. It can be seen from the table that when noises with a signal-to-noise ratio of 40 - 60 dB are added to the original signal, although the accuracy decreases slightly, it basically remains above 95%, and the model of the present invention still has a high diagnostic accuracy. Thus, it can be seen that the designed model structure has strong robustness. This is because the S-transform can transform the vibration signal in the time domain into the time-frequency domain, which enables flexible selection of noise suppression regions in different frequency ranges, thereby effectively reducing the influence of noise, and at the same time can well retain the information of the effective signal, and the convolutional layer and pooling layer have a filtering effect.

[0096] Table 4 Diagnostic Results under Noise Interference

[0097]

[0098]

[0099] Specifically, the present invention uses the S transform to convert one-dimensional vibration signals into two-dimensional time-frequency images, adopts multi-scale feature extraction and attention mechanism to extract fault information, and finally completes the identification and classification of faults through two-dimensional CNN. The beneficial effects are as follows:

[0100] 1) By performing time-frequency analysis on non-linear and non-stationary vibration signals through the S transform, the advantages of the S transform in the comprehensive application of the time domain and frequency domain are fully utilized, providing more comprehensive information for the input of the two-dimensional neural network.

[0101] 2) The multi-scale feature extraction module is disclosed, which can obtain fault information at a longer time scale while achieving a larger receptive field, enhancing the feature extraction ability of the model. The attention mechanism can pay more attention to the important features contained in the fault information, improving the accuracy of fault diagnosis.

[0102] 3) Experimental results show that the method proposed in the present invention is superior to existing traditional machine learning fault diagnosis methods in terms of the accuracy and robustness of fault identification.

[0103] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0104] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism, characterized in that: include: S100: collects vibration signals of different fault types, converts one-dimensional vibration signals into two-dimensional images using S transform, sets classification labels, and divides the data into training sets and test sets; S200: extracting features from the two-dimensional image using a multi-scale convolution layer with convolution kernels of different sizes; S300: Initialize the structure of the network model and set the hyperparameters of the network model; S400: importing the training set into the network model for training; S500: input the test set into the network model to verify the fault diagnosis effect and performance of the network model; S600: Perform fault identification and classification on the CNC machine tool ball screw to be tested based on the network model.

2. According to claim 1, a ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism is characterized in that: The expression for converting a one-dimensional vibration signal into a two-dimensional image using S transform is: Where: τ is time, which controls the position of the window function on the time axis; h(t) is the analysis signal; f is the frequency; S(τ,f) is the time-frequency spectrum matrix obtained by transformation.

3. According to claim 1, a ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism is characterized in that: Convolution kernels of different sizes, including: 1*1 convolution kernel, 3*3 convolution kernel, 5*5 convolution kernel.

4. According to claim 1, a ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism is characterized in that: Also includes: The outputs of multi-scale convolutional layers with convolution kernels of different sizes are fused.

5. According to claim 4, a ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism is characterized in that: The fused output image is input into the attention mechanism module.

6. The ball screw intelligent diagnosis method based on multi-scale convolution and attention mechanism according to claim 5 is characterized in that: The attention mechanism module includes: a channel attention module and a spatial attention module, and the expression is: Among them, F∈R C×W×H , M C ∈R C×1×1 , M S ∈R 1×H×W ; Where, F is the output image after fusion, is the multiplication of the corresponding elements, C, W, H represent the number of channels, width and height respectively; F′ is the image after the channel attention module, F″ is the image after the spatial attention module, M S is a one-dimensional channel attention map, M C is a two-dimensional spatial attention map.