Bearing fault automatic diagnosis method and system based on attention coordinate learnable normal form
By employing an attention coordinate learnable paradigm, this method utilizes the cepstrum and time-Mel spectrum feature maps of bearing vibration signals to optimize network parameters, thus solving the problems of high time cost and inflexible feature extraction in traditional methods and achieving efficient and accurate bearing fault diagnosis.
Patent Information
- Application Number
- CN202511186611.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional bearing fault identification technology suffers from high time costs when learning global information and the inability of convolution kernels to adapt to different shape features, making it difficult to diagnose faults efficiently and accurately.
A method based on the attention coordinate learnable paradigm is adopted to optimize network parameters to achieve automatic fault diagnosis by calculating the cepstrum and time-Mel spectrum feature map of the bearing vibration signal, and combining attention features and target detection loss function.
It improves fault identification efficiency, reduces time costs, enhances the adaptability and flexibility of feature extraction, and significantly improves the accuracy of bearing fault diagnosis.
Smart Images

Figure CN120971026A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of fault diagnosis, in particular to a bearing fault automatic diagnosis method and system based on an attention coordinate learnable paradigm. BACKGROUND
[0002] With the development of deep learning, current popular bearing fault recognition technologies, especially in the field of wind power, more adopt frequency spectrum analysis series or neural network series, especially a transformer-based encoder-decoder to recognize faults. The former is to manually extract frequency domain features of a vibration signal and recognize and extract fault features, and the latter actively learns image features through a neural network convolution module or an attention mechanism module. In the neural network series, a local convolution network can learn local information of an image, and an attention network can learn global information of the image, however, the feature learning time consumption cost of the former is low, and the feature learning time consumption cost of the latter is high. Whether the local convolution module or the attention module has a problem that the visual field range of the two architectures is a continuous range rather than adaptively selected several non-continuous visual field ranges.
[0003] The inventor finds that if global information of a vibration signal is to be learned while reducing the learning time cost, the coordinates of the features can be learned first before learning the global features, and the learning cost of learning the global features is determined by controlling the number of the coordinates of the features to be learned. Meanwhile, the coordinate distribution range learned is based on a bearing vibration time sequence signal rather than manually selected visual field ranges, and therefore can adapt to different bearing vibration signals. Meanwhile, the bearing fault automatic diagnosis method and system based on the attention coordinate learnable paradigm can improve the bearing fault recognition rate. However, the traditional transformer model has the problem of global attention region, and therefore the time cost of the transformer in fault recognition is high, and the convolution kernel can only extract features in a rectangular region at the target position and cannot extract features based on the shape of the features. SUMMARY
[0004] To solve the above-mentioned problems, the application provides a bearing fault automatic diagnosis method and system based on an attention coordinate learnable paradigm.
[0005] In a first aspect, the application provides a bearing fault automatic diagnosis method based on an attention coordinate learnable paradigm, which adopts the following technical scheme:
[0006] A bearing fault automatic diagnosis method based on an attention coordinate learnable paradigm, comprising:
[0007] Obtaining bearing vibration signal data;
[0008] According to the obtained bearing vibration signal data, a cepstrum diagram of the vibration signal is calculated, and a time-Mel spectrum feature diagram is calculated;
[0009] According to the obtained time-Mel spectrum feature diagram, attention coordinates are calculated;
[0010] According to the image feature matrix after the offset coordinates, attention features are calculated;
[0011] According to the attention features, a target detection loss is calculated, network parameters are updated by iteratively reducing the loss function value, and automatic bearing fault diagnosis is realized.
[0012] Further, the bearing vibration signal data is obtained, including setting a sampling frequency f of the vibration signal, dividing the vibration signal into a plurality of frame sequences, setting a frame length as 512 times of a sampling period T s = 512 / f, considering that the number of frames obtained within 10 seconds is N s = 10 / T s and taking the integer part downward, adding a Hamming window to the vibration signal of each frame, and setting a window function
[0013] indicates rounding up; the Hamming weighted vibration signal is subjected to Fourier transform to generate a power spectrum, and the frequency range is 0~f / 2, a total of N f = 257 frequency points, the corresponding power spectrum of each frame is calculated to generate a conventional time spectrum diagram
[0014] Further, the cepstrum diagram of the vibration signal is calculated according to the obtained bearing vibration signal data, including calculating a Mel spectrum-normal spectrum filter response matrix wherein N m = 64 triangular filters are designed, the conversion relationship between the Mel frequency and the normal frequency is f m = 2595*log(10*(1+700f)), the divided Mel frequency f m (z) has a value range of 0~2595*log(10*(1+700f / 2)), z=0,1,…,N m +1, a total of N m +2 frequency points; the divided Mel frequency is converted back to the normal frequency, and the left frequency, the center frequency and the right frequency of the triangular filter in the normal frequency coordinate system are f and f
[0015] Further, the time-Mel spectrum feature diagram is calculated, including calculating the response value of the z-th row and the k-th column of the Mel spectrum-normal spectrum filter response matrix:
[0016] Further, the time-Mel spectrum feature map is calculated, denoted as
[0017] Further, the attention coordinates are calculated according to the obtained time-Mel spectrum feature map, including setting a reference coordinate p0, and using three reference coordinates according to the three times of down-sampling, and the reference coordinates of the image reduced by str=8 times are a 3232 coordinate point set p 01 distributed at equal intervals in the horizontal and vertical coordinates, and the reference coordinates of the image reduced by str=16 times are a 1616 coordinate point set p 02 , and the reference coordinates of the image reduced by str=32 times are a 44 coordinate point set p 03 , and a convolution module is designed to input C feature maps down-sampled to generate an offset coordinate feature map Δp n =S 2N,W,H , and the size of the convolution kernel is w=3.
[0018] Further, the attention coordinates are calculated according to the obtained time-Mel spectrum feature map, including calculating a new image feature matrix x H*W,C after the coordinate offset, wherein the new image feature matrix after the coordinate offset is calculated, and p0+p n +Δp n is the offset plane coordinate of the image feature matrix x C,W,H , and the image feature after the bilinear interpolation is x C (p0+p n +Δp n )=(1―u)(1―v)x C (p 0x +p nx ,p 0y +p ny )+u(1―v)x C (p 0x +p nx +1,p 0y +p ny )+(1―u)vx C (p 0x +p nx ,p 0y +p ny +1)+uvx C (p 0x +p nx +1,p 0y +p ny +1),u and v are the offset amounts of Δp n in the horizontal and vertical directions; and the matrix expression of x C (p0+p n +Δp n ) is xC,H,W Secondly, the second and third dimensions of x C,H,W are compressed and processed x C,H*W ; finally, the image feature matrix x H*W,C is obtained by transposition.
[0019] Further, the image feature matrix after the offset coordinates is calculated to obtain the attention feature, including calculating the image feature according to the attention coordinates, wherein the image feature matrix x H*W,C is respectively multiplied by the weight matrix and to obtain the feature matrix Q H*W,C induced by x H*W,D , K H*W,D and V H*W,D , which are denoted as Q, K and V; the correlation matrix is calculated and normalized and processed by the softmax activation function to obtain the attention feature x H*W,D = α H*W,H*W *V, the dimension expansion processing of the attention feature of the image is performed to obtain the attention feature map x H,W,D .
[0020] Further, the target detection loss is calculated according to the attention feature, including taking the attention feature map x H,W,D as the target detection feature, wherein the detection grid of the target detection has a total of HW, and 5 rectangular frame anchor points are set for each grid i = 1, 2, 3, 4, 5 are used to represent the edge boxes with different length-width ratios Meanwhile, the left upper point and the length-width of the rectangular edge box are represented as The confidence of the fault feature contained in the rectangular frame is Four categories are set for each detection grid, denoted as j = 1, 2, 3, 4, and the attention feature given by each grid is expressed as i = 1, 2, …, 5, j = 1, 2, …, 4, and the size of the feature dimension is D = 29.
[0021] Further, the target detection loss is calculated according to the attention feature, further including setting the calculation method of the loss function to iteratively reduce the loss function value and update the network parameters in the process, wherein the loss function of the category is a cross-entropy loss function, denoted as: The loss function of the edge box intersection over union is denoted as: The loss function of the edge box regression is denoted as: The loss function of the confidence regression is denoted as:
[0022] In a second aspect, a bearing fault automatic diagnosis system based on an attention coordinate learnable paradigm comprises:
[0023] A data acquisition module is configured to acquire bearing vibration signal data.
[0024] A cepstrum diagram module is configured to calculate a cepstrum diagram of a vibration signal according to the acquired bearing vibration signal data, and calculate a time-Mel spectrum feature map.
[0025] An attention coordinate module is configured to calculate an attention coordinate according to the obtained time-Mel spectrum feature map.
[0026] An attention feature module is configured to calculate an attention feature according to an image feature matrix after the offset coordinate.
[0027] A detection module is configured to calculate a target detection loss according to the attention feature, update network parameters by iteratively reducing the loss function value, and realize bearing fault automatic diagnosis.
[0028] In a third aspect, the present application provides a computer readable storage medium, wherein a plurality of instructions are stored, the instructions being adapted to be loaded and executed by a processor of a terminal device to implement the bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm.
[0029] In a fourth aspect, the present application provides a terminal device comprising a processor and a computer readable storage medium, the processor being configured to implement instructions, and the computer readable storage medium being configured to store a plurality of instructions, the instructions being adapted to be loaded and executed by the processor to implement the bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm.
[0030] In summary, the present application has the following beneficial technical effects:
[0031] 1. Improve fault recognition efficiency and reduce time cost: the traditional Transformer model has the problem of high time cost caused by global attention area, while the present application controls the number of coordinates of features to be learned to determine the cost of learning global features by learning the coordinates of features. The designed attention coordinate learnable mechanism can focus on key feature areas and avoid indiscriminate learning of the global, while ensuring the learning of global information of the vibration signal, greatly reducing the time consumption of feature learning and improving the overall efficiency of bearing fault diagnosis.
[0032] 2. Enhance the adaptability and flexibility of feature extraction: Traditional convolution kernels can only extract features within a rectangular region at the target position, cannot capture features based on feature shape, and the field of view of local convolution networks and attention networks is usually continuous, which is difficult to adapt to the feature distribution of different bearing vibration signals. In this invention, the distribution range of attention coordinates is automatically learned based on bearing vibration time series signals, rather than manually selected, which can adapt to bearing vibration signals under different working conditions. At the same time, through coordinate offset and bilinear interpolation processing, key fault features of different shapes and positions can be flexibly captured, and the extraction ability of complex and variable fault features is improved.
[0033] 3. Improve the accuracy of fault diagnosis: This method can more effectively extract features containing fault information from vibration signals by calculating time-mel frequency spectrum feature maps. In the attention feature calculation process, by generating Q, K, and V matrices and performing correlation analysis and softmax processing, the weight of key fault features can be highlighted. In addition, the designed multi-dimensional loss function (category loss, bounding box intersection over union loss, bounding box regression loss, and confidence regression loss) can comprehensively optimize network parameters, making the model's judgment of fault categories and positioning of fault regions more accurate, thereby significantly improving the accuracy of bearing fault automatic diagnosis. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 is a schematic diagram of a bearing fault automatic diagnosis method based on an attention coordinate learnable paradigm according to Embodiment 1 of the present invention.
[0035] Figure 2 is a schematic diagram of a bearing fault automatic diagnosis method based on an attention coordinate learnable paradigm according to Embodiment 1 of the present invention. DETAILED DESCRIPTION
[0036] The present invention will be further described in detail below with reference to the accompanying drawings.
[0037] Embodiment 1
[0038] Referring to Figure 1 , the bearing fault automatic diagnosis method based on an attention coordinate learnable paradigm according to the present embodiment includes:
[0039] Specifically,
[0040] First, the vibration signal data is obtained,
[0041] The vibration signal is mainly obtained by an acceleration sensor installed on the bearing seat. In industrial sites (such as wind power, production lines, rail transportation, etc.), the sensor is usually fixed to the radial and axial measurement points of the bearing by magnetic attraction or threaded connection, and the vibration waveform is recorded by a data acquisition device at a sampling frequency of ≥25.6 kHz. In practical applications, the collection strategy needs to be adjusted according to the device speed and bearing type: long-time continuous monitoring is used for low-speed heavy-load bearings, and transient impact capture is focused on for high-speed bearings. Anti-interference processing (such as shielded cable, optical fiber transmission) and real-time quality verification (peak symmetry, signal-to-noise ratio, etc.) are required during signal collection to ensure that the obtained vibration data contains effective fault feature information. In typical industrial scenarios, the system will automatically optimize the sampling parameters in combination with the device operating state (start-stop, load change, etc.), and preliminary feature extraction will be realized through edge computing nodes, and finally the processed vibration signals will be transmitted to the diagnosis system for analysis.
[0042] Step S1 calculates the cepstrum of the vibration signal, which includes:
[0043] Step S1.1, set the sampling frequency f of the vibration signal, divide the vibration signal into multiple frame sequences, and set the time length of one frame to 512 times the sampling period T s = 512 / f, considering that the number of frames that can be obtained within 10 seconds is N s = 10 / T s and take the integer part, add a Hamming window to the vibration signal of each frame, and set the window function , which represents the integer part.
[0044] Step S1.2, Fourier transform the Hamming weighted vibration signal to generate a power spectrum, and the frequency range is 0-f / 2, a total of N f = 257 frequency points, calculate the corresponding power spectrum for each frame to generate a conventional time-frequency spectrum
[0045] Step S1.3, calculate the Mel-spectrum-normal spectrum filter response matrix
[0046] (1) Design N m = 64 triangular filters, the triangular filter has a response of 1 at the center frequency and linearly decays to 0 at the two bottom angles, and the conversion relationship between the Mel frequency and the normal frequency is f m = 2595*log(10*(1+700f)), which is abbreviated as
[0047] (2) Divide the Mel frequency f m (z) into the range 0-2595*log(10*(1+700f / 2)), z=0,1,…,N m+1, total N m +2 frequency points.
[0048] (3) Convert the equalized mel-frequency back to normal frequency, then the left frequency, center frequency and right frequency of the triangular filter in the normal frequency coordinate system can be obtained as and
[0049] (4) The response value of the z-th row and k-th column of the mel-spectrum-normal-spectrum filter response matrix is calculated as
[0050] Step S1.4, calculate the time-mel-spectrum feature map
[0051] wherein the time-spectrum map is a matrix with dimensions N s ×N f . Among them, N s is the total number of frames after the vibration signal is framed within 10 seconds, and N f =257 is the frequency range corresponding to the number of frequency points of the power spectrum (0-f / 2).
[0052] Transpose of the mel-spectrum-normal-spectrum filter response matrix is a matrix with dimensions N f ×(N m +2) (the original matrix is (N m +2)×N f , and the rows and columns are interchanged after transposition). Among them, N m =64 is the number of triangular filters, so N m +2=66.
[0053] Time-mel-spectrum feature map is the result of matrix multiplication, with dimensions N s ×(N m +2), i.e. the number of frames×the feature dimension after mel filtering.
[0054] The calculation method of each element xi,j (corresponding to the i-th frame and the j-th mel-filtered feature) of the time-mel-spectrum feature map is as follows:
[0055] wherein: Xi,k: element of the i-th row and k-th column of the time-spectrum map, corresponding to the power spectrum value of the i-th frame vibration signal at the k-th frequency point.
[0056] The element in the k-th row and j-th column of the transposed filter response matrix corresponds to the response value of the j-th Mel-triangle filter at the k-th frequency point (calculated by the formula in claim 4).
[0057] The calculation process involves weighted summation of the power spectrum value of each row with the Mel filter response value at the corresponding frequency point, ultimately converting the original frequency domain features into Mel frequency domain features.
[0058] Step S2, calculating the attention coordinates, includes:
[0059] Step S2.1: Set reference coordinates p0. Based on the feature map after three downsamplings, three reference coordinates are used: image scaled down by str = 8, and reference coordinates are a set of 32x32 coordinate points with equal intervals on the horizontal and vertical axes, denoted as p0. 01 Similarly, str = 16 times the reduced reference coordinates is the set of 1616 coordinate points p. 02 str = 32 times reduced reference coordinates, the set of coordinate points p is 44. 03 .
[0060] Step S2.2: Design a convolutional module that can take C downsampled feature maps as input. (For ease of expression, let's denote it as x) W,H After that, an offset coordinate feature map Δp is generated. n =S 2N,W,H The kernel size is w = 3. Let the convolution module be Conv(C, 2w = 6, w = 3).
[0061] Step S2.3: Calculate the new image feature matrix x after coordinate offset. H*W,C Because the attention mechanism's query matrix Q, key matrix K, and eigenvalue matrix V depend on the input image feature matrix x. H*W,C Therefore, we first calculate the new feature matrix of the image after the coordinate shift, and let p0+p n +Δp n The image feature matrix x in step S2.3 C,W,H Given the offset plane coordinates, the image features after bilinear interpolation are x. C (p0+p n +Δp n )=(1―u)(1―v)x C (p 0x +p nx ,p 0y +p ny )+u(1―v)x C (p 0x +p nx +1,p 0y +p ny )+(1―u)vxC (p 0x +p nx ,p 0y +p ny +1)+uvx C (p 0x +p nx +1,p 0y +p ny +1), u and v are Δp n Offset along the horizontal and vertical axes. C (p0+p n +Δp n The matrix representation of ) is x C,H,W Secondly, regarding x C,H,W The second and third dimensions are shrunk. C,H*W Finally, the image feature matrix x is obtained by transposing the matrix. H*W,C .
[0062] Step S3: Calculate attention features, including:
[0063] Step S3.1: Calculate the attention feature map after coordinate offset based on the attention.
[0064] (1) Using the image feature matrix x H*W,C and weight matrix respectively and Calculate x H*W,C The induced feature matrix Q has a feature dimension of D. H*W,D K H*W,D and V H*W,D These are abbreviated as Q, K, and V.
[0065] Wherein, the image feature matrix x H*W,C The dimension is (H*W)×C, where H and W are the height and width of the feature map, respectively, C is the number of channels of the feature map, and HW represents the length of the feature map's spatial dimension (height × width) flattened into a one-dimensional sequence.
[0066] Weight matrix:
[0067] The dimension is CxD, which is the weight matrix used to generate the "Query" features;
[0068] The dimension is CxD, which is the weight matrix used to generate the "key" feature;
[0069] The dimension is CxD, which is the weight matrix used to generate the "value" feature;
[0070] Where D is the transformed feature dimension.
[0071] Specifically,
[0072] By matrix multiplication, the image feature matrix x H*W,C is multiplied by three weight matrices respectively to obtain the induced feature matrix:
[0073]
[0074] The rule of matrix multiplication is that for any element Qi,j (i-th row, j-th column) in the result matrix, its value is equal to the sum of the corresponding elements of the i-th row of x H*W,C and the j-th column of , that is:
[0075] The calculation method of Ki,j and Vi,j is exactly the same, only replacing the corresponding weight matrix.
[0076] The essence of this process is to map the original image features (dimension C) to a new feature space (dimension D) through linear transformation (weight matrix), and generate the three role features required by attention mechanism:
[0077] Q (query): used to calculate the relevance with key features to determine the area that needs to be focused on the current position;
[0078] K (key): used to match with query features to provide association information between features;
[0079] V (value): based on the relevance of query and key, provide value information for generating attention features.
[0080] Through this conversion, the original features are adapted to the calculation framework of attention mechanism, laying a foundation for subsequent relevance analysis and weight allocation.
[0081] (2) Calculate the relevance matrix and perform normalization and softmax activation function processing
[0082] Where Q is the query matrix of (HW)XD (HW is the length of the flattened feature sequence, and D=29 is the feature dimension);
[0083] K T is the transpose of the key matrix of DX(HW) (obtained by transposing the K of (HW)XD dimension). The result of matrix multiplication QK^T is a square matrix of (HW)X(HW), where the element in the i-th row and j-th column represents the original relevance of the i-th query feature and the j-th key feature, and the calculation formula is: The similarity of two features in D-dimensional space is calculated by inner product, and the larger the value is, the stronger the relevance is.
[0084] Normalization: divide each element of the above matrix multiplication result by (D = 29), that is:
[0085] Effect: when D is large, the numerical range of the inner product result will expand with D, which may cause the gradient of the softmax function to disappear (because the function curve tends to be flat when the input value is too large). Divide by The numerical range can be normalized to a reasonable interval to ensure gradient stability.
[0086] Softmax activation function processing: apply the softmax function to the normalized matrix to obtain the correlation weight matrix αHW,HW, where the element in the ith row and jth column is: Properties and effects: the sum of each row of elements is 1, which realizes weight normalization and facilitates subsequent feature aggregation; amplifies the weight of high correlation elements (exponential function property), and suppresses low correlation elements, so that attention is focused on key feature areas.
[0087] (3) Calculate attention feature x H*W,D = α H*W,H*W *V, and perform dimension expansion processing on the attention feature of the image to obtain the attention feature map x H,W,D .
[0088] where, α H*W,H*W : attention weight matrix, dimension (H*W) x (H*W), where each element αi,j represents the attention weight of the ith feature position to the jth feature position (normalized by softmax, and the weight sum of each row is 1).
[0089] V: value feature matrix, dimension (H*W) x D (D = 29), where the jth row represents the "value" feature of the jth feature position (containing the key information of the position).
[0090] Output x H*W,D : attention feature matrix, dimension (H*W) x D.
[0091] Specifically, the calculation formula of the ith row and kth column element xi,k of matrix multiplication is:
[0092]
[0093] where:
[0094] αi,j is the attention weight of the ith row and jth column of the weight matrix;
[0095] Vj,k is the feature value of the jth row and kth column of the value matrix;
[0096] The calculation process is: all attention weights (i-th row) of the i-th position are weighted and summed with the "value" features (j-th row) of the corresponding position to finally obtain the attention features of the i-th position.
[0097] This process aggregates the global "value" features through attention weights. The "value" features (Vj,k) corresponding to positions with high weights (ai,j) have a higher proportion in the results, realizing the core function of the attention mechanism of "focusing on key features and suppressing irrelevant information".
[0098] Dimension expansion: from xHW,D to attention feature map xH,W,D
[0099] Original matrix dimension:
[0100] xHW,D is a one-dimensional sequence after flattening the spatial dimensions (height H x width W) of the feature map, with a dimension of (HW) x D.
[0101] Dimension expansion processing: the flattened one-dimensional sequence is restored to a two-dimensional spatial structure, and the specific operation is:
[0102] Keep the feature dimension D unchanged;
[0103] Split the one-dimensional sequence of length HW into a two-dimensional structure of height H and width W, that is: xH,W,D = reshape(xHW,D,(H,W,D)).
[0104] Dimension expansion maps the abstract one-dimensional attention features back to the spatial coordinates (H is the height direction and W is the width direction) of the original image, so that they recover the spatial structure of the input features Figure 1 , providing feature input conforming to visual logic for subsequent target detection (based on spatial position division detection grid).
[0105] Step S4, calculate the target detection loss, including:
[0106] Step S4.1, divide the feature matrix.
[0107] According to the attention feature map x H,W,D , make target detection features, where the detection grid of target detection has a total of HW, and each grid is set with 5 rectangular frame anchor points i = 1, 2, 3, 4, 5 are used to represent the edge boxes of different length-width ratios (for example, 4:3, 4:3, 2:5, 5:2, 1:1) At the same time, the top-left vertex and length-width of the rectangular frame are represented as The confidence of the fault features contained in the rectangular frame is Because the bearing vibration generally contains inner ring, outer ring and ball fault plus normal vibration mode, so each grid is set with 4 categories represented as j = 1, 2, 3, 4. Finally, the attention feature given by each grid is expressed as i = 1, 2, …, 5, j = 1, 2, …, 4, so the size of the feature dimension is D = 29.
[0108] Step S4.2, in order to calculate the loss value, that is, the difference between the feature expression calculated by the current model parameters and the real feature expression, the artificially labeled feature label is used as the true value, and the subsequent loss function is used to update the model parameters according to the chain rule.
[0109] The artificially labeled label style (z x , z y , z h , z w , l j ), j = 1, 2, …, 4, let the grid matching function be If the target bounding box B g centered in the grid is recorded as 1, and vice versa. The loss function is set to calculate the iterative reduction of the loss function value and update the network parameters in the process.
[0110] (1) The loss function of the category uses the cross-entropy loss function
[0111]
[0112] (2) The loss function of the bounding box intersection-over-union
[0113]
[0114] (3) The bounding box regression loss function
[0115]
[0116] (4) The confidence regression loss function
[0117]
[0118] Step S5, model prediction
[0119] By comparing the prediction results of the public standard data set COCO2017 and the commonly used models (yolov5, DETR), the average accuracy rate, the average accuracy rate of more than 50% intersection-over-union, and the average accuracy rate of more than 75% intersection-over-union can be improved by 2.7%, 2.0% and 2.7% respectively.
[0120] For example Figure 2As shown, the deformable attention-based feature extraction network process starts from the input image and gradually extracts features of different scales through multi-stage processing, including:
[0121] 1. Overall process and scale change
[0122] Input: The leftmost is the original image (1x scale, i.e., the original size of the input image), which is the starting data for the entire process.
[0123] Multi-stage processing: The network processes the image in stages to obtain features of 8x, 16x, and 32x scales. Here, "x" represents the scaling factor relative to the input image, such as 8x, which means the feature map size is 1 / 8 of the input image. As the processing progresses, the feature map scale gradually decreases, and the channel number and other feature expression capabilities generally increase.
[0124] 2. Operation modules at each stage
[0125] First stage (1x→8x)
[0126] convk=patch,s=patch: First, a convolution operation (conv) is performed, where "k=patch,s=patch" indicates that the convolution kernel size and step size are related to "patch" (which can be understood as the image block size). The purpose is to preliminarily downsample the input image and extract basic features, converting the 1x scale image into 8x scale preliminary features.
[0127] deformable attention: Then, the deformable attention module is applied. Deformable attention is an improvement over the standard attention mechanism, which can more flexibly capture the correlation of different positions in the image and adaptively adjust the attention focus area to further explore semantic and structural information on the 8x scale features.
[0128] Second stage (8x→16x)
[0129] convk=2,s=2: Use a convolution with a kernel size of 2 and a step size of 2 to downsample the 8x scale features again, obtaining 16x scale features, further compressing the spatial dimension and improving the feature abstraction degree.
[0130] deformable attention: Again, the deformable attention module is applied to strengthen feature correlation learning on the 16x scale features and capture more global and complex image information.
[0131] Third stage (16x→32x)
[0132] convk=2,s=2: continue to use the convolution kernel 2, step 2, and convert the 16x scale feature into the 32x scale feature, further reducing the spatial size.
[0133] deformableattention: the last deformable attention operation, which deeply mines the global correlation and subtle features of the image on the 32x scale feature, and outputs features with strong expression ability for subsequent tasks (such as detection, recognition, etc.).
[0134] 3. Output feature visualization
[0135] Under each "deformableattention" module, there are three types of visualization graphs: feature, ref, and pos.
[0136] feature: visualization of the feature map output after the module processing, showing the abstract expression of the feature (different colors and brightness represent the intensity of feature activation), reflecting the feature extraction result of the image content at this stage.
[0137] ref (reference): visualization of the reference point and reference correlation information in deformable attention, reflecting the distribution of key points used for correlation and reference in the attention mechanism, and showing the key position logic that the model focuses on and references.
[0138] pos (position): visualization of the position offset and sampling position related to deformable attention, reflecting how deformable attention flexibly adjusts the sampling position to adaptively capture image features, and showing the "deformable" characteristics of attention (such as irregular and flexible sampling area).
[0139] 4. Using convolution to gradually downsample, combined with deformable attention to flexibly capture feature correlation, deeply mine image information at different scales (8x / 16x / 32x), both through convolution to compress space and improve abstraction, and through deformable attention to break through the fixed sampling limit of conventional attention, more accurately and flexibly learn global and local features of the image, and provide multi-level and strongly expressed feature support for downstream tasks (such as target detection, segmentation, etc.).
[0140] Embodiment 2
[0141] The embodiment provides an automatic bearing fault diagnosis system based on an attention coordinate learnable paradigm, comprising:
[0142] The data acquisition module is configured to acquire bearing vibration signal data.
[0143] The cepstrum module is configured to calculate the cepstrum of the vibration signal according to the acquired bearing vibration signal data, and calculate the time-Mel spectrum feature map.
[0144] The attention coordinate module is configured to calculate attention coordinates according to the obtained time-mel frequency spectrum feature map;
[0145] The attention feature module is configured to calculate attention features according to the image feature matrix after the offset coordinates;
[0146] The detection module is configured to calculate a target detection loss according to the attention features, update network parameters by iteratively reducing the loss function value, and realize automatic diagnosis of bearing faults.
[0147] A computer readable storage medium, wherein a plurality of instructions are stored, the instructions are suitable for being loaded and executed by a processor of a terminal device, and the instructions implement the bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm.
[0148] A terminal device, comprising a processor and a computer readable storage medium, the processor is used to implement instructions, and the computer readable storage medium is used to store a plurality of instructions, the instructions are suitable for being loaded and executed by the processor, and the instructions implement the bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm.
[0149] The above are preferred embodiments of the present application, not limited to the protection scope of the present application, therefore: any equivalent changes made according to the structure, shape, principle of the present application should be covered within the protection scope of the present application.
Claims
1. A bearing fault automatic diagnosis method based on an attention coordinate learnable paradigm, characterized in that, The method comprises the following steps: obtaining bearing vibration signal data; calculating a cepstrum of the vibration signal according to the obtained bearing vibration signal data, and obtaining a time-Mel spectrum feature map; calculating attention coordinates according to the obtained time-Mel spectrum feature map; calculating attention features according to the image feature matrix after the shift of the coordinates; calculating a target detection loss according to the attention features, updating network parameters by iteratively reducing the loss function value, and realizing automatic diagnosis of bearing faults.
2. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 1, characterized in that, The bearing vibration signal data is acquired, including setting a sampling frequency f of the vibration signal, dividing the vibration signal into a plurality of frame sequences, setting a frame length as 512 times of a sampling period T s = 512 / f, considering that the number of frames acquired within 10 seconds is N s = 10 / T s and taking the lower integer, adding a Hamming window to the vibration signal of each frame, setting a window function denotes rounding up; performing Fourier transform on the Hamming weighted vibration signal to generate a power spectrum, the frequency range is 0~f / 2, a total of N f = 257 frequency points, calculating the corresponding power spectrum of each frame to generate a conventional time-frequency spectrum diagram 3. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 2, characterized in that, The inverse spectrum diagram of the vibration signal is calculated according to the acquired bearing vibration signal data, and the calculation of the inverse spectrum diagram of the vibration signal comprises calculating a mel-frequency-normal frequency filter response matrix Wherein, design N m =64 triangular filters, the conversion relationship between mel frequency and normal frequency is f m =2595*log(10*(1+700f)), the equal interval mel frequency f m (z) is in the range of 0~2595*log(10*(1+700f / 2)), z=0,1,…,N m +1, a total of N m +2 frequency points; converting the equal interval mel frequency back to normal frequency, the left frequency, center frequency and right frequency of the triangular filter in the normal frequency coordinate system are And 4. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 3, characterized in that, The time-Mel spectrum feature map is calculated, and the response value of the zth row and the kth column of a Mel spectrum-normal spectrum filter response matrix is calculated as follows: Further, a time-Mel-spectral feature map is computed, denoted as 5. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 4, characterized in that, The attention coordinates are calculated according to the obtained time-Mel spectrum feature map, including setting a reference coordinate p0, using three reference coordinates according to the feature map after three times of downsampling, and the reference coordinates of the image reduced by str=8 times are 32*32 coordinate point sets p 01 , the reference coordinates of the image reduced by str=16 times are 16*16 coordinate point sets p 02 , and the reference coordinates of the image reduced by str=32 times are 4*4 coordinate point sets p 03 , and a convolution module is designed to input C feature maps after downsampling to generate an offset coordinate feature map Δp n =S 2N,W,H , and the size of the convolution kernel is w=3.
6. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 5, characterized in that, The step of calculating attention coordinates based on the obtained time-Mel spectrum feature map also includes calculating a new image feature matrix x after coordinate shift. H*W,C First, the new feature matrix of the image after the offset coordinates is calculated, and p0+p n +Δp n For the image feature matrix x C,W,H Given the offset plane coordinates, the image features after bilinear interpolation are x. C (p0+p n +Δp n )=(1―u)(1―v)x C (p 0x +p ny ,p 0y +p ny )+u(1―v)x C (p 0x +p nx +1,p 0y +p ny )+(1―u)vx C (p 0x +p nx ,p 0y +p ny +1)+uvx C (p 0x +p nx +1,p 0y +p ny +1), u and v are Δp n Offset in the horizontal and vertical directions; x C (p0+p n +Δp n The matrix representation of ) is x C,H,W Secondly, regarding x C,H,W The second and third dimensions are shrunk. C,H*W Finally, the image feature matrix x is obtained by transposing the matrix. H*W,C .
7. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 6, characterized in that, The image feature matrix after the offset coordinates is calculated, including calculating the image feature according to the attention coordinates, wherein the image feature matrix x H*W,C is calculated by the weight matrix and . H*W,C The feature matrix Q with a feature dimension of D induced by x H*W,D , K H*W,D and V H*W,D is calculated, which is denoted as Q, K and V; the correlation matrix is calculated and normalized and processed by a softmax activation function α H*W,H*W = softmax(Q*K T / √D); the attention feature x H*W,D = α H*W,H*W *V is calculated, and the attention feature of the image is dimensionally expanded to obtain the attention feature map x H,W,D .
8. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 7, characterized in that, The target detection loss is calculated according to the attention feature, including calculating the target detection loss according to the attention feature map x H,W,D As the target detection feature, where the detection grid of the target detection has a total of HW, and 5 rectangular frame anchor points are set for each grid The bounding box is used to represent different aspect ratios At the same time, the upper left point and the length-width of the rectangular bounding box are represented as The confidence of the fault feature contained in the rectangular frame is Each detection grid is set to represent four categories The attention feature given by each grid is finally expressed as The size of the feature dimension is D=29.
9. The bearing fault automatic diagnosis method based on the attention coordinate learnable paradigm according to claim 8, characterized in that, The target detection loss calculated according to the attention feature further comprises setting a calculation mode of a loss function to iteratively reduce the value of the loss function and update network parameters in the process, wherein the loss function of the category is a cross-entropy loss function, denoted as: The loss function of the bounding box intersection-over-union ratio is denoted as: The bounding box regression loss function is denoted as: The confidence regression loss function is denoted as:
10. An automatic bearing fault diagnosis system based on attention coordinate learnable paradigm, characterized in that, The method comprises the following steps: a data acquisition module configured to obtain bearing vibration signal data; a cepstrum module configured to calculate a cepstrum of the vibration signal according to the obtained bearing vibration signal data, and obtain a time-Mel spectrum feature map; an attention coordinates module configured to calculate attention coordinates according to the obtained time-Mel spectrum feature map; an attention features module configured to calculate attention features according to the image feature matrix after the shift of the coordinates; a detection module configured to calculate a target detection loss according to the attention features, update network parameters by iteratively reducing the loss function value, and realize automatic diagnosis of bearing faults.