Semantic segmentation method for vibration displacement measurement of disc brake

CN119992092AActive Publication Date: 2025-05-13KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510095304.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13
Estimated Expiration
2045-01-21

Smart Images

  • Figure CN119992092A_ABST
    Figure CN119992092A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic segmentation method for vibration displacement measurement of a disc brake, and belongs to the field of visual vibration displacement measurement and computer vision. The method is used for predicting each frame of an obtained vibration video of the disc brake through a semantic segmentation network model to obtain a corresponding segmentation mask; the semantic segmentation network model comprises an encoder and a decoder, Resnet101 is adopted as a network backbone at the encoder, 7 * 7 convolution and maximum pooling are replaced by a down-sampling module based on Harr wavelet transform, and a global-local feature extraction module is inserted into an original residual block of the Resnet101 network; a collaborative attention-based intra-scale feature interaction module and a Laplacian image-guided cross-scale feature fusion module are adopted in a decoder part. The method pays attention to the balance relation between local and global information, effectively emphasizes edge information in the image, makes up for the deficiency of detail and local structure information in a deep network, improves the positioning precision of a network model for the disc brake, and improves the positioning accuracy of the network model for the disc brake. And the sensitivity degree and the capture capability of the semantic segmentation model to the subtle change of the disc brake in a dynamic environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a semantic segmentation method for disc brake vibration displacement measurement, and belongs to the field of visual vibration displacement measurement and computer vision. Background Art

[0002] Disc brakes are widely used in industry and transportation due to their efficient braking force. However, brake disc vibration caused by wear may cause a series of problems, including reduced braking performance, increased braking noise, and increased wear. These problems will not only shorten the service life of components, but may even cause serious safety accidents. Therefore, it is of great significance to accurately obtain the vibration displacement curve of the center point of the disc brake to determine the fault type, which is of great significance for the measurement of disc brake vibration displacement in complex scenarios.

[0003] There are two methods for measuring the vibration displacement of disc brakes: contact and non-contact. The contact measurement method is to install physical sensors on the surface of the brake, which is not only greatly affected by the environment, but also changes the dynamic characteristics of the brake; the non-contact measurement method includes the use of ultrasonic, laser and displacement sensors, etc. These methods are difficult to adapt to high-speed scenes. The visual vibration measurement method based on deep learning can predict the pixel points by learning a large amount of data, thereby measuring the vibration displacement of the disc brake. The existing deep learning method faces the problem of inaccurate boundary recognition, which will cause large errors in the measurement of disc brake displacement. Summary of the invention

[0004] The present invention provides a semantic segmentation method for disc brake vibration displacement measurement, which aims to extract image features by using Resnet101 as the network backbone, adopting a downsampling strategy based on wavelet transform (HWTD) and a global-local feature extraction residual structure (GL-Bottleneck), and using a multi-scale feature fusion strategy guided by Laplacian image (LIFFN) to effectively integrate features from different coding layers to output the final segmentation result.

[0005] The technical solution of the present invention is:

[0006] According to a first aspect of the present invention, a semantic segmentation method for disc brake vibration displacement measurement is provided, comprising the following steps: obtaining disc brake vibration video data; each frame of the video is predicted by a semantic segmentation network model to obtain a corresponding segmentation mask; the semantic segmentation network model comprises an encoder and a decoder, wherein the encoder uses Resnet101 as the network backbone, and replaces the 7×7 convolution and maximum pooling with a downsampling module based on Harr wavelet transform, so that the output feature map has more detailed information of the disc brake in the image frame, and inserts a global-local feature extraction module into the original residual block of the Resnet101 network to improve the semantic segmentation network's ability to extract detailed information of the disc brake. The representation ability of segmentation and local structure information is enhanced, and feature layers f2, f3, f4 and f5 are obtained through the encoder; in the decoder part, the intra-scale feature interaction module based on collaborative attention and the cross-scale feature fusion module guided by the Laplacian image are adopted, and the feature layer f5 is input into the intra-scale feature interaction module based on collaborative attention. The spatial semantic information of the disc brake image is extracted from the feature layer through the intra-scale feature interaction module based on collaborative attention and the spatial semantic difference is reduced to obtain the feature layer P5. The original input image f1, feature layer f2, feature layer f3, feature layer f4, feature layer f5 and feature layer P5 are input into the cross-scale feature fusion module guided by the Laplacian image to obtain the corresponding segmentation mask.

[0007] Furthermore, the encoder is specifically as follows: first, the input image f1 is processed twice by a Haar wavelet-based downsampling module, and after being processed by the first Haar wavelet-based downsampling module, a feature layer f2 is obtained; the feature layer f2 is used as the input of the second Haar wavelet-based downsampling module, and the output of the second Haar wavelet-based downsampling module is sequentially passed through three residual blocks inserted into the global-local feature extraction module, and the output results of the three residual blocks inserted into the global-local feature extraction module are feature layers f3, f4 and f5 respectively.

[0008] Furthermore, the specific process of the Haar wavelet-based downsampling module is as follows: first, the resolution of the disc brake image is halved; then, its signal is decomposed into two parts, namely, a low-frequency L component and high-frequency HH, HV and HD components, and they are spliced ​​in the channel dimension; finally, a feature map is obtained through convolution, normalization and Relu activation function processing.

[0009] Furthermore, the normalization selects to use Group Normalization.

[0010] Furthermore, the residual block inserted with the global-local feature extraction module is specifically as follows: a global-local feature extraction module is added after the 1×1 and 3×3 convolutions of the original residual block part, and the global-local feature extraction module performs local adaptive average pooling on the input image to obtain local spatial information; it is then divided into two branches, the first branch performs global adaptive average pooling on the obtained local features, performs a convolution operation when the spatial dimension is reduced to 1, and then expands to the original input size of the module; the second branch performs shape adjustment on the obtained local features, extracts deeper features through one-dimensional convolution, and then returns to the original input size of the module through shape adjustment; the features obtained from the two branches are weighted fused, and then restored to the same size as the input features through adaptive inverse average pooling to obtain an attention map; finally, the input features are multiplied by the generated attention map to obtain the output of the global-local feature extraction module.

[0011] Furthermore, the intra-scale feature interaction module based on collaborative attention is specifically as follows: the input feature layer f5 is average pooled, and the average values ​​are taken along the height and width dimensions respectively, and the feature maps of the two dimensions are obtained, and each of them is evenly divided into 4 groups of non-overlapping sub-features along the corresponding dimensions; then, a separable one-dimensional convolution with multi-receptive field depth sharing is applied to these sub-features, and they are respectively spliced, normalized and activated, and then re-adjusted back to the original dimensions to form a spatial attention map in two dimensions; and then the obtained result is element-wise multiplied with the initial input feature layer f5 of the intra-scale feature interaction module based on collaborative attention to obtain The weighted features that integrate global contextual dependencies and multi-semantic spatial information are average pooled and normalized, and the dimensions are transformed through two-dimensional convolution to generate query q, key k and value v; then, the self-attention score between k and q is calculated along the channel dimension using the dot product method, and the channel attention relationship weight is obtained through softmax and dot product calculation is performed with v; finally, the output dimension of the self-attention is restored to the input shape, and after average pooling and activation function, the output feature is multiplied with the input "weighted features that integrate global contextual dependencies and multi-semantic spatial information" to obtain the feature layer P5.

[0012] Furthermore, the Laplacian image guided cross-scale feature fusion module is specifically: the feature layer f j+1 Perform 2 times upsampling and then combine with the feature layer f j Perform subtraction operations to obtain Laplacian images L of various scales j; Wherein, j = 1, 2, 3, 4; feature layer P5 is fused with feature layer f4 in the first fusion module to obtain output feature layer x4; feature layer P5 is subjected to five first convolution modules to obtain feature layer R5; feature layer R5 is upsampled twice and then stacked with x4 and L4 to obtain feature layer P4; feature layer P4 is subjected to four first convolution modules and then added to L4 to obtain R4; after P4 is subjected to one first convolution module, it is subjected to one first fusion module and then fused with feature layer f3 to obtain output feature layer x3; feature layer R4 is upsampled twice and then stacked with x3 and L3 to obtain feature layer P3; feature layer P3 is subjected to three first convolution modules and then added to L3 to obtain R3; after P3 is subjected to one first convolution module, it is subjected to one first fusion module and then combined with feature layer f2 is fused to obtain the output feature layer x2; the feature layer R3 is upsampled twice and then stacked with x2 and L2 to obtain the feature layer P2; the feature layer P2 passes through the first convolution module twice and is added with L2 to obtain R2; P2 passes through the first convolution module once and then passes through the first fusion module to obtain the feature layer x1 by fusing with the feature layer f1; the feature layer R2 is upsampled twice and then stacked with x1 and L1 to obtain the feature layer P1; the feature layer P1 passes through three convolution modules and is added with L1 to obtain R1; R5 is upsampled twice and then added with R4 to obtain D4; D4 is upsampled twice and then added with R3 to obtain D3; D3 is upsampled twice and then added with R2 to obtain D2; D2 is upsampled twice and then added with R1 to obtain D1, which is the final output result of the semantic segmentation network.

[0013] According to a second aspect of the present invention, a processor is provided, the processor being configured to execute an operation, wherein the operation comprises executing any one of the semantic segmentation methods for disc brake vibration displacement measurement described above.

[0014] The beneficial effects of the present invention are as follows: the present invention uses a high-speed industrial camera as an image acquisition medium, and uses a disc brake in a high-speed video as an object of vibration displacement measurement, introduces a semantic segmentation method based on deep learning into the field of visual vibration measurement of disc brakes, and verifies the feasibility of deep learning methods in visual vibration measurement of disc brakes from multiple angles. Specifically: the present invention uses the classic semantic segmentation network Resnet101 as a basic framework, adopts a downsampling strategy based on wavelet transform (HWTD) and a global-local feature extraction residual structure (GL-Bottleneck) to extract image features, and uses a multi-scale feature fusion strategy (LIFFN) guided by a Laplacian image to effectively integrate features from different coding layers to output the final segmentation result; according to the segmentation result, the coordinates of the center point of the image are extracted frame by frame, so as to more accurately measure the vibration displacement offset. The new semantic segmentation method proposed in the present invention focuses on the balance between local and global information, effectively emphasizes the edge information in the image, makes up for the lack of details and local structure information in the deep network, improves the positioning accuracy of the network model for the disc brake, and improves the sensitivity and capture ability of the semantic segmentation network model to subtle changes of the disc brake in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a work flow chart;

[0016] Figure 2 This is the encoder structure diagram for semantic segmentation;

[0017] Figure 3 It is a downsampling module based on Haar wavelet transform;

[0018] Figure 4 It is a diagram of the intra-scale feature interaction module based on collaborative attention;

[0019] Figure 5 Graph of the process of generating Laplacian images;

[0020] Figure 6 The semantic segmentation network model diagram for introducing the Laplacian image-guided multi-scale feature fusion module;

[0021] Figure 7 Comparison chart of vibration displacement curves of different algorithms;

[0022] Figure 8 This is a comparison chart of axis trajectories of different algorithms;

[0023] Fig. 9 This is a comparison chart of amplitude-frequency curves of different algorithms. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. It should be noted that the embodiments in this application and the features in the embodiments can be combined with each other arbitrarily without conflict.

[0025] Example 1: Figure 1-9 As shown, according to a first aspect of an embodiment of the present invention, a semantic segmentation method for disc brake vibration displacement measurement is provided, comprising: obtaining disc brake vibration video data; each frame of the video is predicted by a semantic segmentation network model to obtain a corresponding segmentation mask; the semantic segmentation network model comprises an encoder and a decoder, wherein Resnet101 is used as the network backbone in the encoder, and the 7×7 convolution and maximum pooling therein are replaced with a downsampling module (HWTD) based on Harr wavelet transform, so that the output feature map has more detailed information of the disc brake in the image frame, and a global-local feature extraction module (Global-Local Feature Extractor, GL-Bottleneck) is inserted into the original residual block of the Resnet101 network to enhance the semantic segmentation network's ability to represent disc brake details and local structural information, and feature layers f2, f3, f4 and f5 are obtained through the encoder; a Laplace Image-guided Feature Fusion Network (Laplace Image-guided Feature Fusion) is used in the decoder part. A novel LIFFT Network (LIFFN) is proposed, which includes a collaborative attention-based intra-scale feature interaction module (SAIFI) and a Laplacian image-guided cross-scale feature fusion module (LICFF). The feature layer f5 is input into the collaborative attention-based intra-scale feature interaction module. The spatial semantic information of the disc brake image is extracted from the feature layer through the collaborative attention-based intra-scale feature interaction module and the spatial semantic difference is mitigated to obtain the feature layer P5. The original input image f1, feature layer f2, feature layer f3, feature layer f4, feature layer f5 and feature layer P5 are input into the Laplacian image-guided cross-scale feature fusion module to obtain the corresponding segmentation mask.

[0026] Furthermore, if Figure 2As shown in the figure, the Resnet101 backbone feature extraction network is constructed, and the 7×7 convolution and maximum pooling are replaced with a downsampling module (HWTD) based on Haar wavelet transform, and a global-local feature extraction module (GL-Bottleneck) is inserted into the original residual block of the Resnet101 network to reduce the spatial resolution of the image while retaining more effective semantic information. The encoder is specifically as follows: first, the input image f1 is processed twice by a downsampling module based on Haar wavelet, and after the first downsampling module based on Haar wavelet is processed, a feature layer f2 is obtained; the feature layer f2 is used as the input of the second downsampling module based on Haar wavelet, and the output of the second downsampling module based on Haar wavelet is sequentially passed through three residual blocks inserted into the global-local feature extraction module, and the output results of the three residual blocks inserted into the global-local feature extraction module are feature layers f3, f4 and f5 respectively.

[0027] Furthermore, Resnet101 is used as the backbone network to perform preliminary feature extraction on the input disc brake image. First, the input image f1 is downsampled twice based on Haar wavelet. After the first downsampling based on Haar wavelet, the feature layer f2 is obtained. Figure 3 As shown in the figure, the specific process of the downsampling strategy based on Haar wavelet is as follows: first, the resolution of the disc brake image is halved; then its signal is decomposed into two parts: low-frequency (L) component and high-frequency component (HH, HV and HD), and they are spliced ​​in the channel dimension, and the channel is expanded. Finally, the feature map is obtained through convolution, normalization and Relu activation function processing, in which Group Normalization is selected for normalization. Figure 2As shown in the figure, the residual block adds a global-local feature extraction module after 1×1 and 3×3 convolutions. This module performs local adaptive average pooling on the input image to obtain local spatial information. It is then divided into two branches. The first branch performs global adaptive average pooling on the local features obtained, performs convolution when the spatial dimension is reduced to 1, and then expands to the original input size of the module; the second branch adjusts the shape of the local features obtained, extracts deeper features through one-dimensional convolution, and then adjusts the shape back to the original input size of the module. The features obtained from the two branches are weighted fused, and then restored to the same size as the input features through adaptive anti-average pooling to obtain the attention map; finally, the input features are multiplied with the generated attention map to obtain the output of the global-local feature extraction module. The residual block first uses 1×1 to reduce the number of channels of the image, and then uses 3×3 convolution to extract features from the image. The extracted features are multiplied by the input features and the generated attention map through the global-local feature extraction module, and then the dimensions are restored using 1×1 convolution, and finally added to the input image. The output results of the three residual blocks are feature layers f3, f4 and f5 respectively.

[0028] In the above, Group Normalization is selected for normalization. The purpose of using it is: considering that when the Batch size is small, the calculated mean and variance of BatchNormalization are not accurate enough, resulting in poor normalization effect. Although LayerNormalization does not depend on the Batch size, it ignores the information between channels and only normalizes between neurons in the same layer. Therefore, the present invention chooses to use Group Normalization, which not only retains the dependency between channels, but also can maintain a good normalization effect in small batches.

[0029] Furthermore, the decoder part adopts a Laplacian image guided multi-scale feature fusion strategy (LIFFN), which includes a collaborative attention-based intra-scale feature interaction module (SAIFI) and a Laplacian image guided cross-scale feature fusion module (LICFF) to enhance the feature extraction of images.

[0030] Furthermore, if Figure 4As shown in the figure, the extracted effective feature layer f5 is input into the scale feature interaction module based on collaborative attention, which is designed to guide channel attention learning: first, the input feature f5 is average pooled, and the average values ​​are taken along the height and width dimensions respectively, and the feature maps of the two dimensions are obtained, which are divided into 4 groups of non-overlapping sub-features along the corresponding dimensions; then, the separable one-dimensional convolution DWConv1d with multi-receptive field depth sharing is applied to these sub-features, and they are spliced, normalized and activated respectively, and then adjusted back to the original dimension to form a spatial attention map in two dimensions; then the obtained result is compared with the initial input feature map of the scale feature interaction module based on collaborative attention. f5 obtains the weighted features that integrate global contextual dependencies and multi-semantic spatial information by element-wise multiplication, performs average pooling and normalization on the features, and transforms the dimensions through two-dimensional convolution DWConv to generate query (q), key (k) and value (v); then, the self-attention score between k and q is calculated along the channel dimension using the dot product method, and the channel attention relationship weight is obtained through softmax and the dot product is calculated with v; finally, the output dimension of the self-attention is restored to the input shape, and after average pooling and activation function, the output feature is multiplied with the input "weighted features that integrate global contextual dependencies and multi-semantic spatial information" to obtain the feature layer P5.

[0031] Furthermore, the results of the feature layers f1, f2, f3, f4 and f5 and the intra-scale feature interaction module based on collaborative attention are input into the Laplacian image-guided cross-scale feature fusion module.

[0032] First, if Figure 5 As shown, the method for calculating the Laplace residuals of each level of the input disc brake RGB image is as follows: j traverses 1 to 4 and converts the feature layer f j+1 Perform 2 times upsampling and then combine with the feature layer f j Perform subtraction operations to obtain Laplacian images L of various scales j .

[0033] L j =f j -Up(f j+1 ),j=1,2,3,4

[0034] From the feature extraction part, we can know that Up is a 2x upsampling function, and the feature layer f j+1 The size of f j 1 / 2 times.

[0035] Then, if Figure 6As shown, the feature layer P5 is fused with the feature layer f4 after being processed by standardization, activation function, upsampling and convolution in the first fusion module, and then outputs the feature layer x4 through one convolution; the feature layer P5 passes through the first convolution module five times, that is, after being processed by standardization, activation function, weight normalization and convolution five times, the feature layer R5 is obtained;

[0036] The feature layer R5 is upsampled twice and then stacked with x4 and L4 to obtain the feature layer P4; the feature layer P4 undergoes four first convolution modules, namely standardization, activation function, weight standardization and convolution processing, and then is added with L4 to obtain R4; P4 undergoes one first convolution module and then one first fusion module, in which it undergoes standardization, activation function, upsampling and convolution processing, and then is fused with the feature layer f3, and then outputs the feature layer x3 through one convolution;

[0037] The feature layer R4 is upsampled twice and then stacked with x3 and L3 to obtain the feature layer P3; the feature layer P3 passes through the first convolution module three times and is added with L3 to obtain R3; P3 passes through the first convolution module once and then the first fusion module once, where it is fused with the feature layer f2 after being normalized, activated, upsampled and convolved, and then outputs the feature layer x2 through one more convolution;

[0038] The feature layer R3 is upsampled twice and then stacked with x2 and L2 to obtain the feature layer P2; the feature layer P2 passes through the first convolution module twice and is added with L2 to obtain R2; P2 passes through the first convolution module once and then passes through the first fusion module once, where it is fused with the feature layer f1 after being normalized, activated, upsampled and convolved, and then outputs the feature layer x1 through another convolution;

[0039] The feature layer R2 is upsampled twice and then stacked with x1 and L1 to obtain the feature layer P1; the feature layer P1 passes through three convolution modules and is added to L1 to obtain R1.

[0040] In the above, about R j The expression of R is as follows: j =B j ([x j ,L j ,Up(R j+1 )])+L j ,j=1,2,3,4; where Bj means stacking and Up means double upsampling.

[0041] Finally, R5 is upsampled twice and added to R4 to obtain D4; D4 is upsampled twice and added to R3 to obtain D3; D3 is upsampled twice and added to R2 to obtain D2; D2 is upsampled twice and added to R1 to obtain D1, which is the final output result of the semantic segmentation network.

[0042] In the above, about D j The expression is as follows:

[0043] This process proceeds layer by layer, mapping the features back to the size of the original disc brake image from coarse to fine, successfully restoring the local details and global layout of the disc brake.

[0044] In specific applications, the model proposed in the present invention can be trained first, and the semantic segmentation model of the disc brake is added with the Focal-Loss loss function during training.

[0045] The deep learning algorithm used in the present invention is run on a desktop computer (equipped with an Intel Core i7-13700K processor, NVIDIA GeForce RTX 4080 graphics card, and 32G video memory) equipped with a unified operating environment (Windows 10, Cuda 11.8, Pytorch 2.0.1, torch-vision 0.15.2).

[0046] The disc brake vibration displacement data set is formed by taking disc brake vibration data with the help of a high-speed camera. In the constructed data acquisition system, a high-speed industrial camera is used as an image data sensor, and an eddy current sensor is used as a detection sensor for real vibration displacement, and the two synchronously collect signals. The data acquisition system is mainly composed of three parts: a vibration test and control test bench, a vibration measurement system, and an image acquisition system; the vibration test and control test bench covers the rotor test bench, disc brake, and power supply equipment; the vibration measurement system consists of two eddy current sensors with a range of two millimeters, two signal acquisition cards, a DH5923 dynamic signal acquisition analyzer, and a desktop computer equipped with special software DHDAS; the image acquisition system includes a Thousand Eye Wolf high-speed camera M220M, a corresponding light source, and a laptop computer equipped with high-speed acquisition software. Before the data acquisition work is carried out, the two ends of the signal acquisition card need to be connected to the DH5923 dynamic signal acquisition analyzer and the eddy current sensor connected to the desktop computer, and at the same time, ensure that the direction of the eddy current sensor probe used to measure radial and tangential displacement coincides with the horizontal and vertical center lines of the disc brake. The signal sampling rate of the eddy current sensor is set to 2000fps; the high-speed camera is placed 1500mm in front of the disc brake, with a resolution of 960×960 and a frame rate of 2000fps; the light source is placed diagonally in front of the disc brake, with the brightness adjusted to 75%. The voltage and current values ​​of the power supply are adjusted to ensure that the rotor can run smoothly. With the help of the synchronous sampling program, the high-speed industrial camera and the eddy current sensor can synchronously collect the disc brake image sequence and voltage displacement signal.

[0047] Furthermore, 500 frames were randomly selected from the collected video, and the labeling software Labelme was used to label the disc brake target in the image, i.e., the rotor boundary box, to obtain coordinate information; the training data set and the validation data set accounted for 90% and 10% of the total disc brake vibration displacement data set, respectively. The semantic segmentation network model was trained using the training data set, and then the model was evaluated using the test data set to obtain the optimal weight file and load the optimal weight file into the semantic segmentation network model.

[0048] Furthermore, the collected data is input into the trained semantic segmentation model for prediction, such as Figure 6 As presented, the segmentation mask belonging to the disc brake category is extracted and the coordinates of each pixel are obtained from it. Then the average of these coordinates is calculated to calculate the coordinates of the center point of the segmentation mask for each frame. The average value of the coordinates of the center point of the disc brake in each frame is used as the benchmark reference value to ensure that all displacement offsets of the vibration signal meet the zero mean distribution requirements. By calculating the displacement offset between each frame and the reference value, the vibration displacement curve of the disc brake during rotation can be obtained.

[0049] Furthermore, the present invention uses three evaluation indicators to evaluate the performance of the semantic segmentation network model: recall rate (Recall), category average pixel accuracy (mean PixelAccuracy, mPA) and mean intersection over Union (meanIntersection over Union, mIoU). The calculation formula is as follows:

[0050]

[0051] Among them, TP represents the correctly predicted positive sample, FP represents the incorrectly predicted positive sample, TN represents the correctly predicted negative sample, and FN represents the incorrectly predicted negative sample. Disc The intersection over union (IoU) of the target representing the disc brake background The intersection-over-combination ratio of the background representing the disc brake; PA Disc Represents the pixel accuracy of the disc brake target, PA background Pixel accuracy of the background representing the disc brake.

[0052] Furthermore, frequency domain indicators and time domain indicators are two commonly used indicators in vibration detection and evaluation. Therefore, the present invention combines the time domain and frequency domain indicators to construct an evaluation system for visual vibration measurement tasks, thereby more accurately and comprehensively reflecting the changes in the motion state of the disc brake.

[0053] The time domain index evaluates the vibration signal of time, which reflects the change of the vibration signal on the time axis. The present invention selects waveform factor error (WFE), clearance factor error (CFE), and normalized displacement root mean square error (d-NRMSE) as time domain evaluation indicators.

[0054] The calculation process of WFE, CFE, and d-NRMSE is as follows:

[0055]

[0056] WFE=|WF1-WF2|

[0057] Where n is the number of sampling points, x i is the displacement data of the i-th sampling point in the vibration signal, WF1 and WF2 are the waveform factors of the vibration displacement signal measured by the semantic segmentation method of the present invention and the eddy current sensor, respectively.

[0058]

[0059] CFE=|CF1-CF2|

[0060] Among them, CF1 and CF2 are the margin factors of the semantic segmentation method of the present invention and the vibration displacement signal measured by the eddy current sensor, respectively.

[0061]

[0062] in, and x i are the displacement data of the i-th sampling point in the vibration displacement signal measured by the semantic segmentation method of the present invention and the eddy current sensor, respectively, max and x min They are the maximum and minimum values ​​of the real displacement measured by the eddy current sensor. The smaller the value of d-NRMSE, the better the fitting effect between the vibration displacement curve predicted by the network model and the real displacement curve, which better meets the actual needs of this experimental scenario.

[0063] Frequency domain indicators are methods for evaluating the frequency components of vibration signals. The present invention selects frequency center of gravity error (CGFE), frequency-root mean square error (f-RMSE), and fundamental frequency amplitude error (FFAE) as frequency domain evaluation indicators.

[0064] The calculation process of CGFE, f-RMSE, and FFAE is as follows:

[0065]

[0066] CGFE=|CGF1-CGF2|

[0067] Where m is the number of main frequency components, f k is the frequency value of the kth frequency component, and s(k) is the corresponding amplitude.

[0068]

[0069] f-RMSE=|f-RMS1-f-RMS2|

[0070] Among them, f-RMS1 and f-RMS2 are the frequency root mean square of the vibration displacement signal measured by the semantic segmentation method and the eddy current sensor, respectively.

[0071] FFAE=|FFA1-FFA2|

[0072] Among them, FFA1 and FFA2 are the fundamental frequency amplitudes of the vibration displacement signals measured by the semantic segmentation method and the eddy current sensor, respectively.

[0073] Furthermore, in order to evaluate and understand the effectiveness and contribution of this method to the overall performance of the semantic segmentation network model and the vibration displacement measurement of the disc brake, five groups of ablation experiments were conducted on the self-made data set using the UNet network as the baseline network model (Baseline), and each group was superimposed on the previous group. That is, in Table 1, "+Layeradjust" means adding "Layer adjust" on the basis of "Baseline", that is, Resnet50 in the Unet encoder part is replaced with Resnet101; "+LIFFN" means adding "LIFFN" on the basis of "Baseline+Layer adjust", that is, adding the LIFFN module to the encoder part; "+HWTD" means adding HWTD on the basis of "Baseline+Layer adjust+LIFFN", that is, using the HWTD module in the encoder part; "+GL-Bottleneck" means adding GL-Bottleneck on the basis of "Baseline+Layer adjust+LIFFN+HWTD", that is, using GL-Bottleneck in the decoder part. In the ablation experiment, the present invention adjusts the number of encoder layers and uses a deeper network structure to adapt to the disc brake micro-vibration measurement task that requires high precision and complex feature extraction. The detailed results of the ablation experiment are shown in Table 1.

[0074] Table 1 Statistics of ablation experiment results

[0075]

[0076] From the ablation experiment results in Table 1, the model proposed in the present invention can achieve the network performance of the existing benchmark model. Furthermore, in vibration analysis, time domain indicators are used to evaluate the time domain characteristics in the vibration signal, such as the degree of impact and periodic vibration; while frequency domain indicators are used to evaluate the frequency components of the vibration signal, such as resonance, natural frequency, and specific fault frequency. Combining these two indicators can more accurately diagnose the health status of the disc brake, detect problems early, and take maintenance measures. Among the evaluation indicators of the visual vibration measurement task of the present invention (NRMSE, WFE, CGFE and f-RMSE), only the WFE in the y direction is slightly lower than the baseline model, and the remaining time domain and frequency domain indicators have achieved significant improvements. This shows that the adjustment of the number of network layers significantly improves the accuracy of vibration displacement measurement.

[0077] Specifically, it can be seen from Table 1 that after introducing LIFFN into the semantic segmentation model, the d-NRMSE index is reduced by 1.5% in the x direction and 0.09% in the y direction; the CGFE index is reduced by 32.0370347 in the x direction and 4.63383015 in the y direction; the f-RMSE index is reduced by 30.6498079 in the x direction and 2.73538304 in the y direction. The overall performance of the network and the vibration displacement measurement accuracy are further enhanced.

[0078] After introducing HWTD into the semantic segmentation model, the Recall index increased to 99.67%; the mPA index increased to 99.83%; the mIoU index increased to 99.65%; and the NRMSE index decreased to 0.049 and 0.0259 in the x and y directions respectively.

[0079] After introducing GL-Bottleneck into the semantic segmentation model, all vibration displacement measurement evaluation indicators performed well. Compared with the baseline model, the final algorithm reduced the NRMSE, CGFE and f-RMSE indicators in the x direction by 5.7%, 69.51427, and 60.5425635, respectively; and reduced the NRMSE, WFE, CGFE and f-RMSE in the y direction by 3.23%, 0.014596783, 30.41015573, and 31.15941723, respectively. GL-Bottleneck improves the model's attention to the detailed structural information of the disc brake image by balancing global and local information.

[0080] The improvement of the indicators shows that the innovative parts of the semantic segmentation algorithm proposed in the present invention are feasible and effective in improving the vibration displacement measurement accuracy of the disc brake.

[0081] Furthermore, in order to test the performance of the method of the present invention in terms of semantic segmentation performance and vibration displacement measurement accuracy, the proposed new semantic segmentation algorithm is compared with the current five representative semantic segmentation algorithms (FCN, DeepLabV3+, PSPNet, HRNet, SegFormer) under the same evaluation system. To ensure the fairness of the comparative test, all algorithms are trained for 300 rounds of iterations on the self-made data set, and the batch size is uniformly set to 4, while other parameters maintain the default configuration.

[0082] The present invention inputs the self-made data set into the various semantic segmentation algorithms mentioned above for training and prediction, and performs time domain and frequency domain analysis on the vibration displacement data of the disc brake, and regresses the vibration displacement curves, axis trajectory diagrams and spectrum diagrams in the x and y directions. The prediction results obtained by various semantic segmentation algorithms are compared with the standard vibration displacement signals collected by the eddy current sensor for qualitative comparison.

[0083] like Figure 7 As shown ( Figure 7 The first to sixth rows are FCN, DeepLabV3+, PSPNet, HRNet, SegFormer, and the present invention, with the left side of each row being the x direction and the right side being the y direction). The vibration displacement curve in the y direction generally shows a better fitting effect than that in the x direction. This is because the rotor is affected by its own gravity when rotating, so that the displacement change in the y direction is greater than that in the x direction, so the semantic segmentation algorithm is relatively easy to capture the position change in the y direction. The overall performance of FCN is the worst. Although it can roughly fit the vibration cycle, there is a lot of noise in its vibration displacement signal, which causes peaks and burrs to appear in the curve, and the peak and trough are significantly offset, which ultimately makes the overall vibration displacement amplitude lower than the standard signal. Although the fitting effect of the vibration displacement curve of SegFormer, PSPNet and DeepLabV3+ networks is better than that of FCN, it is still disturbed by a small amount of noise information, and the measurement accuracy at the displacement peak is low. In addition, from the displacement curve obtained by regression, it can be seen that on the right side of the y-direction displacement curve of SegFormer, the x-direction displacement of PSPNet below the mean value, and the x-direction displacement above and below the mean value of DeepLabV3+ are significantly lower than the actual displacement value, and fail to effectively fit the standard signal. Compared with the previous algorithms, the HRNet algorithm performs well in fitting the vibration period with the real displacement curve, and the overall curve is relatively smooth, but the performance at the peaks and troughs still needs to be improved, and it fails to effectively approach the real amplitude.

[0084] By comparing the vibration displacement curves, it can be found that the vibration displacement signal fitted by the method proposed in the present invention has less interference information and the curve has the best smoothness, successfully overcoming the shortcomings of peaks and troughs commonly found in other comparison algorithms, thereby being as close to the actual displacement data as possible.

[0085] Furthermore, the vibration displacement data in the x and y directions are synthesized to obtain an axis trajectory diagram to represent the actual dynamic path of the disc brake, such as Figure 8 As shown ( Figure 8The first row from left to right is FCN, DeepLabV3+, PSPNet, and the second row from right to right is HRNet, SegFormer, and the present invention). The waveform trajectory regressed from the displacement measured by the eddy current sensor is elliptical. The main reason is that the support stiffness in the x and y directions is asymmetric or the force is unbalanced, which is consistent with the experimental setting. It can be seen that the trajectory fitted by the FCN algorithm is no longer an ellipse, but an irregular and chaotic closed curve. Although the trajectories of other algorithms are approximately elliptical, there is a lot of noise, and it is almost wrapped inside the actual axis trajectory, which is quite different from the standard trajectory. The algorithm proposed in the present invention has the highest degree of fit with the real axis trajectory and the least noise information.

[0086] Furthermore, by using fast Fourier transform to conduct spectrum analysis on the vibration signal of the disc brake, the frequency components such as the fundamental frequency and each harmonic component contained in the signal can be clearly observed. At around 10Hz, the spectrum graph shows a peak, and this peak has the highest intensity in the signal, representing the basic vibration frequency of the rotor, that is, the fundamental frequency, which is directly related to the rotation speed of the rotor. The amplitude of the harmonics is significantly smaller than the fundamental frequency, which indicates that the vibration of the disc brake is mainly concentrated on the rotation frequency and is not seriously disturbed by vibration. However, the disc brake has a slight uneven mass distribution. This unbalanced state will cause additional vibration frequencies, and these harmonic frequencies are usually integer multiples of the fundamental frequency. Based on Fig. 9 The comparison results of the spectrum graph show that (( Fig. 9 The first column is FCN-x, FCN-y, HRNet-x, HRNet-y from top to bottom; the second column is DeepLabV3+-x, DeepLabV3+-y, SegFormer-x, SegFormer-y from top to bottom; the third column is PSPNet-x, PSPNet-y, the present invention-x, the present invention-y from top to bottom). The fundamental frequency of each algorithm can roughly match the standard signal, but there is a significant difference in the degree of fit of the amplitude. The amplitude of the algorithm proposed in the present invention is closer to the standard eddy current signal, which can more accurately reflect the energy or power of the frequency component in the vibration signal. In addition, it can be clearly seen from the spectrum diagram that a large number of irrelevant frequency components appear in the spectrum diagram (especially in the x direction) of each comparison algorithm, while the method proposed in the present invention performs best at each harmonic, and almost no such situation occurs. These erroneous harmonic components will largely cause the center of gravity of the frequency to shift to the right, and the root mean square frequency will be higher than the actual value, which will seriously mislead the observer's accurate judgment of the health of the disc brake.

[0087] Furthermore, based on the evaluation index of the ablation experiment, the present invention adds CFE and FFAE to more comprehensively analyze the time domain and frequency domain performance of the proposed method, thereby quantitatively analyzing different algorithms. The comparison results are shown in Table 2.

[0088] Table 2: Comparison results of different algorithms

[0089]

[0090] In Table 2, from the evaluation indicators of the semantic segmentation network, the indicators mPA and mIOU of the method proposed in the present invention perform best, among which the indicator Recall is as high as 99.65%; from the vibration displacement measurement results, the vibration displacement measurement performance of the algorithm of the present invention in the x and y directions almost crushes other semantic segmentation algorithms with absolute advantages in all indicators in the time domain and frequency domain.

[0091] Specifically:

[0092] FCN has the worst vibration displacement measurement effect, and the error with the standard signal is much higher than other algorithms in all evaluation indicators. This is because FCN uses simple deconvolution for upsampling in the decoder part, which makes the segmentation result fuzzy and smooth, and cannot effectively find the relationship between pixels. In addition, FCN is not good at integrating long-distance dependencies and context information, which also leads to the inability to effectively segment disc brakes;

[0093] SegFormer generally performs poorly in the x and y directions, especially in d-NRMSE, which is as high as 7.14% and 6.78%, CGFE, which is as high as 212.0786695 and 131.6053557, and f-RMSE, which is as high as 250.4417194 and 149.1209699. SegFormer uses the Transformer architecture to capture long-distance dependencies. For very subtle changes, its ability to capture local features is often not as strong as CNN.

[0094] PSPNet also has large differences from the real displacement signal in the time domain and frequency domain, and the d-NRMSE in the x direction is as high as 0.0857. PSPNet uses the pyramid pooling module to obtain global context information of different scales, which is helpful for understanding large areas, but the pooling operation itself will lose some spatial information and often ignore key local features for subtle changes;

[0095] DeepLabV3+ has a high FFAE of 0.17592 in the x direction. Although DeepLabV3 can extract multi-scale information through dilated convolution, the multi-scale fusion of feature maps in the decoder is not sufficient, resulting in low segmentation precision of the final semantic segmentation map. The vibration displacement of disc brakes is usually small, and this network design is prone to blurred edges, affecting the accuracy of displacement measurement.

[0096] The method of the present invention effectively overcomes the shortcomings of the above-mentioned comparative algorithms and achieves the most outstanding results in vibration displacement measurement accuracy. This is due to the downsampling method that effectively retains high-resolution features, the excellent global-local feature information coordination ability, and the multi-scale fusion strategy rich in edge feature information guided by the Laplace residual, which enables the method of the present invention to effectively overcome the shortcomings of the above-mentioned comparative algorithms and achieve the most significant results in vibration displacement measurement accuracy.

[0097] In summary, the proposed method not only has excellent network performance, but also performs outstandingly in visual vibration measurement tasks. By comparing with the most representative semantic segmentation algorithm currently, the method of the present invention shows obvious advantages in multiple key evaluation indicators, which proves its ability to handle subtle changes under dynamic conditions and also provides a solid data foundation for subsequent fault diagnosis and maintenance decisions.

[0098] According to a second aspect of an embodiment of the present invention, a processor is provided, wherein the processor is used to execute an operation, wherein the operation includes executing any one of the semantic segmentation methods for disc brake vibration displacement measurement described above.

[0099] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented: obtaining disc brake vibration video data; and obtaining a corresponding segmentation mask for each frame of the video after prediction by a semantic segmentation network model.

[0100] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A semantic segmentation method for disc brake vibration displacement measurement, characterized in that: The following steps are involved: Obtain disc brake vibration video data; Each frame of the video is predicted by the semantic segmentation network model to obtain the corresponding segmentation mask; the semantic segmentation network model includes an encoder and a decoder. The encoder uses Resnet101 as the network backbone, and replaces the 7×7 convolution and maximum pooling with a downsampling module based on Harr wavelet transform, so that the output feature map has more detailed information of the disc brake in the image frame, and inserts a global-local feature extraction module into the original residual block of the Resnet101 network to enhance the semantic segmentation network's ability to represent the disc brake details and local structure information. The encoder obtains feature layers f2 and f3 , f4 and f5; in the decoder part, the intra-scale feature interaction module based on collaborative attention and the cross-scale feature fusion module guided by the Laplacian image are adopted, and the feature layer f5 is input into the intra-scale feature interaction module based on collaborative attention. The spatial semantic information of the disc brake image is extracted from the feature layer through the intra-scale feature interaction module based on collaborative attention and the spatial semantic difference is reduced to obtain the feature layer P5. The original input image f1, feature layer f2, feature layer f3, feature layer f4, feature layer f5 and feature layer P5 are input into the cross-scale feature fusion module guided by the Laplacian image to obtain the corresponding segmentation mask.

2. The semantic segmentation method for disc brake vibration displacement measurement according to claim 1, characterized in that: The encoder specifically comprises the following steps: first, the input image f1 is processed twice by a downsampling module based on Haar wavelet, and a feature layer f2 is obtained after the first downsampling module based on Haar wavelet is processed; the feature layer f2 is used as the input of the second downsampling module based on Haar wavelet, and the output of the second downsampling module based on Haar wavelet is sequentially passed through three residual blocks inserted into the global-local feature extraction module, and the output results of the three residual blocks inserted into the global-local feature extraction module are feature layers f3, f4 and f5 respectively.

3. The semantic segmentation method for disc brake vibration displacement measurement according to claim 2, characterized in that: The specific process of the downsampling module based on Haar wavelet is as follows: first, the resolution of the disc brake image is halved; then its signal is decomposed into two parts: a low-frequency L component and high-frequency HH, HV and HD components, and they are spliced ​​in the channel dimension; finally, the feature map is obtained through convolution, normalization and Relu activation function processing.

4. The semantic segmentation method for disc brake vibration displacement measurement according to claim 3, characterized in that: The normalization option uses Group Normalization.

5. The semantic segmentation method for disc brake vibration displacement measurement according to claim 2, characterized in that: The residual block inserted with the global-local feature extraction module is specifically as follows: a global-local feature extraction module is added after the 1×1 and 3×3 convolutions of the original residual block part, and the global-local feature extraction module performs local adaptive average pooling on the input image to obtain local spatial information; then it is divided into two branches, the first branch performs global adaptive average pooling on the obtained local features, performs a convolution operation when the spatial dimension is reduced to 1, and then expands to the original input size of the module; the second branch performs shape adjustment on the obtained local features, extracts deeper features through one-dimensional convolution, and then returns to the original input size of the module through shape adjustment; the features obtained from the two branches are weighted fused, and then restored to the same size as the input features through adaptive anti-average pooling to obtain an attention map; finally, the input features are multiplied by the generated attention map to obtain the output of the global-local feature extraction module.

6. The semantic segmentation method for disc brake vibration displacement measurement according to claim 1, characterized in that: The intra-scale feature interaction module based on collaborative attention is specifically as follows: the input feature layer f5 is average pooled, and the average values ​​are taken along the height and width dimensions respectively, and the feature maps of the two dimensions are obtained, and each of them is evenly divided into 4 groups of non-overlapping sub-features along the corresponding dimensions; then, a separable one-dimensional convolution with multi-receptive field depth sharing is applied to these sub-features, and they are respectively spliced, normalized and activated, and then re-adjusted back to the original dimensions to form a spatial attention map in two dimensions; and then the obtained result is element-wise multiplied with the initial input feature layer f5 of the intra-scale feature interaction module based on collaborative attention to obtain a fusion. The weighted features that combine global context dependencies and multi-semantic spatial information are average pooled and normalized, and the dimensions are transformed through two-dimensional convolution to generate query q, key k and value v; then, the self-attention score between k and q is calculated along the channel dimension using the dot product method, and the channel attention relationship weight is obtained through softmax and dot product calculation is performed with v; finally, the output dimension of the self-attention is restored to the input shape, and after average pooling and activation function, the output feature is multiplied with the input "weighted features that combine global context dependencies and multi-semantic spatial information" to obtain the feature layer P5.

7. The semantic segmentation method for disc brake vibration displacement measurement according to claim 1, characterized in that: The Laplacian image-guided cross-scale feature fusion module is specifically: The feature layer f j+1 Perform 2 times upsampling and then combine with the feature layer f j Perform subtraction operations to obtain Laplacian images L of various scales j ; where j = 1, 2, 3, 4; The feature layer P5 is fused with the feature layer f4 in the first fusion module to obtain the output feature layer x4; the feature layer P5 is subjected to five first convolution modules to obtain the feature layer R5; The feature layer R5 is upsampled twice and then stacked with x4 and L4 to obtain the feature layer P4; the feature layer P4 undergoes four first convolution modules and is added with L4 to obtain R4; after P4 undergoes one first convolution module, it passes through the first fusion module once and is fused with the feature layer f3 to obtain the output feature layer x3; The feature layer R4 is upsampled twice and then stacked with x3 and L3 to obtain the feature layer P3; the feature layer P3 passes through the first convolution module three times and is added with L3 to obtain R3; P3 passes through the first convolution module once and then passes through the first fusion module once to obtain the output feature layer x2 by fusing with the feature layer f2; The feature layer R3 is upsampled twice and then stacked with x2 and L2 to obtain the feature layer P2; the feature layer P2 passes through the first convolution module twice and is added with L2 to obtain R2; P2 passes through the first convolution module once and then passes through the first fusion module once to obtain the feature layer x1 by fusing with the feature layer f1; The feature layer R2 is upsampled twice and then stacked with x1 and L1 to obtain the feature layer P1; the feature layer P1 passes through three convolution modules and is added with L1 to obtain R1; R5 is upsampled twice and added to R4 to obtain D4; D4 is upsampled twice and added to R3 to obtain D3; D3 is upsampled twice and added to R2 to obtain D2; D2 is upsampled twice and added to R1 to obtain D1, which is the final output result of the semantic segmentation network.

8. A processor, characterized in that: The processor is used to execute operations, and the operations include executing the semantic segmentation method for disc brake vibration displacement measurement according to any one of claims 1-7.

Citation Information

Patent Citations

  • Semantic segmentation method and visual positioning method for vibration image

    CN116229468A

  • Global and local feature reconstruction network-based medical image segmentation method

    US20230274531A1