Train bogie bolt intelligent detection system and method based on machine vision
Through the improved YOLOv8 network and position tracking system, combined with ODConv and CARAFE modules, the efficiency and accuracy of train bogie bolt detection are solved, and efficient and accurate bolt status monitoring is achieved.
Patent Information
- Application Number
- CN202510333578.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to efficiently and accurately detect the loose state of train bogies, especially in complex backgrounds and high speed environments, with high model resource requirements and low detection efficiency.
Using a bolt detection method based on machine vision, combined with the improved YOLOv8 network, an ODConv dynamic convolution network and a CARAFE module are introduced, and a position tracking system is equipped with a real-time detection through high-precision sensors and image processing technology.
It improves the accuracy and efficiency of bolt detection, reduces false detection and missed detection rates, optimizes the use of computing resources, and enhances the adaptability and generalization capabilities of the model.
Smart Images

Figure CN120259235A_ABST
Abstract
Description
Technical Field
[0001] For the bolt loosening target detection technology based on machine vision Background Art
[0002] China's transportation industry is in a period of rapid development. Whether the transportation is developed affects the economic development to a certain extent. Train tools play a crucial role in the process of transportation development. As an important tool for railway transportation, high-speed trains must first ensure their safe operation. As one of the most important fasteners on railways, bolts play a fixing role on workpieces through clamping force, and the forces borne by trains can easily cause bolts to fail. Therefore, fault detection of train bolts is an important part of ensuring train operation safety
[0003] For the visual target detection of small bolts, due to their small size, similar shapes, and complex backgrounds, a series of challenges will be faced when detecting bolts. For example, in terms of the images obtained of bolts, small bolts often occupy very few pixels, resulting in low contrast and resolution, making it difficult for the model to accurately identify and locate them; they are often densely arranged on equipment or components, and the mutual occlusion and overlap will increase the difficulty of image detection; complex or cluttered image backgrounds may also interfere with the recognition rate of detection. For example, in terms of the model for target detection, in order to improve the detection accuracy of small bolts, the model structure often requires higher complexity, leading to overfitting; high-complexity models require more computing resources and time, and shorter time is often pursued in real-time detection. Therefore, how to optimize resources and time has become an urgent problem to be solved Summary of the Invention
[0004] In order to solve the above problems, the present invention provides a bolt detection method based on machine vision, with high detection efficiency and accurate detection results. The specific technical solutions are as follows
[0005] The bolt defect detection method based on machine vision includes the following steps
[0006] S1. Monitor the approaching train in real time, trigger the image data acquisition module, and collect image data of the train number
[0007] S2. Track the position of the train bogie and number the bogie
[0008] S3. Locate the bolts on the bogie and collect image data of the bolts
[0009] S4. Process the collected image data and transmit it to the improved YOLOv8
[0010] S5. Introduce the ODConv dynamic convolution network on the basis of YOLOv8 to replace the traditional convolution network in the Neck layer, and modify the upsampling operator of the Neck layer to the CARAFE module;
[0011] Furthermore, the specific steps of step S1 include:
[0012] S11. When the system trigger module detects the approaching of a train, start image data acquisition, and set different light sources for different environments
[0013] S12. Collect data on the train number
[0014] Furthermore, in step S11,
[0015] The sensor needs to collect environmental data (such as temperature, humidity, etc.) and train movement parameters (such as speed, acceleration, distance, etc.) in real time. The data acquisition module converts these signals into digital signals and transmits them to the trigger control module. Detect the arrival of the train and trigger image acquisition.
[0016] Common sensors are used, including infrared sensors, radars, laser rangefinders, etc. These devices are installed along the railway or on the platform to detect the arrival of trains. Infrared sensors emit and receive infrared signals to detect whether an object (such as a train) passes by. Radar radio waves are used to measure the distance and speed of the target, and then the distance between the train and the monitoring point is measured by laser pulses. When the sensor detects the approaching of the train, the trigger control module makes a logical judgment based on preset conditions (such as distance threshold, speed range). If the trigger condition is met (for example: the train is within 50 meters of the monitoring point), the system will send a trigger signal to the image acquisition module. The trigger signal can be a digital signal or an analog signal, which is used to control the startup of the image acquisition module. This process requires quick response to ensure accurate capture of the train number information when the train passes by.
[0017] The system trigger module is a highly integrated component that covers the complete process from data acquisition to condition judgment and then to execution control. Its structure includes multiple functional modules such as a monitoring unit, a trigger condition judgment logic, an execution unit, etc. Each part has its specific role and the cooperation relationship between them.
[0018] Set parameters such as sampling frequency, signal conditioning (such as amplification, filtering), etc. to adapt to the sensor output characteristics and subsequent data analysis requirements. Install the data acquisition equipment on the train or ground control center and ensure a stable connection with the sensor. Use wired communication (such as Ethernet, optical fiber) or wireless communication technology (such as 5G, Wi-Fi, ZigBee, etc.) to select the appropriate method according to the application scenario. Deploy communication equipment between the train and the monitoring center to ensure real-time data transmission. Use a database (such as MySQL, MongoDB) or cloud storage system to classify and store the collected data for subsequent query and analysis.
[0019] Furthermore, in step S12,
[0020] Select the appropriate industrial camera or ordinary camera equipment according to the needs. Industrial cameras have high frame rate and high resolution, which are suitable for shooting high-speed moving objects. The camera is installed in a suitable position (such as on a fixed bracket beside the railroad track) to ensure that the train number information can be clearly captured. Set parameters such as exposure time, white balance, and focus to adapt to ambient light and background conditions. Adjust the frame rate according to the train speed to ensure that multiple images can be captured in a short time. After receiving the trigger signal, the camera starts the photo function to obtain the image data at the current moment. If necessary, you can set a continuous shooting mode (such as taking a picture every 0.1 seconds) to record multiple states during the train passing. Perform pre-processing operations such as denoising and enhancement on the collected images to improve image quality. Extract text and pattern features related to the train number information through binarization or edge detection algorithms. Use optical character recognition (OCR) technology to identify text information such as vehicle number and vehicle model in the image. Store the collected image data and recognition results in a local or cloud database for subsequent query and analysis. In addition, the data management system supports efficient retrieval functions, such as quickly finding relevant records by vehicle number, time range, etc.
[0021] The method of real-time monitoring of trains mainly relies on advanced sensor technology, communication technology and data analysis methods. By deploying a variety of sensors on the train, key parameters are collected in real time; data acquisition systems and communication technologies are used to achieve data transmission and storage; machine learning algorithms and visualization tools are used to analyze and display data, and provide feedback and intervention. This method can effectively improve the safety and reliability of train operation and provide important technical support for the rail transit industry.
[0022] Furthermore, the step S2 specifically includes:
[0023] S21. Equipped with a position tracking system to locate the bogie in real time
[0024] S22. Number the positioned bogies
[0025] Further, in step S21,
[0026] The purpose of the bogie position tracking system is to obtain the position information (such as coordinates) of the bogie in real time through various sensors and technical means, and combine the vehicle operation status and environmental data to achieve high-precision positioning of the bogie position. This system can be applied to the safety monitoring, fault diagnosis, and operation management of rail transit vehicles.
[0027] The bogie position tracking system mainly consists of the following parts:
[0028] Sensor module: Used to collect data such as the position, attitude, acceleration, and vibration of the bogie. Data acquisition and transmission module: Responsible for converting the sensor signals into digital signals and transmitting them to the main control unit through the communication network. Positioning algorithm module: Based on the sensor data, combined with the kinematic model and positioning technology (such as inertial navigation, GPS, etc.), to calculate the position of the bogie. Data processing and analysis module: Store, analyze, and visually display the positioning results.
[0029] To achieve high-precision position tracking, a combination of multiple sensors needs to be selected:
[0030] (1) Accelerometer
[0031] The accelerometer is used to measure the linear acceleration of the bogie, and its output signal can reflect the motion state of the vehicle. Assuming the measured value of the accelerometer is a x , a y , a z , then the speed and position of the bogie can be obtained through integration:
[0032]
[0033] (2) Gyroscope
[0034] The gyroscope is used to measure the angular velocity of the bogie, and its output signal is ω x , ω y , ω z . The attitude angle of the bogie can be obtained by integrating the angular velocity:
[0035]
[0036] (3) Global Positioning System (GPS)
[0037] GPS is used to provide the absolute position information of the bogie, and its output is the longitude and latitude coordinates (x, y) and altitude z.
[0038] (4) Light Detection and Ranging (LiDAR)
[0039] LiDAR can be used for high-precision positioning by scanning the environment and matching the map to achieve positioning. Assume that the measurement value of LiDAR is point cloud data Then the position and attitude of the bogie can be calculated through feature matching algorithms.
[0040] Data acquisition and preprocessing: Data acquisition: Sensor data needs to be sampled by an analog-to-digital converter (ADC) and transmitted to the main control unit in the form of digital signals. Assume the sampling frequency is f s , then N = f samples can be collected per second s Data preprocessing: Noise reduction: Use a low-pass filter to eliminate high-frequency noise. For example, use a Butterworth filter:
[0041]
[0042] where ω c is the cut-off frequency and n is the filter order.
[0043] Signal fusion: Align the signals from different sensors in time and adjust the scale. The positioning algorithm is the core of the system and requires the combination of multiple technologies to achieve high-precision positioning.
[0044] (1) Inertial Navigation System (INS)
[0045] Based on the data from accelerometers and gyroscopes, calculate the position and attitude of the bogie through integration:
[0046]
[0047] (2) GPS-aided positioning
[0048] The inertial navigation result data
[0049] x k+1 = Ax k + Bu k
[0050] Observation equation:
[0051] y k = Cx k + Dv k
[0052] (3) Vision-based positioning
[0053] Through cameras and image processing techniques, identify the environmental features around the bogie and combine with the SLAM (Simultaneous Localization and Mapping) algorithm to achieve positioning.
[0054] The overall architecture of the bogie position tracking system is as follows: Sensor module: accelerometers, gyroscopes, GPS, lidar, etc. Data acquisition and transmission module: analog-to-digital converter (ADC), communication networks (such as CAN bus, wireless communication). Positioning algorithm module: inertial navigation, Kalman filtering, SLAM, etc. Data processing and analysis module: data storage, visualization interface, alarm system.
[0055] Marking points are arranged on the track, and the position of the bogie is measured by GPS and lidar and compared with the actual position. A vibration table is used to simulate the vehicle running state to test the measurement accuracy of accelerometers and gyroscopes.
[0056] Furthermore, in step S22,
[0057] The geographical location information of the bogie is collected in real time using sensors or readers / writers. The position data is transmitted to the central control system or database through wired or wireless networks to ensure the real-time and accuracy of the data. A numbering rule is designed according to actual needs and encoded in sequence. Ensure the uniqueness of the numbering to avoid duplication or confusion. The collected real-time position data is processed in the central control system. The processed data is stored in the database and associated with the bogie number information to establish a one-to-one correspondence. The position and number information of each bogie are displayed in real time on the user interface for easy management and monitoring. A search function is provided to allow users to quickly find the position information of a specific bogie according to the number. Regularly check the stability of the positioning system to ensure the normal operation of sensors or readers / writers. Regularly back up and clean the database to prevent data redundancy or damage.
[0058] Furthermore, step S3 specifically includes:
[0059] S31. By extracting the unique features of the track bolts (such as shape, texture, color, etc.), bolts can be effectively distinguished from other objects, and the positioning algorithm is used for accurate positioning.
[0060] S32. The image acquisition device includes industrial cameras (such as CCD cameras or CMOS cameras) and cameras (the light source also needs to be considered)
[0061] Furthermore, in step S31,
[0062] Arrange high-precision sensors for the tram in the running state, such as laser displacement sensors, image acquisition devices or ultrasonic detection devices. Use the Canny operator or Sobel operator to detect the edges in the image and capture the contour of the bolt. Adopt sliding window detection and positioning. Traverse the image through the sliding window, and use feature matching or classifiers (such as SVM, AdaBoost, etc.) to judge whether each window contains a bolt. Scan each key bolt on the track one by one, and record information such as the position coordinates, tilt angle and fastening status of each bolt. Since the train is running, factors such as vibration and temperature that affect the data also need to be considered. Denoise and filter the collected raw data to eliminate environmental interference and sensor noise. Establish a standard model of the bolt position and compare the actual measurement values with the standard values.
[0063] Further, in step S32,
[0064] Select a suitable industrial camera or ordinary imaging device according to requirements. Industrial cameras have high frame rates and high resolutions, and are suitable for photographing high-speed moving objects. The camera is installed at a suitable position (such as on a fixed bracket beside the railway track) to ensure that the train number information can be clearly captured. Set parameters such as exposure time, white balance, and focus to adapt to the ambient light and background conditions.
[0065] Further, step S4 specifically includes:
[0066] S41. Introduce the ODConv dynamic convolution network to replace the traditional convolution network in the Neck layer, and perform image enhancement on the collected sample images
[0067] S42. After the output of the Darknet53 network, add an SPP network structure. Perform max-pooling operations on the feature map through three max-pooling layers of different scales, and fuse the multi-scale output local feature maps of the track bolts with the input global feature maps of the track bolts;
[0068] Further, in step S41,
[0069] Adjust the brightness of the track bolt sample pictures. The adjustment method is: f(x) = g(x) + β, where f(x) is the adjusted track bolt image, g(x) is the original track bolt image, and β is the brightness adjustment coefficient. Perform 90° and 180° clockwise and counterclockwise flips on the brightness-adjusted track bolt samples to achieve sample expansion. Blur the expanded images, and use Gaussian blur for the sample data. The mathematics of Gaussian blur:
[0070] Expressed as:
[0071] Among them, σ is the standard deviation, r is the blur radius, and N is the normalization constant of the Gaussian distribution, which ensures that the sum of the Gaussian functions is 1, so that the brightness of the image remains unchanged and no information is lost due to the blur processing.
[0072] ODConv full-dimensional dynamic convolution realizes dynamic adjustment of the shape and size of the convolution kernel according to different input data features by introducing a learnable deformation module. This method processes from the spatial dimension, input channel dimension, and output channel dimension, and can also pay attention to the shape and size of the convolution kernel to better adapt to different input data. The main formula of ODConv full-dimensional dynamic convolution is as follows:
[0073] y = (α W1 W1 +... + α Wn W n ) * x
[0074] In the formula, W i is the i-th convolution kernel, and α wi is the attention scalar for weighting W i . * represents the convolution operation, and ⊙ represents the multiplication operation in different dimensions. ODConv will perform a multiplication operation on the attention scalar and the convolution kernel to adjust the size and shape of the convolution kernel, and then perform a convolution operation with the input feature x to obtain the final output feature y. First, use the global average pooling layer GAP (global average pooling) to compress the target input feature, then reduce the complexity of the dynamic convolution through the fully connected layer FC layer (fully connected layers), and then use the ReLU function to remove the negative values of the feature vector to generate 4 head branches to construct 4 types of attention scalars of the ODConv convolution module. Finally, use the Sigmoid activation function and Softmax function to normalize the processed features to generate the normalized attention scalars α si , α ci , α fi , α wi . Each attention scalar is multiplied and summed with the original convolution kernel W i , and then combined with the input feature to obtain the output feature.
[0075] The formula of the Softmax function is as follows:
[0076]
[0077] Using the Softmax function to convert each value in the input feature x i into a numerical value between 0 and 1, and the sum of all output values is 1. In image detection, it can distinguish and identify multiple objects or object categories existing in the image, enhance the interpretability of model prediction, and simplify the loss calculation and model training process.
[0078] The formula of the ReLU activation function is as follows:
[0079] ReLU(x) = max(0, x)
[0080] The ReLU activation function enhances the nonlinear characteristics of the model by setting the negative value part to 0, which helps the model capture complex features in the image, thereby improving its expression ability. At the same time, it adds 0 element outputs in the output to represent the sparsity of features, so it can help the model ignore irrelevant background information in the image and focus on meaningful target features.
[0081] The formula of the Sigmoid function in the improved structure is as follows:
[0082]
[0083] The Sigmoid function is used to generate attention weights on each dimension, and these weights reflect the importance of different positions and channels, so as to be able to process the input features more precisely. The Sigmoid function maps each generated attention scalar to between 0 and 1, which can effectively integrate different feature information and improve the overall performance of the model.
[0084] Furthermore, in step S42,
[0085] For the extraction of deeper semantic information of the detection features of the track bolt and the need to realize multi-scale feature prediction, the Darknet53 network layer is improved. Since there is a large size difference between the track nut and the overall track bolt in the input image of the track bolt, after the output of the Darknet53 network, referring to the method of Spatital Pyramid Pooling (SPP), an SPP network structure is added to strengthen the network feature performance.
[0086] The SPPF model has a smaller computational amount and a faster operation speed, and can solve the problem of inconsistent input image sizes. The SPPF module extracts features at different scales through max-pooling operations. Specifically, it has four parallel branches: The first branch: directly passes the input features to the output without any operation. The second branch: performs a max-pooling operation with a kernel size of 5×5 to obtain medium-scale features through downsampling.
[0087] The third branch: uses a max-pooling operation with a kernel size of 9×9 to further extract larger-scale features.
[0088] The fourth branch: adopts a max-pooling operation with a kernel size of 13×13 to extract even larger-scale features.
[0089] The stride of all pooling operations is 1, which means that padding is required during the convolution operation to keep the size of the feature map unchanged. Finally, the feature maps of these four branches are concatenated to complete the fusion of features at different scales. This structure realizes efficient multi-scale information integration and helps to capture objects of different sizes in the object detection task. Max-pooling operations are performed on the feature map through three max-pooling layers of different scales to achieve the fusion of the local features and global features of the input track bolt feature map. Fusing the multi-scale output local feature map of the track bolt with the input global feature map of the track bolt enriches the expressive power of the feature map.
[0090] Further, the step S5 specifically includes:
[0091] S51. Modify the upsampling operator in the Neck layer to the CARAFE module;
[0092] S52. At the same time, add the CBAM attention mechanism
[0093] Further, in the step S51.
[0094] The core module of this structure includes a kernel prediction module (KernelPredictionModule) and a content-aware reassembly module (Content-awareReassemblyModule). The kernel prediction module consists of three parts: a channel compressor (ChannelCompressor), a content encoder (ContentEncoder), and a kernel normalizer (Kernel Normalizer); the content-aware reassembly module is composed of upsampling kernel normalization, feature mapping back to the input feature map, and output value calculation.
[0095] First, the channel compressor uses a 1*1 convolution to compress the number of channels of the input feature map to reduce the computational complexity of subsequent steps. The feature channel compression formula is as follows:
[0096] y = Conv1*1(x)
[0097] where x and y are the input and compressed feature maps respectively, and Conv 1*1 represents applying a 1*1 convolution operation. Then, the content encoder encodes the compressed feature map so that each upsampling kernel can represent the content features at its corresponding position. It usually contains a convolution kernel of size k up ×k up (the size of the upsampling kernel). The content encoder formula is as follows:
[0098] y = Convk encoder *k encoder (x)
[0099] where x and y are the feature maps of the input and output respectively, and the upsampling kernel size is k up ×k up Predict the upsampling kernel through a convolutional layer of k encoder *k encoder The number of output channels is Unfold the channel dimension in the spatial dimension to obtain a shape of upsampling kernel.
[0100] Secondly, the kernel normalizer performs Softmax normalization on the generated upsampling kernel to ensure numerical stability and rationality during the upsampling process. Finally, the feature recombination module uses the obtained upsampling kernel to recombine the content of the input feature map, realizing instance-specific content-aware processing and dynamically generating adaptive upsampling weights for each position.
[0101] Furthermore, in step S52.,
[0102] CBAM is a module that integrates channel and spatial attention mechanisms, enhancing the network's ability to capture important features while suppressing unnecessary information. The CBAM module, as Figure 5 shown, fine-tunes features in two dimensions by combining the channel attention module CAM and the spatial attention module SAM. The channel attention module and the spatial attention module respectively emphasize the importance of each pixel in different channel features and in spatial features.
[0103] The input features enter CAM and SAM successively, and a global average pooling (AvgPool) and a global max pooling (MaxPool) are performed in parallel in these two modules. Among them, the global average pooling is used to obtain the average value of each channel / spatial dimension, reflecting the global information of the feature map; the global max pooling captures the maximum value of each channel / spatial dimension, highlighting the local maximum amplitude in the feature map. The combination of the two can take into account both the average information and the maximum amplitude information of the feature map.
[0104] After passing through the global average pooling and the global max pooling in the channel attention module CAM, a multi-layer perceptron (MLP) composed of two fully connected layers is used to generate the final channel attention weights. The attention weights of each channel are obtained through the Sigmoid activation function and normalized to between 0 and 1 so that they can be used as scaling coefficients for the input feature map. Finally, the obtained attention weights are multiplied by the input feature map to obtain M c , and the calculation formula of the channel attention module is as follows:
[0105]
[0106] In the Spatial Attention Module (SAM), after global average pooling and global maximum pooling, the results are reduced to single-channel features through a convolution with a kernel size of 7*7, and the final feature map M is obtained through processing with the Sigmoid activation function. s , and the calculation formula of the spatial attention module is as follows:
[0107]
[0108] Compared with the prior art, one or more of the above technical solutions can achieve at least one of the following beneficial effects:
[0109] The present invention locates the area of the track bolts, speeds up the positioning speed, improves the detection efficiency, and reduces the time.
[0110] The present invention creates a storage and transmission system, which can perform advanced retrieval of relevant information such as train numbers, bogies, bolts, etc. in the system and can transmit the information quickly.
[0111] The present invention creates a position tracking system to track the train bogie in real time, avoid confusing bogies during the detection process, increase the accuracy rate, and speed up the positioning speed of the bolts.
[0112] In the bolt detection of the present invention, the accuracy and mAP value are effectively improved. This module can enhance the feature expression ability, adaptability, and generalization ability, reduce the false detection and missed detection rates, optimize the use of computing resources, accelerate the training convergence, and improve the multi-scale detection ability. Description of the Drawings
[0113] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0114] Figure 1 Intelligent Detection System for Train Bogie Bolts Based on Machine Vision
[0115] Figure 2 ODConv Full-Dimensional Dynamic Network Structure
[0116] Figure 3 Improved Structure Diagram of YOLOv8
[0117] Figure 4 Intelligent Detection Flow Chart of Train Bogie Bolts
[0118] Figure 5 Module Diagram of Intelligent Detection System for Train Bogie Bolts Specific Embodiments
[0119] The following provides a detailed description of the specific embodiments of the present invention. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0120] Embodiment 1
[0121] This paper presents an intelligent detection system for train bogie bolts based on machine vision: The system includes: a system trigger module that triggers the image data acquisition module by detecting whether a train is approaching; a position tracking module that monitors each bogie of the train in real time when the train approaches; a visual positioning module that performs visual positioning on the bolts of each bogie of the train; an image data acquisition module that acquires the image data of the positioned bogie bolts through a camera; a data storage and transmission module that processes and stores the acquired image data and transmits it to YOLOv8; a YOLOv8 improvement module that introduces an ODConv dynamic convolution network to replace the traditional convolution network in the Neck layer and modifies the upsampling operator in the Neck layer to a CARAFE module.
[0122] The bogie position tracking system mainly consists of the following parts: a sensor module for collecting data such as the position, attitude, acceleration, and vibration of the bogie; a data acquisition and transmission module responsible for converting sensor signals into digital signals and transmitting them to the main control unit through a communication network; a positioning algorithm module that calculates the position of the bogie based on sensor data in combination with kinematic models and positioning technologies (such as inertial navigation, GPS, etc.); a data processing and analysis module that stores, analyzes, and visually displays the positioning results.
[0123] The bolt positioning implementation module usually consists of the following parts: data acquisition equipment for obtaining information such as the position and fastening status of each bolt on the bogie, and common equipment includes laser measuring instruments, image sensors, torque sensors, etc.; a positioning algorithm that analyzes the collected data through mathematical models and algorithms to determine the specific position of the bolt; a positioning accuracy evaluation that verifies the accuracy of the positioning result to ensure that the positioning error is within the allowable range; data storage and management that stores the positioning data in a database and provides support for subsequent data processing.
[0124] The specific implementation steps are as follows:
[0125] S1. Monitor the approaching train in real time, trigger the image data acquisition module, and acquire the image data of the train number.
[0126] S2. Track the position of the train bogie and number the bogies.
[0127] S3. Locate the bolts on the bogie and collect image data of the bolts.
[0128] S4. Process the collected image data and transmit it to the improved YOLOv8.
[0129] S5. Introduce the ODConv dynamic convolution network on the basis of YOLOv8 to replace the traditional convolution network in the Neck layer, and modify the upsampling operator of the Neck layer to the CARAFE module;
[0130] S11. Determine the train number and its location to be photographed, and ensure that equipment is arranged at appropriate locations (such as platforms, beside the tracks or specific detection points). When the train is detected approaching, transmit the signal to the system trigger module to start collecting image data
[0131] S12. Collect data on the train number
[0132] After the train enters the specified area, the triggering device starts, and continuous shooting or multi-angle shooting begins. The shooting frequency needs to be adjusted according to the train speed and detection requirements to ensure that key parts (such as bogies, bolts, etc.) are clearly captured.
[0133] Store the photos in the local server or cloud database, and add information such as time stamps and train number to each photo. Establish an image index file for easy subsequent retrieval and analysis.
[0134] S21. Install a position tracking system to locate the bogie in real time
[0135] Install positioning devices (such as RFID tags, QR code signs or other sensors) on the train or the track to identify the position information of the bogie. Deploy tracking devices (such as cameras, laser scanners or infrared sensors) to capture the position changes of the bogie in real time. Use computer vision technology (such as based on YOLO or other object detection models) to achieve automatic identification and positioning of the bogie.
[0136] S22. Number the located bogies
[0137] Annotate the bogie in the image and record its position information through a coordinate system or other coding methods. According to the train number and bogie numbering rules, assign a unique identification code to each bogie. Collect the position data of the bogie in real time (such as x, y coordinates), and analyze it in combination with the time series. Associate and store the position information with the photo data for easy subsequent analysis of the state changes of the bogie.
[0138] S31. By extracting the unique features of the track bolts (such as shape, texture, color, etc.), the bolts can be effectively distinguished from other objects, and positioning algorithms are used to accurately locate them.
[0139] Rail bolts usually have specific geometries, such as hexagonal or round heads, which can be extracted by image processing techniques. The bolt surface may have specific texture features (such as scratches, oxide layers, etc.), which can help distinguish the bolt from other objects. The color of rail bolts is usually relatively uniform (such as silver, gray, etc.), which is different from the surrounding environment or background. The positioning algorithms include traditional methods based on image processing (such as edge detection, template matching) and deep learning-based methods (such as object detection models using convolutional neural networks (CNNs)).
[0140] S32. The image acquisition device includes industrial cameras (such as CCD cameras or CMOS cameras) and cameras (the light source also needs to be considered).
[0141] Take pictures using an industrial camera or a high-resolution camera. Deploy fill lights or other light sources (if necessary) to ensure stable photo quality under different lighting conditions. Install a triggering device (such as a photoelectric sensor or a timer) to automatically start the photo-taking function.
[0142] S41 After image enhancement of the collected sample images, including operations such as randomly rotating the picture by 0 to 360 degrees and / or 5 to 10 random rotation times, etc.;
[0143] During the training process, randomly select an angle (ranging from 0° to 360°) for each image and rotate it clockwise or counterclockwise. Repeat the rotation of the same image 5 to 10 times (the angle of each rotation is randomly selected).
[0144] This operation can further increase the diversity of the data, enabling the model to more comprehensively learn the local characteristics of rail bolts that are not S42. After the output of the Darknet53 network, add an SPP network structure, perform max-pooling operations on the feature maps through three max-pooling layers of different scales, and fuse the multi-scale output local feature maps of rail bolts with the input global feature maps of rail bolts;
[0145] By introducing the SPP (Spatial Pyramid Pooling) network structure, further extract multi-scale features and enhance the model's detection ability for rail bolts. Insert an SPP module after the Darknet53 network. On the basis of these feature maps, add an SPP module. SPP extracts features of different scales by applying max-pooling operations (Max-Pooling) on each feature map. Fuse the extracted local features (multi-scale pyramid features) with the input global feature maps of rail bolts in terms of channels or dimensions.
[0146] Improve based on YOLOv8, and introduce the ODConv (Omni-Dimensional Dynamic Convolution) dynamic convolution network to replace the traditional convolution network in the Neck layer.
[0147] ODConv full-dimensional dynamic convolution is a dynamic convolution that uses a multi-dimensional attention mechanism to learn four complementary attention types of convolution kernels along four dimensions of the kernel space in parallel. For the dynamic convolution layer, it uses a linear combination of n convolution kernels, dynamically weighted by the attention mechanism, making the convolution operation dependent on the input
[23] . Compared with conventional convolution, ODConv full-dimensional dynamic convolution introduces a learnable deformation module to dynamically adjust the shape and size of the convolution kernel according to different input data features. This method processes from the spatial dimension, input channel dimension, and output channel dimension, and can also pay attention to the shape and size of the convolution kernel to better adapt to different input data.
[0148] S52. Replace some Conv convolution structures in YOLOv8 with ODConv full-dimensional dynamic convolution structures;
[0149] In this paper, the Upsample operator in the YOLOv8 structure is replaced with the CAEAFE upsampling module. Compared with the Upsample operator, CARAFE can dynamically adjust the kernel weights according to the specific content of the input features by using an adaptive upsampling kernel, enabling the model to have less computational cost and higher upsampling efficiency when processing high-resolution images, supporting content-aware processing, and predicting the upsampling kernel according to the content of each target location.
[0150] S53. Modify the upsampling operator in the Neck layer to the CARAFE (Content-Aware ReAssembly of Features) module to better retain details by aggregating context content; at the same time, add the CBAM (Convolutional Block Attention Module) attention mechanism.
[0151] CBAM is a module that integrates channel and spatial attention mechanisms, enhancing the network's ability to capture important features while suppressing unnecessary information. The CBAM module fine-tunes features in two dimensions by combining the channel attention module CAM and the spatial attention module SAM. The channel attention module and the spatial attention module respectively emphasize the importance of each pixel in different channel features and spatial features.
[0152] Through the improvement of the YOLOv8 model structure, this paper has enhanced the performance of the model in the task of abnormal detection of bolts on the vehicle side. These improvements not only improve the accuracy of the model but also enhance its ability to handle complex scenarios, demonstrating its efficiency, reliability, and developability in practical applications.
Claims
1. An intelligent detection system and method for train bogie bolts based on machine vision, characterized in that, The method includes the following steps: S0. Construct an intelligent detection system for bogie bolts of a vehicle: The system includes: a system trigger module that triggers an image data acquisition module by detecting whether a train is approaching; a position tracking module that, when a train approaches, monitors each bogie of the train in real time; a visual positioning module that performs visual positioning on the bolts of each bogie of the train; an image data acquisition module that acquires image data of the positioned bogie bolts through a camera; a data storage and transmission module that processes and stores the acquired image data and transmits it to YOLOv8; a YOLOv8 improvement module that introduces an ODConv dynamic convolution network to replace the traditional convolution network in the Neck layer and modifies the upsampling operator in the Neck layer to a CARAFE module.
2. An intelligent detection system and method for train bogie bolts based on machine vision, characterized in that, The method includes the following steps: S1. Monitor the approaching train in real time, trigger the image data acquisition module, and acquire image data of the train number. S2. Track the position of the train bogies and number the bogies. S3. Locate the bolts on the bogies and acquire image data of the bolts. S4. Process the acquired image data and transmit it to the improved YOLOv8. S5. On the basis of YOLOv8, introduce an ODConv dynamic convolution network to replace the traditional convolution network in the Neck layer and modify the upsampling operator in the Neck layer to a CARAFE module.
3. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that, The specific steps of step S1 include: S11. When the system trigger module detects the approach of a train, start acquiring image data and set different light sources for different environments. S12. Acquire data of the train number.
4. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, wherein The specific steps of step S2 include: S21. Install a position tracking system to perform real-time positioning on the bogies. S22. Number the positioned bogies.
5. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that, The specific steps of step S3 include: S31. By extracting unique features (such as shape, texture, color, etc.) of the track bolts, bolts can be effectively distinguished from other objects, and a positioning algorithm is used for accurate positioning. S32. The image acquisition device includes an industrial camera (such as a CCD camera or a CMOS camera) and a camera (the light source also needs to be considered). According to the intelligent detection system and method for bogie bolts of a train based on machine vision described in claim 1, the specific steps of step S4 include: S41. Introduce an ODConv dynamic convolution network to replace the traditional convolution network in the Neck layer and perform image enhancement on the acquired sample images. S42. After the output of the Darknet53 network, add an SPP network structure, perform max-pooling operations on the feature map through three max-pooling layers of different scales, and fuse the multi-scale output local feature maps of the track bolts with the input global feature map of the track bolts.
6. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that, The specific steps of step S5 include: S51. Modify the upsampling operator in the Neck layer to a CARAFE module. S52. At the same time, add a CBAM attention mechanism.
7. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S11. Specifically includes: The sensor needs to collect environmental data (such as temperature, humidity, etc.) and train motion parameters (such as speed, acceleration, distance, etc.) in real time; the data acquisition module converts these signals into digital signals and transmits them to the trigger control module; detects the arrival of the train and triggers image acquisition; Common sensors used include infrared sensors, radars, laser rangefinders, etc. These devices are installed along the railway or on the platform to detect the arrival of the train; infrared signals are emitted and received by the infrared sensor to detect whether an object (such as a train) passes by; the radar radio wave is used to measure the distance and speed of the target, and then the distance between the train and the monitoring point is measured through laser pulses; when the sensor detects the approach of the train, the trigger control module will make a logical judgment according to preset conditions (such as distance threshold, speed range); if the trigger condition is met (for example: the train is within 50 meters of the monitoring point), the system will send a trigger signal to the image acquisition module; the trigger signal can be a digital signal or an analog signal, which is used to control the startup of the image acquisition module; this process requires a quick response to ensure that the train number information can be accurately captured when the train passes by; The system trigger module is a highly integrated component that covers the complete process from data acquisition to condition judgment and then to execution control; its structure includes multiple functional modules such as a monitoring unit, a trigger condition judgment logic, an execution unit, etc., and each part has its specific role and the cooperation relationship between them; Set parameters such as sampling frequency, signal conditioning (such as amplification, filtering), etc. to adapt to the output characteristics of the sensor and the subsequent data analysis requirements; install the data acquisition device on the train or the ground control center and ensure a stable connection with the sensor; adopt wired communication (such as Ethernet, optical fiber) or wireless communication technology (such as 5G, Wi-Fi, ZigBee, etc.), and select the appropriate method according to the application scenario; deploy communication devices between the train and the monitoring center to ensure real-time data transmission. Use a database (such as MySQL, MongoDB) or a cloud storage system to classify and store the collected data for subsequent query and analysis.
8. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S12. Specifically include: Select an appropriate industrial camera or ordinary imaging device according to requirements; industrial cameras have high frame rates and high resolutions and are suitable for photographing high-speed moving objects; the camera is installed in a suitable position (such as on a fixed bracket beside the railway track) to ensure that the train number information can be clearly captured; set parameters such as exposure time, white balance, and focus to adapt to the ambient light and background conditions; adjust the frame rate according to the train speed to ensure that multiple images can be taken within a short time; after receiving the trigger signal, the camera activates the photographing function to obtain the image data at the current moment; if necessary, a continuous shooting mode can be set (such as taking one picture every 0.1 seconds) to record multiple states during the train's passing; perform preprocessing operations such as denoising and enhancement on the collected images to improve the image quality; extract the text and pattern features related to the train number information through binarization or edge detection algorithms; use optical character recognition (OCR) technology to recognize the text information such as the train number and train type in the image; store the collected image data and recognition results in a local or cloud database for subsequent query and analysis; and the data management system supports an efficient retrieval function, such as quickly finding relevant records according to conditions such as train number and time range; The method for real-time monitoring of trains mainly relies on advanced sensor technologies, communication technologies, and data analysis methods; by arranging various sensors on the train, key parameters are collected in real time; the data acquisition system and communication technologies are used to achieve data transmission and storage; machine learning algorithms and visualization tools are used to analyze and display the data and provide feedback and intervention; this method can effectively improve the safety and reliability of train operation and provide important technical support for the rail transit industry.
9. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S21. Specifically, it includes: The purpose of the bogie position tracking system is to obtain the position information (such as coordinates) of the bogie in real time through various sensors and technical means, and combine the vehicle operation status and environmental data to achieve high-precision positioning of the bogie position; this system can be applied to the safety monitoring, fault diagnosis, and operation management of rail transit vehicles; The bogie position tracking system mainly consists of the following parts: Sensor module: used to collect data such as the position, attitude, acceleration, and vibration of the bogie. Data acquisition and transmission module: responsible for converting the sensor signal into a digital signal and transmitting it to the main control unit through a communication network; positioning algorithm module: based on the sensor data, combined with kinematic models and positioning technologies (such as inertial navigation, GPS, etc.), to calculate the position of the bogie; data processing and analysis module: store, analyze, and visually display the positioning results; To achieve high-precision position tracking, it is necessary to select a combination of multiple sensors: (1) Accelerometer The accelerometer is used to measure the linear acceleration of the bogie, and its output signal can reflect the motion state of the vehicle. Assume that the measured value of the accelerometer is a x ,a y ,a z , then the speed and position of the bogie can be obtained by integration: (2) Gyroscope The gyroscope is used to measure the angular velocity of the bogie, and its output signal is ω x , ω y , ω z . The attitude angle of the bogie can be obtained by integrating the angular velocity: (3) Global Positioning System (GPS) GPS is used to provide the absolute position information of the bogie, and its output is the longitude and latitude coordinates (x, y) and altitude z. (4) Light Detection and Ranging (LiDAR) LiDAR can be used for high-precision positioning by scanning the environment and matching the map to achieve positioning; assuming that the measurement value of LiDAR is point cloud data Then the position and attitude of the bogie can be calculated through the feature matching algorithm; Data acquisition and preprocessing: Data acquisition: Sensor data needs to be sampled by an analog-to-digital converter (ADC) and transmitted to the main control unit in the form of digital signals; assuming the sampling frequency is f s , then N = f s samples can be collected per second; Data preprocessing: Denoising: Use a low-pass filter to eliminate high-frequency noise; for example, use a Butterworth filter: where ω c is the cut-off frequency and n is the filter order. Signal fusion: Align the signals from different sensors in time and adjust the scale. The positioning algorithm is the core of the system and requires the combination of multiple technologies to achieve high-precision positioning; (1) Inertial Navigation System (INS) Based on the data from accelerometers and gyroscopes, calculate the position and attitude of the bogie through integration: (2) GPS-aided positioning The inertial navigation result data x k+1 = Ax k + Bu k Observation equation: y k = Cx k + Dv k (3) Vision-based positioning Through cameras and image processing technologies, identify the environmental features around the bogie, and combine with the SLAM (Simultaneous Localization and Mapping) algorithm to achieve positioning; The overall architecture of the bogie position tracking system is as follows: Sensor module: accelerometers, gyroscopes, GPS, lidar, etc.; Data acquisition and transmission module: analog-to-digital converter (ADC), communication network (such as CAN bus, wireless communication); Positioning algorithm module: inertial navigation, Kalman filter, SLAM, etc. Data processing and analysis module: data storage, visualization interface, alarm system; Arrange marker points on the track, measure the position of the bogie through GPS and lidar, and compare it with the actual position; Use a shaker table to simulate the vehicle running state and test the measurement accuracy of accelerometers and gyroscopes.
10. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S22. Specifically include: Use sensors or readers to collect the geographical location information of the bogie in real time; Transmit the position data to the central control system or database through wired or wireless networks to ensure the real-time and accuracy of the data; Design a numbering rule according to actual needs and encode in sequence. Ensure the uniqueness of the numbering to avoid duplication or confusion; Process the real-time position data collected in the central control system; Store the processed data in the database and associate it with the numbering information of the bogie to establish a one-to-one correspondence; Display the position and numbering information of each bogie in real time on the user interface for easy management and monitoring; Provide a search function to allow users to quickly find the position information of a specific bogie according to the number; Regularly check the stability of the positioning system to ensure the normal operation of sensors or readers; Regularly back up and clean the database to prevent data redundancy or damage.
11. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S31. Specifically include: Arrange high-precision sensors for the running tram, such as laser displacement sensors, image acquisition devices or ultrasonic detection devices; Use the Canny operator or Sobel operator to detect the edges in the image and capture the contour of the bolt; Adopt sliding window detection and positioning, traverse the image through the sliding window, and use feature matching or classifiers (such as SVM, AdaBoost, etc.) to judge whether each window contains a bolt. Scan each key bolt on the track one by one and record information such as the position coordinates, tilt angle and tightening state of each bolt; Because the train is running, it is also necessary to consider the influence of factors such as vibration and temperature on the data; Denoise and filter the collected raw data to eliminate environmental interference and sensor noise; Establish a standard model of the bolt position and compare the actual measurement values with the standard values.
12. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S32. Specifically include: Select a suitable industrial camera or ordinary imaging device according to requirements; industrial cameras have high frame rates and high resolutions and are suitable for photographing high-speed moving objects; the camera is installed in a suitable position (such as on a fixed bracket beside the railway track) to ensure that the train number information can be clearly captured; set parameters such as exposure time, white balance, and focus to adapt to the ambient light and background conditions.
13. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S41. Specifically, it includes: Adjust the brightness of the track bolt sample picture. The adjustment method is: f(x) = g(x) + β, where f(x) is the adjusted track bolt image, g(x) is the original track bolt image, and β is the brightness adjustment coefficient; perform 90° and 180° clockwise and counterclockwise flips on the brightness-adjusted track bolt samples to achieve sample expansion; perform blurring on the expanded images, and use Gaussian blurring for the sample data. The mathematics of Gaussian blurring: Expressed as: Among them, σ is the standard deviation, r is the blurring radius, and N is the normalization constant of the Gaussian distribution, which ensures that the sum of the Gaussian function is 1, so that the brightness of the image remains unchanged and no information is lost due to blurring. ODConv full-dimensional dynamic convolution realizes dynamic adjustment of the shape and size of the convolution kernel according to different input data features by introducing a learnable deformation module. This method processes from the spatial dimension, input channel dimension, and output channel dimension, and can pay attention to the shape and size of the convolution kernel to better adapt to different input data; the main formula of ODConv full-dimensional dynamic convolution is as follows: y = (α W1 W1 +... + α Wn W n ) * x where W i is the i-th convolution kernel, and α wi is the attention scalar for weighting W i . * represents the convolution operation, and ⊙ represents the multiplication operation of different dimensions. ODConv multiplies the attention scalar with the convolution kernel to adjust the size and shape of the convolution kernel, and then performs a convolution operation with the input feature x to obtain the final output feature y. First, the global average pooling layer GAP (global average pooling) is used to compress the target input feature. Then, the fully connected layer FC (fully connected layers) is used to reduce the complexity of the dynamic convolution. Next, the ReLU function is used to remove the negative values of the feature vector, generating 4 head branches to construct 4 types of attention scalars of the ODConv convolution module. Finally, the Sigmoid activation function and the Softmax function are used to normalize the processed features to generate the normalized attention scalars α si , α ci , α fi , α wi . Each attention scalar is multiplied and summed with the original convolution kernel W i , and then combined with the input feature to obtain the output feature; Among them, the formula of the Softmax function is as follows: Use the Softmax function to convert each value in the input feature x i into a numerical value between 0 and 1, and the sum of all output values is 1. In image detection, it can distinguish and identify multiple objects or object categories existing in the image, enhance the interpretability of model prediction, and simplify the loss calculation and model training process; The formula of the ReLU activation function is as follows: ReLU(x) = max(0, x) The ReLU activation function enhances the nonlinear characteristics of the model by setting the negative value part to 0, which helps the model capture complex features in the image, thereby improving its expression ability. At the same time, adding 0 element outputs in the output represents the sparsity of the features, so it can help the model ignore irrelevant background information in the image and focus on meaningful target features. The formula of the Sigmoid function in the improved structure is as follows: The Sigmoid function is used to generate attention weights in each dimension, and these weights reflect the importance of different positions and channels, so that the input features can be processed more precisely. The Sigmoid function maps each generated attention scalar to between 0 and 1, which can effectively integrate different feature information and improve the overall performance of the model.
14. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S42. Specifically, it includes: For the need to extract deeper semantic information of track bolt detection features and achieve multi-scale feature prediction, improve the Darknet53 network layer; since there is a large size difference between the track nut and the overall track bolt in the input track bolt image, after the output of the Darknet53 network, refer to the method of Spatital PyrmidPooling (SPP) and add an SPP network structure to strengthen the network feature performance; The SPPF model has a smaller computational cost and faster operation speed, and can solve the problem of inconsistent input image sizes; the SPPF module extracts features at different scales through max-pooling operations; specifically, it has four parallel branches: the first branch: directly passes the input features to the output without any operation. The second branch: performs a max-pooling operation with a kernel size of 5×5 to obtain medium-scale features through downsampling; the third branch: uses a max-pooling operation with a kernel size of 9×9 to further extract larger-scale features; the fourth branch: adopts a max-pooling operation with a kernel size of 13×13 to extract even larger-scale features; The stride of all pooling operations is 1, which means that padding is required during the convolution operation to keep the feature map size unchanged; finally, the feature maps of these four branches are concatenated to complete the fusion of features at different scales; this structure realizes efficient multi-scale information integration, which helps to capture targets of different sizes in the object detection task; by performing max-pooling operations on the feature map through three max-pooling layers of different scales, the fusion of local features and global features of the input track bolt feature map is achieved; the fusion of the multi-scale output track bolt local feature map and the input track bolt global feature map enriches the expression ability of the feature map.
15. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S51. Specifically, it includes: The core module of this structure includes a Kernel Prediction Module and a Content-aware Reassembly Module. The Kernel Prediction Module consists of three parts: a Channel Compressor, a Content Encoder, and a Kernel Normalizer; the Content-aware Reassembly Module is composed of upsampling kernel normalization, feature mapping back to the input feature map, and output value calculation; First, the Channel Compressor uses a 1*1 convolution to compress the number of channels of the input feature map to reduce the computational cost of subsequent steps. The feature channel compression formula is as follows: y = Conv1*1(x) where x and y are the input and compressed feature maps respectively, and Conv 1*1 represents applying a 1×1 convolution operation. Then, the content encoder encodes the compressed feature map so that each upsampling kernel can represent the content features at its corresponding position, and it usually contains a convolution kernel of size k up ×k up (the size of the upsampling kernel). The formula of the content encoder is as follows: where x and y are the feature maps of the input and output respectively, and the upsampling kernel size is k up ×k up Predict the upsampling kernel through a convolutional layer of k encoder *k encoder The number of output channels is Unfold the channel dimension in the spatial dimension to obtain an upsampling kernel with a shape of ; Secondly, the Kernel Normalizer performs Softmax normalization on the generated upsampling kernel to ensure the numerical stability and rationality during the upsampling process; finally, the Content-aware Reassembly Module uses the obtained upsampling kernel to recombine the content of the input feature map, realizing instance-specific content-aware processing and dynamically generating adaptive upsampling weights for each position.
16. The intelligent detection system and method for train bogie bolts based on machine vision according to claim 1, characterized in that S52. Specifically, it includes: CBAM is a module that integrates channel and spatial attention mechanisms, enhancing the network's ability to capture important features while suppressing unnecessary information; the CBAM module, as shown in Figure 5, fine-tunes features in two dimensions by combining a Channel Attention Module (CAM) and a Spatial Attention Module (SAM); the Channel Attention Module and the Spatial Attention Module respectively emphasize the importance of each pixel in different channel features and spatial features; The input features enter the CAM and SAM successively, and a global average pooling (AvgPool) and a global maximum pooling (MaxPool) are both executed in parallel in these two modules. Among them, the global average pooling is used to obtain the average value of each channel / space, reflecting the global information of the feature map; the global maximum pooling captures the maximum value of each channel / space, highlighting the local maximum amplitude in the feature map. The combination of the two can take into account both the average information and the maximum amplitude information of the feature map at the same time. In the channel attention module CAM, after global average pooling and global max pooling, a multi-layer perceptron (MLP) consisting of two fully connected layers is used to generate the final channel attention weights; the Sigmoid activation function is used to obtain the attention weights for each channel, and the weights are normalized to the range of 0 to 1 so that they can be used as scaling factors for the input feature map; finally, the obtained attention weights are multiplied by the input feature map to obtain M c , and the calculation formula of the channel attention module is as follows: In the Spatial Attention Module (SAM), after global average pooling and global max pooling, the results are reduced to single-channel features through a convolution with a kernel size of 7×7, and the final feature map M is obtained through processing by the Sigmoid activation function. s , and the calculation formula of the spatial attention module is as follows:
Citation Information
Cited By
Rail bolt maintenance system and method
CN121061567A
Fault detection system and method based on fusion of visual perception and physical perception of bolts under vehicle
CN122451657A