A turnout detection method based on computer vision
By improving the YOLOv8 model, enhancing the C2f module and network structure, and combining with improving the CIoU loss function, the accuracy of small and medium-sized defect detection of railway switch detection is solved, and the detection accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510171996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The prior art is difficult to accurately detect small-scale defects in railway switch detection, and it is difficult to extract features in complex environments, resulting in false alarms and missed alarms.
Using the improved YOLOv8 model, the C2f module is enhanced by the state space model, the C2f-VPU structure is designed to capture long-distance dependence and context features, and a small object detection enhanced pyramid structure is introduced at the net layer of the network, using an improved CIoU loss function with scaling factors and auxiliary bounding boxes.
It improves feature extraction capability and small object detection accuracy in complex environments, significantly enhances the detection capability of small-scale defects, and reduces false alarms and missed response rates.
Smart Images

Figure CN119648700B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of turnout detection, and in particular to a turnout detection method and an intelligent monitoring pile based on computer vision. Background Art
[0002] Railway is an important infrastructure of modern economy and the main mode of public transportation. As a key component for switching train tracks, turnout devices play a vital role in the railway system. With the increase of train mileage, turnout devices will inevitably have defects such as peeling, cracks and wear. If these defects cannot be detected and processed in time, they will pose a serious threat to the safety of train operation and passenger safety. Therefore, timely and accurate detection of rail defects is crucial to ensure the safety of railway operations. In addition to traditional image processing and deep learning-based object detection methods, vertical projection method, using adaptive threshold segmentation to extract defect areas, most traditional methods rely on the texture and color of defects to locate and identify defects. When the color and shape features of the surrounding environment are similar to the defects, it will lead to large detection errors. The YOLOv8 network model strikes a good balance between speed and accuracy. However, in a complex railway environment, factors such as weather and shooting angles may make it difficult to extract target features, resulting in information loss. In addition, the shape, size and specifications of defects vary greatly. Most images contain small, numerous and unevenly distributed small-scale defects, which may lead to false positives and false negatives, affecting the overall detection accuracy. Therefore, how to design an effective method for small object detection in railroad tracks is a major challenge. Summary of the invention
[0003] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a turnout detection method and intelligent monitoring pile based on computer vision. In order to improve the detection accuracy, the present invention proposes an improved YOLOv8 model, which adopts a state space model to enhance the C2f module in the YOLOv8 backbone network, and proposes C2f-VPU to better capture long-distance dependencies and context features, thereby improving the feature extraction capability in complex environments. In the neck layer of the network, the present invention uses a spatial-depth convolution module and a feature enhancement-fusion module to jointly improve the original path aggregation feature pyramid network structure, and proposes a small target detection enhancement pyramid structure to enhance the small target detection capability. In addition, an improved loss function with a scaling factor and an auxiliary bounding box is used to further enhance the detection capability of the model.
[0004] The present invention first discloses a turnout detection method based on computer vision, and the specific steps are as follows:
[0005] Step 1: Construct a rail defect dataset; collect images of the turnout device by a visual displacement measurement unit; the images include images with defects and images without defects, and the images are taken at different times, various weather conditions, and different lighting environments;
[0006] Step 2: All data sets are labeled to obtain small-scale defect data; the data sets are divided into training sets, test sets, and validation sets; and data augmentation techniques are used to randomly expand the data of the training sets and validation sets;
[0007] Step 3: Use the improved YOLOv8 model to train the defect data of the turnout device to achieve better feature extraction and feature fusion, and enhance the ability to locate and detect small-scale defects; the improved YOLOv8 model is specifically:
[0008] 1) Based on the original YOLOv8, the visual processing unit VPU is used to improve the C2f structure of the seventh and ninth layers in the backbone network;
[0009] 2) Use a small target detection enhancement pyramid structure to enhance the small target detection capability; the small target detection enhancement pyramid structure includes a spatial-depth convolution module and a feature enhancement-fusion module; the P2 layer features of the backbone network are processed by the spatial-depth convolution module to retain more fine-grained features, and the processed features are fused with the output features of the P3 layer and input into the feature enhancement-fusion module to enhance the feature extraction and fusion capabilities, thereby more effectively capturing and identifying defects;
[0010] 3) Use the improved CIoU with scaling factor and auxiliary bounding box as the loss function to adapt to defects of different shapes and sizes and improve the localization effect of the detection box;
[0011] Step 4: Control the sensor to collect data and upload the sensor measurement data to the industrial computer; the industrial computer outputs the defect data in the collected data through the improved YOLOv8 model and issues an alarm.
[0012] Furthermore, the structure of the visual processing unit VPU in step 3 is specifically as follows:
[0013] Layer normalization and branching, the input data is divided into two branches after layer normalization; the first branch is processed by a linear layer and an activation function; the second branch is processed by a linear layer, a depthwise separable convolution, and an activation function; and then input to the selective scanning module for feature extraction;
[0014] Feature merging, the data after feature extraction is layer normalized and element-wise multiplied with the output of the first branch to merge the two paths;
[0015] Linear layers and residual connections, use linear layers to mix features and combine this result with residual connections to form the output of the VPU.
[0016] Furthermore, the specific method of the selective scanning module for feature extraction is:
[0017] Given the input features, the output features of the selective scanning module can be expressed as:
[0018]
[0019] Among them, v represents four different scanning directions; the expansion operator expands the input feature z along these four directions to generate a set of one-dimensional sequences ; S6 is the core operator in the VPU block, which is used to perform the expanded features Processing to generate new features ; The merge operator merges the four processed one-dimensional sequences into a new complete two-dimensional feature map .
[0020] Furthermore, the space-depth convolution module in step 3 is composed of a space-depth layer and a non-strided convolution layer, specifically:
[0021] First, the feature map is input into the space-depth layer, which divides the feature map into four sub-feature maps along the x and y directions. The mapping method is:
[0022]
[0023] in, represents the sub-feature map in the i-th row and j-th column; X represents the input feature map; i and j represent the index in the horizontal and vertical directions respectively; S represents the step size; scale represents the scaling factor, which is used to determine the size of the sub-feature map;
[0024] Subsequently, the sub-feature maps are concatenated together to obtain an intermediate feature map, which reduces the spatial size of the feature map and increases the feature information in the channel dimension;
[0025] Finally, the intermediate feature maps are sent to the non-strided convolutional layer to obtain the final feature maps.
[0026] Furthermore, the feature enhancement-fusion module in step 3 has the following specific process:
[0027] In the feature enhancement-fusion module, the input features are divided into two parts; one part is processed by the full-core module and then fused with the unprocessed feature map in the other part;
[0028] Among them, the full-core module includes global branch, macro branch and micro branch; the dual-domain channel attention module and frequency-based spatial attention module in the global branch enhance the feature representation of the track area; the macro branch uses three large-scale deep convolutions of 63×1, 63×63, and 1×63 to enhance the recognition ability of medium and large-scale targets; the micro branch extracts local features through 1×1 deep convolution operation to ensure sensitivity to small-scale defects; after the input information is transmitted and accumulated through the three branches, the detection ability of small-scale defects is enhanced.
[0029] On the other hand, the present invention further provides an intelligent monitoring pile, which can execute any of the above-mentioned intelligent rail detection methods based on computer vision, and the intelligent monitoring pile comprises:
[0030] 1 visual displacement measurement unit, 2 laser ranging units, 1 edge computing analysis unit, 4G transmission unit, solar power supply unit, integrated monitoring pile structure, installation base; among them,
[0031] Visual displacement measurement unit: collects rail images and measures the longitudinal displacement and relative height difference between the left and right rails; it is equipped with a near-infrared light source to provide supplementary light for the camera in dim light or at night;
[0032] Laser distance measuring unit: measure the gauge of two rails;
[0033] Edge computing analysis unit: controls sensors to collect data and uploads sensor measurement data to industrial computers;
[0034] 4G transmission unit: Upload measurement data and image data.
[0035] Furthermore, the integrated monitoring pile structure is:
[0036] Camera cabin: used to place visual cameras in the pile body; Camera single-axis bracket: used to install camera equipment, the camera single-axis bracket is installed in the camera cabin; Protective glass: used to protect the camera cabin from dust and rain; Data processing cabin: used to place edge computing industrial computers; Battery unit: used to place batteries; Solar panel module: power supply; Mains access port: connected to the mains; Cable trough: used to place equipment lines; Expansion cabin: to expand other functions;
[0037] Dual-axis bracket: for mounting the camera on the top of the pile, and for expanding multiple cameras; Waterproof unit: for preventing rainwater immersion.
[0038] The present invention relates to a turnout detection method and intelligent monitoring pile based on computer vision, which has the following beneficial effects:
[0039] 1) Based on the original YOLOv8, the present invention improves the backbone network, uses the visual processing unit VPU to improve the C2f structure of the seventh and ninth layers in the backbone network, and designs the C2f-VPU structure, which enhances the model's ability to identify features and capture long-distance dependencies in complex environments;
[0040] 2) A small target detection enhancement pyramid structure is proposed to enhance the small target detection capability. The neck layer of the network is improved by introducing a spatial-depth convolution module and a feature enhancement-fusion module. This significantly enhances the model's ability to detect small-scale defects;
[0041] 3) In order to enhance the loss function, the present invention introduces an improved CIoU with a scale factor to replace CIoU, allowing the size of the auxiliary bounding box to be dynamically controlled, thereby achieving adaptive adjustment of defect detection boxes of different scales. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a schematic diagram of the original YOLOv8 structure;
[0043] Figure 2 Schematic diagram of railway tracks under different weather conditions and times;
[0044] Figure 3 This is the structure diagram of the improved YOLOv8 algorithm model;
[0045] Figure 4 It is the structural block diagram of VPU;
[0046] Figure 5 Schematic diagram of the integrated monitoring pile structure. DETAILED DESCRIPTION
[0047] The detailed description set forth below in conjunction with the accompanying drawings is intended to be a description of various exemplary embodiments of the present invention, and is not intended to represent the only embodiment that can practice the present invention. For the purpose of providing a thorough understanding of the present invention, the detailed description includes specific details. However, it is apparent to those skilled in the art that the present invention can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid blurring the concept of the present invention.
[0048] YOLOv8 is designed to perform tasks including object detection, instance segmentation, and image classification. It has five versions of different sizes, and the present invention uses the compact YOLOv8 as an improved baseline model.
[0049] Depend on Figure 1As shown in the figure, the YOLOv8 model mainly consists of three components: backbone network, neck network and head network. The backbone network uses Darknet53 as its framework and introduces the gradient-rich C2f module to replace the C3 module in YOLOv5. This significantly improves the convergence speed and effect of the model. In the neck network part, YOLOv8 adopts the structure of PAN (Pyramid Attention Network)-FPN; in the head network, it can output small-size data of 80×80, medium-size data of 40×40, and large-size data of 20×20 respectively. Among them, CBS is the abbreviation of Conv-BatchNorm2d-Silu (Convolution-2d Batch Normalization-Activation Function); C2f is a convolution block with feature fusion; Upsample is an upsampling module; Concat is a splicing module; SPPF is a spatial pyramid pooling-fixed version.
[0050] Step 1: Construct a rail defect dataset; collect images of the turnout device by a visual displacement measurement unit; the images include images with defects and images without defects, and the images are taken at different times, various weather conditions, and different lighting environments.
[0051] The existing datasets are mainly composed of close-up images of railway tracks, which have a relatively simple structure and are mainly composed of two rails. However, since the structure of the turnout device is more complex, mainly including the frog core, wing rails, guard rails and other connecting parts, these existing datasets cannot meet the experimental requirements. In order to better complete the defect detection task of the turnout device, the present invention constructs a rail defect dataset. The images of the turnout device are collected by the visual displacement measurement unit. These images are taken in JPEG format with a resolution ranging from 1200×1600 to 3000×4000 pixels. The camera angle is a bird's-eye view, and the distance from the turnout device is kept between 0.5 and 1.5 meters.
[0052] In order to ensure the diversity and availability of data samples, the present invention selects 550 images, including 500 images with defects and 50 images without defects, which are taken at different times, in various weather conditions and under different lighting environments. Figure 2 shown.
[0053] Step 2: All data sets are labeled to obtain small-scale defect data; the data sets are divided into training sets, test sets, and validation sets; and data augmentation techniques are used to randomly expand the data of the training sets and validation sets;
[0054] All data sets were annotated using software, and a total of 1,440 defects were found. Most of them are small-scale defects, and the defect boxes are mostly long strips; the data set is divided into training set, test set, and validation set in a ratio of 7:2:1. In order to expand the data set and improve the generalization ability of the model, the present invention uses data enhancement techniques such as random scaling, flipping, brightness enhancement, and grayscale conversion to randomly expand the data of the training set and validation set. The test set only retains real data to avoid introducing human errors. After enhancement, the training set is expanded to 1,750 images and the validation set is expanded to 250 images.
[0055] The main features of this dataset include the following: 1) complex and diverse image backgrounds, including various weather conditions and environmental scenes; 2) a wide range of defect sizes, with significant size differences between defects in the same image, ranging from 30 to 300 pixels, and small-scale defects dominating; 3) images taken from multiple angles, which also causes the shape and size of defects to be affected by changes in shooting angles, resulting in diverse shapes and target appearances with larger aspect ratios. These features make the dataset challenging and help train models with good generalization performance.
[0056] Step 3: Use the improved YOLOv8 model to train the defect data of the turnout device to achieve better feature extraction and feature fusion, and enhance the ability to locate and detect small-scale defects;
[0057] Due to the small pixel size of defect targets in the turnout device defect dataset, large scale differences between defects, complex background, and similarities between certain backgrounds and defects, this paper proposes a new improved YOLOv8 algorithm model for detecting turnout device defects to achieve better feature extraction and feature fusion, and enhance the ability to locate and detect small-scale defects.
[0058] and Figure 1 Compared with the baseline model in , the present invention improves some modules based on YOLOv8. Figure 3 As shown, first, based on the original YOLOv8, the present invention uses the visual processing unit VPU to improve the C2f structure of the seventh and ninth layers in the backbone network, and designs a feature fusion convolution block C2f-VPU based on the improved visual processing unit. The visual processing unit VPU converts the image into four groups of one-dimensional sequences, and uses a state space model to capture long-term dependencies and contextual information in the image, thereby enhancing feature extraction in complex environments.
[0059] Secondly, in the neck layer of the network, the present invention improves the original PAN structure and introduces a small target detection enhanced pyramid structure. By integrating the spatial-depth convolution module (SDC) and the feature enhancement-fusion module (FEF), Figure 3 The P2 layer (the first C2f structure) of the backbone network is rich in small target information. After passing through the spatial-depth convolution module, more fine-grained features are retained, and the processed features are fused with the output features of the P3 layer (the second C2f structure) and input into the feature enhancement-fusion module. The feature enhancement-fusion module includes global, macro and micro branches, which effectively learns features from global to local and improves the performance of small-scale defect detection.
[0060] Finally, a modified CIoU with a scale factor and auxiliary bounding boxes is used as the loss function to adapt to defects of different shapes and sizes, improve the localization effect of the detection box, and further improve the detection accuracy.
[0061] Step 4: Control the sensor to collect data and upload the sensor measurement data to the industrial computer; the industrial computer outputs the defect data in the collected data through the improved YOLOv8 model and issues an alarm.
[0062] The structure of the visual processing unit VPU is as follows:
[0063] Layer normalization and branching, the input data is divided into two branches after layer normalization; the first branch is processed by a linear layer and an activation function; the second branch is processed by a linear layer, a depthwise separable convolution, and an activation function; and then input to the selective scanning module for feature extraction;
[0064] Feature merging, the data after feature extraction is layer normalized and element-wise multiplied with the output of the first branch to merge the two paths;
[0065] Linear layers and residual connections, use linear layers to mix features and combine this result with residual connections to form the output of the VPU.
[0066] The defect dataset of turnout devices is affected by outdoor environmental factors, resulting in high randomness and unclear features, which makes feature extraction challenging and leads to information loss that affects detection accuracy. CNN operations can only perceive local features at each layer, making it difficult to capture long-term dependencies. Transformers excel in global modeling and can effectively capture long-term dependencies, but their self-attention mechanism requires a lot of computation. To overcome the limitations of CNN and Transformers, state-space models establish long-term dependencies while maintaining linear complexity. Integrating time-varying parameters into the state-space model excels in capturing long-term dependencies and enables efficient parallel training.
[0067] The present invention integrates VPU into YOLOv8 to improve C2f and proposes a C2f-VPU module. The state-space model-based architecture is applied to YOLO to process one-dimensional data, combined with its advantage of capturing global dependencies, to improve the model's feature extraction ability and understanding of complex scenes, thereby improving the model's detection accuracy. When processing sequence data, the state-space model can automatically adjust weights to better capture important information, focus on turnout devices, reduce the impact of environmental factors, and maintain linear complexity without increasing the computational burden.
[0068] The working principle of VPU is as follows Figure 4 As shown in the figure. After layer normalization (Layer Norm), the input is divided into two independent information streams. One branch passes through a linear layer (Linear) and an activation function (SiLU), while the other branch is processed by a linear layer (Linear), a depthwise separable convolution (DW Conv) and an activation function (SiLU), and then reaches the core component of the VPU block: the selective scanning module (BD-SS). Finally, these streams are merged through another layer of normalization layer (LayerNorm), and the final output is generated after the linear layer (Linear) mixes the features, and with a residual connection, it forms the output of the VPU.
[0069] The VPU addresses the challenges associated with planar images through the selective scanning module. It unfolds the image into four different sequences, processes each sequence using a state-space model, and then merges the output features to form a new, complete planar feature map. Given the input features, the output features of the selective scanning module are represented as:
[0070]
[0071] Among them, v represents four different scanning directions; the expansion operator expands the input feature z along these four directions to generate a set of one-dimensional sequences , the purpose is to capture the long-distance dependencies and contextual information in the image; S6 is the core state space model operator in the VPU block, which represents the expanded features Processing to generate new features , which aims to extract and enhance feature information in order to better perform subsequent feature fusion and detection tasks; the merge operator merges the processed four one-dimensional sequences into a new complete two-dimensional feature map ; The scanning operation in four directions covers all areas of the image and provides rich multi-dimensional information, which improves the comprehensiveness of capturing image features; Figure 4 The numbers 1-25 represent the sequence data of the input feature graph.
[0072] The C2f-VPU module replaces the original C2f at the seventh and ninth layers at the end of the backbone network, enabling the model to learn features associated with defect targets and backgrounds, effectively capturing complex details and broader semantic context in images, thereby improving detection accuracy.
[0073] A small target detection enhancement pyramid structure is used to enhance the small target detection capability; wherein the small target detection enhancement pyramid structure includes a spatial-depth convolution module and a feature enhancement-fusion module;
[0074] In the task of defect detection of turnout devices, there are many small-scale defects, which are difficult to detect. In the YOLOv8 model, there are only three types of detection heads: small, medium and large. The downsampling and pooling process in the backbone network may result in a lower resolution of the final high-level feature map, resulting in fewer small target pixels in the deep feature map. Detecting small-scale defects in the normal P3, P4 and P5 detection layers can be challenging, often resulting in missed detections. A fourth layer has been added in the prior art specifically for small target detection, but this significantly increases the number of parameters and affects the efficiency of the model in practical applications. Therefore, the present invention modifies the network structure of the neck layer and proposes a small target detection enhanced pyramid structure to improve the model's ability to detect small-scale defects.
[0075] The present invention first uses the features of the P2 layer, which have not been over-downsampled and retain more fine-grained information, and then processes them through the spatial-depth convolution module, which contains rich small target information and merges them with the features of the P3 layer. CNN models usually use strided convolution and pooling to downsample and output feature maps of a specific size, but this may result in the loss of fine-grained information. The spatial-depth convolution module consists of spatial-depth layers and non-stride convolution layers, and performs well in low-resolution images and small target detection tasks.
[0076] First, the feature map is input to the space-depth layer. When down-sampled by a ratio of 2, the space-depth layer divides the feature map into four sub-feature maps along the x and y directions. The mapping method is shown in formulas (4) and (5):
[0077]
[0078] in, represents the sub-feature map in the i-th row and j-th column; X represents the input feature map; i and j represent the index in the horizontal and vertical directions respectively; S represents the step size; scale represents the scaling factor, which is used to determine the size of the sub-feature map; is the sub-feature map in the first row and the first column;
[0079] The sub-feature maps are then concatenated together to obtain an intermediate feature map, which reduces the spatial size of the feature map and increases the feature information in the channel dimension. Finally, the feature map is sent to the non-strided convolution layer to obtain the final feature map. In the above operation, compared with the traditional convolution, the spatial-depth convolution module reduces the number of channels while retaining the global spatial feature information in the channel dimension, has a higher degree of information retention, and outputs image information with finer granularity, which helps improve the model's ability to detect small-scale defects.
[0080] The output feature information of the P2 and P3 layers is then fed into the feature enhancement-fusion module to enhance the feature extraction and fusion capabilities, thereby more effectively capturing and identifying defects. The feature enhancement-fusion module is an improvement based on the full-core module, which reduces the computational load while retaining more feature information. In the feature enhancement-fusion module, the input features are divided into two parts. One part is processed by the full-core module and then fused with the unprocessed feature map, which reduces the consumption of some computing resources. The full-core module consists of three branches: global, macro, and micro, which effectively learns features from global to local, thereby improving the detection accuracy of small target defects. In the microscopic branch, local features are extracted through 1×1 deep convolution operations to ensure sensitivity to small-scale defects; in the macroscopic branch, three large-scale convolutions are used to enhance the recognition ability of medium and large-scale targets; in the global branch, the feature representation of the track area is enhanced through the dual-domain channel attention module and the frequency-based spatial attention module, allowing the model to pay more attention to potential defect areas. After the input information is passed through the three branches and accumulated and fused, it captures the global context information in a wider range, improves the defect detection accuracy, and enhances the detection capability of small-scale defects.
[0081] Use the improved CIoU with scaling factor and auxiliary bounding box as the loss function to adapt to defects of different shapes and sizes and improve the positioning effect of the detection box;
[0082] In YOLOv8, CIoU loss is used as the loss function for the bounding box task to evaluate the accuracy of bounding box prediction, as shown in formula (6):
[0083]
[0084] Among them, IoU is the intersection over union ratio; The predicted bounding box b and the true bounding box The square of the Euclidean distance between the center points; The square of the diagonal length of the minimum enclosed area between the predicted bounding box and the true bounding box; is the dynamic adjustment coefficient; v includes the aspect ratio prediction.
[0085] However, since the defect sizes and shapes in the turnout defect dataset vary, some defects are elongated, and there are many small objects, the CIoU loss cannot accurately reflect the actual situation. This limitation affects the accuracy of the model bounding box prediction. To solve this problem, the present invention introduces an improved IoU based on auxiliary bounding boxes to replace CIoU as the loss function. A scaling factor is also introduced to control the generation of auxiliary boxes of different sizes for loss calculation. Combining the improved IoU with CIoU forms a new loss function - improved CIoU, whose calculation formula is shown in formulas (7)-(13):
[0086]
[0087]
[0088] Among them, the coordinates of the center point of the target bounding box are given by express; are the coordinates of the left, right, top, and bottom boundary points of the target bounding box, respectively; and are the width and height of the target bounding box; are the coordinates of the left, right, top, and bottom boundary points of the predicted bounding box; the coordinates of the center point of the predicted bounding box are represents; w and h are the width and height of the predicted bounding box; ratio is the scaling factor used to adjust the size of the auxiliary bounding box;
[0089] The improved CIoU loss function can dynamically control the size of the auxiliary bounding box by introducing a scaling factor ratio, which enhances the sensitivity to the target boundary; for defects of different shapes, sizes and scales, it can adaptively adjust the position and size of the border to better match targets of different scales; for smaller targets, by adaptively adjusting the ratio, the range of the auxiliary boundary can be effectively expanded, so that when calculating the loss, the model can capture more small target feature information; this flexibility gives the model stronger generalization ability, enabling it to better adapt to diverse detection scenarios.
[0090] When the ratio is less than 1, the size of the auxiliary bounding box is smaller than the actual bounding box, which results in a smaller effective regression range than the IoU loss function, but produces a larger absolute value of the gradient. This accelerates the convergence speed of high IoU samples. On the contrary, when the ratio is greater than 1, the size of the auxiliary bounding box is larger than the actual bounding box. This expands the effective regression range and enhances the regression effect of low IoU samples. By adjusting the ratio value, the improved CIoU loss function can adaptively select the appropriate auxiliary bounding box size according to different sample IoU levels, achieve effective convergence and improve the generalization ability of the model.
[0091] Example 2
[0092] Depend on Figure 5 It can be seen that the present invention also provides an intelligent monitoring pile, which can execute any of the above-mentioned intelligent rail detection methods based on computer vision, and the intelligent monitoring pile includes:
[0093] 1 visual displacement measurement unit, 2 laser ranging units, 1 edge computing analysis unit, 4G transmission unit, solar power supply unit, integrated monitoring pile structure, and installation base;
[0094] Visual displacement measurement unit: collects rail images and measures the longitudinal displacement and relative height difference between the left and right rails; it is equipped with a near-infrared light source to provide supplementary light for the camera in dim light or at night;
[0095] Laser distance measuring unit: measure the gauge of two rails;
[0096] Edge computing analysis unit: controls sensors to collect data and uploads sensor measurement data to industrial computers;
[0097] 4G transmission unit: Upload measurement data and image data.
[0098] Furthermore, the integrated monitoring pile structure is:
[0099] Camera cabin: used to place visual cameras inside the pile body; Camera single-axis bracket: for installing camera equipment, the camera single-axis support is installed in the camera cabin; Protective glass: for protecting the camera cabin from dust and rain; Data processing cabin: for placing edge computing industrial computers; Battery unit: for placing batteries; Solar panel module: for power supply; Mains access port: for connecting to the mains; Cable trough: for placing equipment lines; Expansion cabin: for expanding other functions.
[0100] Dual-axis bracket: for mounting the camera on the top of the pile, and for expanding multiple cameras; Waterproof unit: for preventing rainwater immersion.
[0101] In summary, the present invention proposes a turnout detection method and intelligent monitoring pile based on computer vision, collects images from intelligent detection equipment, and constructs a defect dataset of turnout devices in complex environments. In order to improve the detection accuracy, the present invention proposes an improved YOLOv8 model, which adopts a state space model to enhance the C2f module in the YOLOv8 backbone network, and proposes a C2f improvement module to better capture long-distance dependencies and context features, thereby improving the feature extraction capability in complex environments. In the neck layer of the network, the present invention uses a spatial-depth convolution module and a feature enhancement-fusion module to jointly improve the original path aggregation feature pyramid network structure, and proposes a small target detection enhancement pyramid structure to enhance the small target detection capability. In addition, a loss function with a scaling factor is applied to further enhance the detection capability of the model. Compared with the baseline model, the average accuracy of the improved YOLOv8 model on the rail detection dataset is improved by 3.5%, showing higher accuracy and robustness, indicating that the improved YOLOv8 model has good generalization ability.
[0102] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A turnout detection method based on computer vision, the specific steps are as follows: Step 1: Construct a rail defect dataset; collect images of the turnout device by a visual displacement measurement unit; the images include images with defects and images without defects, and the images are taken at different times, various weather conditions, and different lighting environments; Step 2: All data sets are labeled to obtain small-scale defect data; the data sets are divided into training sets, test sets, and validation sets; and data augmentation techniques are used to randomly expand the data of the training sets and validation sets; Step 3: Use the improved YOLOv8 model to train the defect data of the turnout device to achieve better feature extraction and feature fusion, and enhance the ability to locate and detect small-scale defects; The improved YOLOv8 model is as follows: 1) Based on the original YOLOv8, the visual processing unit VPU is used to improve the C2f structure of the seventh and ninth layers in the backbone network; Among them, the structure of the visual processing unit VPU is as follows: Layer normalization and branching, the input data is divided into two branches after layer normalization; the first branch is processed by a linear layer and an activation function; the second branch is processed by a linear layer, a depthwise separable convolution, and an activation function; and then input to the selective scanning module for feature extraction; Feature merging, the data after feature extraction is layer normalized and element-wise multiplied with the output of the first branch to merge the two paths; Linear layer and residual connection, use linear layer to mix features and combine this result with residual connection to form the output of VPU; 2) Use a small target detection enhancement pyramid structure to enhance the small target detection capability; the small target detection enhancement pyramid structure includes a spatial-depth convolution module and a feature enhancement-fusion module; the P2 layer features of the backbone network are processed by the spatial-depth convolution module to retain more fine-grained features, and the processed features are fused with the output features of the P3 layer and input into the feature enhancement-fusion module to enhance the feature extraction and fusion capabilities, thereby more effectively capturing and identifying defects; 3) Use the improved CIoU with scaling factor and auxiliary bounding box as the loss function to adapt to defects of different shapes and sizes and improve the localization effect of the detection box; Step 4: Control the sensor to collect data and upload the sensor measurement data to the industrial computer; the industrial computer outputs the defect data in the collected data through the improved YOLOv8 model and issues an alarm.
2. A turnout detection method based on computer vision according to claim 1, characterized in that: The specific method of feature extraction by the selective scanning module is: Given the input features, the output features of the selective scanning module are expressed as: Among them, v represents four different scanning directions; the expansion operator The input features are transformed along these four directions Expand to generate a set of one-dimensional sequences ; S6 is the core operator in the VPU block, which is used to perform the expanded features Processing to generate new features ; Merge operator The four processed one-dimensional sequences Merge into a new complete 2D feature map .
3. A turnout detection method based on computer vision according to claim 1, characterized in that: The spatial-depth convolution module in step 3 consists of a spatial-depth layer and a non-strided convolution layer, specifically: First, the feature map is input into the space-depth layer, which divides the feature map into four sub-feature maps along the x and y directions. The mapping method is: in, Indicated in Row and Sub-feature maps of columns; A feature map representing the input; and Represents the index in the horizontal and vertical directions respectively; It is expressed as step length; Represents the scaling factor, which is used to determine the size of the sub-feature map; Subsequently, the sub-feature maps are concatenated together to obtain an intermediate feature map, which reduces the spatial size of the feature map and increases the feature information in the channel dimension; Finally, the intermediate feature maps are sent to the non-strided convolutional layer to obtain the final feature maps.
4. A turnout detection method based on computer vision according to claim 1, characterized in that: The specific process of the feature enhancement-fusion module in step 3 is as follows: In the feature enhancement-fusion module, the input features are divided into two parts; one part is processed by the full-core module and then fused with the unprocessed feature map in the other part; Among them, the full-core module includes global branch, macro branch and micro branch; in the global branch, the feature representation of the track area is enhanced through the dual-domain channel attention module and the frequency-based spatial attention module; the macro branch uses three large-scale deep convolutions of 63×1, 63×63, and 1×63 to enhance the recognition ability of medium and large-scale targets; the micro branch extracts local features through 1×1 deep convolution operation to ensure sensitivity to small-scale defects; after the input information is transmitted and accumulated through the three branches, the detection ability of small-scale defects is enhanced.
Citation Information
Patent Citations
Insulator pollution flashover rapid detection method, system and equipment
CN116503396A
Visual inspection method for small-size bearing
CN117523245A