A Small Target Detection Method for Orbital Obstacles with Regional Pre-Search
Through area pre-search and improved YOLOX-S network, the accuracy and false alarm problems of small and medium-sized target obstacle detection in rail transit are solved, and high-precision and real-time obstacle detection are achieved.
Patent Information
- Application Number
- CN202211296306.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-10-21
AI Technical Summary
The prior art is difficult to effectively detect small target obstacles in rail transit scenarios, and it is easy to generate false alarms in complex contexts.
The small object detection method of orbital obstacles with regional pre-search is adopted to perform small object detection through the division of the region of interest, semi-distortion scaling reconstruction algorithm and the improved YOLOX-S network, including the use of the Dilated_Block residual structure and the ASFF-CBAM feature fusion structure to improve detection accuracy and reduce false alarms.
It realizes high-precision detection of small orbital target obstacles, reduces false alarms, and can detect small long-distance targets in real time in real orbital scenarios, which has good application value.
Smart Images

Figure CN115661786B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real-time obstacle detection, and particularly to a small target detection method for track obstacles with regional pre-search. Background Art
[0002] Railway transportation, as the most important transportation mode in China, undertakes the mission of most freight transportation and personnel transportation. Due to the characteristics of high speed and long braking distance in railway operation, the problem of train operation safety has always been difficult to avoid. Especially for the detection and prevention of intrusion behavior, under the existing technical conditions, it is difficult to effectively detect intrusion behavior and combine the braking ability of trains.
[0003] Based on existing different hardware facilities, the research methods for track scene obstacles can be divided into obstacle detection based on lidar and obstacle detection based on visual images. Obstacle detection based on lidar relies on on-vehicle radar to obtain data of target objects in front of the train and constructs a regional model to identify obstacles. Before the popularization of deep learning technology, obstacle detection methods based on visual images generally used algorithms such as Hough transform and ViBe extractor for obstacle detection. With the continuous development of deep learning technology, some excellent object detection algorithms have replaced traditional image processing methods and become the mainstream methods for obstacle detection. Compared with the obstacle detection method based on lidar, the obstacle detection method based on visual images can better balance the real-time performance and accuracy of detection and has good application prospects.
[0004] However, there are still two key problems to be solved in actual application scenarios: (1) Obstacles in rail transit often appear in the form of small and medium-sized targets, so the detection accuracy of small and medium-sized targets is crucial; (2) When the algorithm detection accuracy is relatively high, it is easy to misidentify objects in non-track areas as obstacles in track areas and generate false alarms. Summary of the Invention
[0005] The purpose of the present invention is to provide a small target detection method for track obstacles with regional pre-search, which is used to solve the problems that small target obstacles in rail transit scenarios are difficult to detect and the detection background is prone to interference and false alarms.
[0006] The technical solution adopted by a small target detection method for track obstacles with regional pre-search disclosed by the present invention is as follows:
[0007] A small target detection method for track obstacles with regional pre-search includes the following steps:
[0008] S1. Region pre-search and region of interest division stage:
[0009] Determination of the region of interest: The relevant region centered on the track in the train's forward direction is used as the region of interest. The labelimage software is used to construct the region of interest and form the final track region of interest dataset. The target detection algorithm can be used to train on this dataset to infer the track region of interest;
[0010] S2. Reconstruction stage of the region of interest algorithm:
[0011] Semi-distortion scaling reconstruction algorithm: First, calculate the difference between the existing size and the expected size of the image and take half of it as the upsampling expansion size. At the same time, calculate the critical distortion size through the representative size of small targets generated by K-Mean clustering. The critical distortion size depends on the proportional relationship between the representative small target size and the cropping size. If the upsampling size is less than the critical distortion size, the strategy of first using bilinear interpolation for distortion expansion and then using the Letterbox algorithm is selected; otherwise, the Letterbox algorithm is directly used;
[0012] S3. Small target detection stage of the region of interest:
[0013] Selection and improvement of the detection network: The YOLOX-S, the most lightweight detection network with the highest accuracy in the YOLO series, is selected as the basis for the detection network, aiming at the detection accuracy of small target objects;
[0014] Improvement of the backbone network: The Dilated_Block structure improves all the residual structures in the YOLOX-S backbone network. There are 4 layers of residual structures in the YOLOX-S network. Since the Dilated_Block residual structure can provide sufficient receptive fields, the last layer of the residual structure is removed in the improved network. At the same time, the improved network relies on a separate convolution for downsampling, and the residual structure does not change the width and height of the input feature map;
[0015] Adaptation between algorithms: For the region pre-search task, an image with an input size of 416×416 is determined as the input. The small input size can quickly perform network inference. Considering the aspect ratio of the track distribution, the semi-distortion scaling algorithm is used to reconstruct the region of interest size to 480×640. Finally, the image with a size of 480×640 after reconstruction is input into the small target detection network for obstacle detection. The entire network is finally named RPSNet.
[0016] As a preferred solution, in S1, in order to enable the detection network to infer the region of interest (ROI) of the track more quickly, a highly efficient fast residual structure composed of depthwise separable convolutions is proposed. This structure uses an upsampling structure on the main branch to convolutionally transform the feature map into a high-dimensional channel and uses depthwise separable convolutions for feature extraction. On the residual branch, after adjusting the channels using 1×1 convolutions, it is fused with the main branch. First, this structure is used to replace the residual structure in the YOLOv4-Tiny backbone network. Second, the decoding prediction of YOLOv4-Tiny is improved to AnchorsFree decoding prediction. The improved network changes the output of the picture to output four parameters, that is, the position coordinates of the detected object are redefined as (x, y, w, h), laying a foundation for subsequent image reconstruction.
[0017] As a preferred solution, according to the requirements of ROI division, the Labimage software is used to construct a track ROI dataset, and on this basis, the training division of the ROI is carried out; according to the depthwise separable convolution, a highly efficient residual feature extraction structure is proposed. This structure realizes highly efficient feature extraction by convolutionally transforming the feature map into a high-dimensional channel and using depthwise separable convolutions for feature extraction, and based on this structure, the improvement of the YOLOv4-Tiny backbone network is realized. The input features of this structure Figure X ∈R C×H×W and the output feature map Y ∈ R C′×H′×W′ The relationship can be expressed by the following formula:
[0018]
[0019] where C 1×1 (·) represents a standard convolution block with a convolution kernel size of 1, and D 3×3 (·) represents a depthwise separable convolution block with a convolution kernel size of 3. represents the channel dimension addition operation, and MP(·) represents the max pooling operation.
[0020] The AnchorsFree decoding network is used to adapt the output of the improved network, and at the same time, the form of the network output prediction picture is changed to output the ROI coordinates, forming the YOLOv4-ASNet network.
[0021] As a preferred solution, the concept of image semi-distortion is introduced. This algorithm adds semi-distortion augmentation on the basis of the Letterbox algorithm. Compared with the original algorithm, it can augment more image detail information while maintaining the original data distribution of the picture. The calculation formula of this algorithm is as follows:
[0022]
[0023]
[0024]
[0025]
[0026] (lim h ,lim w )=(λ×clu h ,λ×clu w )
[0027] Among them, the function f(x) represents a clustering algorithm.
[0028] As a preferred solution, a detection network YOLOX-STDNet for small targets improved based on YOLOX-S realizes a residual feature extraction structure Dilated_Block that can greatly expand the receptive field of the feature map by designing repeatedly stacked dilated convolutions. In this structure, the HDC design principle is used to design repeatedly stacked dilated convolutions with dilation rates of 1, 2, and 4 respectively, avoiding the sawtooth effect caused by repeatedly stacking dilated convolutions. A spatial channel self-attention mechanism CBAM is introduced in the same rescaling stage of the adaptive feature fusion structure ASFF, realizing an adaptive self-attention feature fusion structure ASFF-CBAM.
[0029] The above two structures are used to improve the backbone network and the feature fusion structure of the YOLOX-S network respectively: the Dilated_Block residual structure is used to replace the residual structure in the original algorithm, and the last layer of the residual structure is removed to realize the improvement of the backbone network. For the feature extraction network, the FPN structure is removed and the ASFF-CBAM structure is directly used, and its channel dimensions are set to {256, 512, 1024} to realize the improvement of the feature extraction network.
[0030] The beneficial effect of a method for detecting small targets of track obstacles with regional pre-search disclosed by the present invention is that: by improving the residual structure of the YOLOv4-Tiny network, it can perform the regional pre-search task more quickly, and according to the constructed image semi-distortion scaling algorithm, the searched region of interest is extracted and the non-track region is suppressed. Finally, the detection of track small target obstacles is realized through the YOLOX-STDNet network improved for small target detection. The invention can effectively detect the detection of long-distance small target obstacles that appear in the real track scene, and can avoid the detection false alarm behavior caused by the complex background, and can perform the detection of obstacles in real time with high precision, having good application value. Description of the Drawings
[0031] Figure 1It is an analysis diagram of different orbital scenarios of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0032] Figure 2 It is a diagram of the division size of the orbital region of interest of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0033] Figure 3 It is a block diagram of the DS_Block residual structure of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0034] Figure 4 It is a block diagram of the Letterbox image scaling algorithm of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0035] Figure 5 It is a block diagram of the semi-distortion image scaling algorithm of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0036] Figure 6 It is a block diagram of the Dilated_Block residual structure of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0037] Figure 7 It is a network structure diagram of the ASFF-CBAM of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0038] Figure 8 It is a general architecture block diagram of the RPSNet algorithm of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0039] Figure 9 It is an effect diagram of the division of the region of interest and the image semi-distortion scaling algorithm of the YOLOv4-ASNet network of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0040] Figure 10 It is a scatter diagram of the dataset and the size distribution of the region of interest of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0041] Figure 11 It is a comparison diagram of the training accuracy between the YOLOX-STDNet network and other similar networks of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0042] Figure 12 It is a comparison experiment diagram between the improved YOLOX-STDNet small target detection network and the classical network of a small target detection method for orbital obstacles with regional pre-search in the present invention.
[0043] Figure 13It is a partial enlarged view of the detection of a small target of an obstacle on a track by the method of regional pre-search of the present invention.
[0044] Figure 14 It is the detection effect diagram of the video frame of the method for detecting small targets of obstacles on a track by regional pre-search of the present invention. Specific implementation manners
[0045] The present invention will be further described and explained below in conjunction with specific embodiments and the accompanying drawings of the specification:
[0046] The present invention adopts the following technical solution as an algorithm for detecting small targets of obstacles on a track by regional pre-search. The implementation steps of this method are as follows:
[0047] Region pre-search and region of interest division stage:
[0048] S1. Determination of the region of interest: Usually, the determination of obstacles in the rail transit scene depends on whether the object is in the relevant region centered on the running track. Therefore, the region pre-search method proposed by the present invention will use this region as the region of interest, and use the rail scene image and the divided region size as shown in Figure 1 , Figure 2 as an example.
[0049] S1-1. Improvement of the deep network: Existing object detection algorithms can be divided into two types according to different detection methods, two-stage object detection networks and single-stage object detection networks. The two-stage object detection network locates the target region and determines the type of the target separately. This type of method has high detection accuracy, but due to the characteristics of the algorithm itself, the real-time performance of this type of method is poor in actual detection applications. Representative methods include Fast-RCNN, Faster-RCNN, etc. The single-stage object monitoring network predicts the target location information and the target category information together and performs decoding and construction. This type of method not only has a fast detection speed but also has high detection accuracy, and has become the mainstream method in the field of object detection in recent years. Representative methods include the YOLO (You Only Look Once, YOLO) series and the SSD (Single Shot MultiBox Detector, SSD) series. Considering the characteristics of regional pre-search, YOLOv4-Tiny, one of the currently fastest detection networks, is selected as the basic detection network, and the improved network is YOLOv4-ASNet (YOLOv4 Area Search Network). At the same time, a highly efficient residual structure (Depthwise Separable Convolution Block, DS_Block) is introduced to improve the main residual structure in YOLOv4-Tiny. This residual structure is as shown in Figure 3 as shown.
[0050] The DS_Block is mainly composed of depthwise separable convolution and is divided into a backbone channel and a residual channel. For an input feature map with an input size of X ∈ R C×H×W for the input feature map, first use a 1x1 standard convolution on the residual channel to reduce the channel dimension from C to C' / 2 dimension to obtain X1 ∈ R C' / 2×H×W waiting for fusion with the backbone channel. First, use a 1x1 standard convolution on the backbone channel to increase the channel dimension of the input feature map from C to λC to obtain X2 ∈ R λC×H×W , where λ is the dimension increase coefficient. Subsequently, use depthwise separable convolution to extract features from the feature map X2, and then use depthwise separable convolution again to reduce the channel dimension from λC to C' / 2 dimension and fuse it with the feature map on the residual channel to obtain X3 ∈ R C′×H×W After that, use a max pooling layer (Max Pooling, MP) to adjust the width and height dimensions of the output feature map to obtain an output feature map with a size of Y ∈ R C′×H′×W′ of the feature map.
[0051] The input features of the DS_Block residual structure Figure X ∈ R C×H×W and the relationship with the output feature map Y ∈ R C′×H′×W′ can be expressed by the formula:
[0052]
[0053] where C 1×1 (·) represents a standard convolution block with a convolution kernel size of 1, and D 3×3 (·) represents a depthwise separable convolution block with a convolution kernel size of 3, represents the channel dimension addition operation, and MP(·) represents the max pooling operation.
[0054] S2. Region of Interest Algorithm Reconstruction Stage:
[0055] S2-1. Image Semi-Distortion Scaling Algorithm: When the picture is input into the YOLOv4-ASNet region pre-search network, it will divide the region of interest related to the track. At this time, it is necessary to extract the region of interest. The extracted region is the region to be detected by the small target detection algorithm in the next stage. However, since the size of the extracted region is small and not uniform, an image reconstruction algorithm is needed for reconstruction. For this reason, an image semi-distortion scaling algorithm is proposed based on the Letterbox algorithm. The Letterbox algorithm is a reconstruction algorithm for image scaling in proportion. Because the reconstructed picture has gray edges like an envelope (Letterbox), it is named the Letterbox algorithm. The flowchart of this algorithm is as Figure 4As shown, the algorithm first calculates the scale factor between the input image and the scaled size, obtains the shrunk size according to the scale factor, and uses the bilinear interpolation method to scale the image to the shrunk size. Then, it calculates the size of the filled gray border and finally scales the image to the required image size.
[0056] S2-2. The semi-distorted image scaling algorithm proposed based on the Letterbox algorithm introduces the concept of semi-distorted size on the basis of the Letterbox algorithm. The flowchart of this algorithm is as Figure 5 shown. The algorithm first obtains the region size (h, w) from the output result of the region pre-search network, calculates the size (dim h , dim w ) that should be upsampled and expanded through formula (2), and then uses formula (3) to obtain the representative size of small targets (clu , clu h , clu w ) according to the size of the small target dataset. Since there is a proportional relationship between the representative size of small targets and the size of the cropped region, finally, the representative size of small targets is multiplied by the scale factor λ to obtain the critical distortion size (lim h , lim w ). When the upsampled expansion size is smaller than the critical distortion size, the algorithm first uses bilinear interpolation to expand the extracted region of interest and then uses the Letterbox algorithm; otherwise, it directly uses the Letterbox algorithm.
[0057]
[0058]
[0059]
[0060]
[0061] (lim h , lim w ) = (λ × clu h , λ × clu w )
[0062] where the function f(x) represents the clustering algorithm.
[0063] S3. Region of interest detection for small target detection stage:
[0064] S3-1. Determination of the Detection Network: Considering the comprehensive performance of small target detection, the lightweight standard detection network YOLOX-S of the YOLOX series is selected as the basic detection network, and the improved network is YOLOX-STDNet (YOLOX SmallTarget Detection Network).
[0065] S3-2. Enhancing Context Information of Small Targets: The receptive field represents the spatial range of the input image corresponding to a unit pixel on the output feature map. In the backbone network of YOLOX-S, high-resolution feature maps are accompanied by small-sized receptive fields. At this time, although the feature maps have relatively complete fine-grained information of small targets, they cannot be fully represented due to the limited size of the receptive field, resulting in a low detection accuracy of the network for small target objects. Therefore, learning feature representations with large receptive fields and high resolutions is beneficial to enhancing the fine-grained feature representations of small targets. Dilated convolution is a type of convolution that can quickly increase the receptive field of the feature map. It realizes skip convolution by introducing a dilation rate parameter into the standard convolution (this parameter defines the spacing between values when the convolution kernel processes data). Thus, a residual structure composed of dilated convolution (Dilated ConvolutionBlock, Dilated_Block) is constructed. Dilated convolution is used to expand the receptive field of the feature map while maintaining the high-resolution feature map to improve the context information of the small target feature map. This residual structure is as Figure 6 shown.
[0066] S3-3. Enhancing Small Target Localization Information: The feature fusion part of the YOLOX-S network uses the PAFPN structure for multi-scale feature map fusion. However, this pyramid-style feature fusion method will have a certain interference on small target features during prediction training. This inconsistency will interfere with the gradient calculation during the training process and reduce the effectiveness of the feature pyramid. The Adaptive Spatial Feature Fusion (ASFF) structure can adaptively learn the weight bias after fusion at different scales. When there are more small target objects, the ASFF structure will generate larger weight coefficients to focus on the small target feature map. This adaptive feature fusion method is more conducive to the detection of small target objects.
[0067] The ASFF adaptive feature fusion structure is divided into two steps: identically rescaling and adaptively fusing. In identically rescaling, feature maps of different scales are first rescaled to the same scale and then concatenated and fused to form an initial feature map. In adaptively fusing, the initial feature map is passed through a softmax operation to calculate the weight coefficients α, β, γ of different scales, which are then incorporated into the initial feature map to form the final fusion output. To enhance the feature representation of the initial feature map after identically rescaling, a channel-spatial attention mechanism (Convolutional Block Attention Module, CBAM) can be added before forming the initial feature map to form an adaptive attention feature fusion structure ASFF-CBAM. This can make full use of CBAM to separately learn the target positions that need to be focused on the channel and spatial axes, achieving the purpose of enhancing the target location information of each scale. The algorithm block diagram of this structure is as shown in Figure 7 shown.
[0068] The algorithm of the present invention mainly uses a deep neural network constructed with the Pytorch framework and a deep learning experimental platform built with the Ubuntu system, and cooperates with a self-built dataset of regions of interest on the track and a self-built dataset of small target obstacles on the track to verify the effectiveness of the algorithm. Among them, the dataset of regions of interest on the track is made by intercepting driving pictures from the train's driving recorder, which contains 2117 pictures of track regions in 7 different scenarios such as wilderness, snowfield, mountain area, and city. A part of the dataset of small target obstacles on the track comes from track pedestrian data taken with a telephoto camera within a distance of 2000 meters on the track, and the other part consists of small target data related to the track collected from the network, including a total of 5329 pictures in 5 categories. The training set and validation set of both datasets are divided according to a ratio of 0.9:0.1.
[0069] The algorithm of the present invention uses the general evaluation index of object detection, Mean Average Precision (MAP), to measure the detection performance of the model, and uses the average detection precision of small targets MAP defined under the COCO dataset small as the evaluation index for small target object detection. The inference speed of the model is measured by frames per second (FPS), and the calculation is as follows:
[0070]
[0071] MAP small = area < 32 2 / px
[0072]
[0073] AP represents the detection accuracy of a certain category, and MAP represents the average detection accuracy of all categories. px is in pixel units. MAPsmall represents the average detection accuracy of all objects smaller than 32*32 pixel size based on the MAP metric. FPS represents the number of pictures that the model can infer per second, and sec represents the time taken by the network model to infer one picture.
[0074] The method of the present invention realizes small target detection of track obstacles, mainly including three parts: region of interest division, image algorithm reconstruction, and small target object detection, as Figure 8 shown in the overall architecture diagram of the algorithm. The specific implementation process of the algorithm is as follows:
[0075] Region of interest division stage:
[0076] S1: Randomly divide the prepared track region of interest dataset into a training set and a test set according to a ratio of 0.9:0.1. Then divide the training set into training use and validation use according to a ratio of 0.9:0.1.
[0077] S1-1: Improve the residual structure of the backbone network of the YOLOv4-Tiny network using the DS_Block residual structure. The DS_Block residual structure is as Figure 3 shown. This structure can quickly and effectively extract feature maps. Then set the training parameters for the improved network YOLOv4-ASNet. Set the total number of training loops to 200, batch size to 16, initial learning rate to 0.0001, and use the SGD optimization strategy to optimize the training parameters.
[0078] S1-1: Conduct experimental verification and result analysis on the network before and after improvement. First, determine the network performance benchmark through experiments. Use the YOLOv4-Tiny network to train on the track region of interest dataset to obtain a MAP of 93.87% and a frames per second of 101.76 FPS. Based on the performance benchmark of the YOLOv4-Tiny network, experimentally select the dimensionality increase coefficient λ of the DS_Block residual structure in the improved network YOLOv4-ASNet of the present invention.
[0079] Table 1 Performance indicators under different dimensionality increase coefficients
[0080] Dimensionality increase coefficient λ Average detection accuracy MAP / % Number of frames transmitted per second / FPS 3C 91.53 121.2 4C 94.92(+3.39) 117.7(-3.5) 5C 95.43(+0.51) 97.7(-20.0) 6C 95.62(+0.19) 84.7(-13.0)
[0081] Table 1 lists the performance indicators of the improved network YOLOv4-ASNet of the present invention when different dimension-increasing coefficients are selected. Experiments have found that when the dimension-increasing coefficient λ = 4C (C is the feature map channel dimension of the input DS_Block residual structure), the detection accuracy of the YOLOv4-ASNet network is similar to that of the YOLOv4-Tiny network before improvement, but the detection speed is higher. At the same time, experiments show that while a higher dimension-increasing coefficient brings an improvement in accuracy, the increased number of parameters also leads to a decrease in the inference speed of the model. Therefore, the dimension-increasing coefficient λ = 4C is selected to construct the YOLOv4-ASNet network to balance detection accuracy and speed.
[0082] Table 2 Comparison of Performance Indicators between YOLOv4-ASNet Network and Lightweight Networks
[0083]
[0084]
[0085] Table 2 shows the comparative experiments between the YOLOv4-ASNet network and classical lightweight networks. Thanks to the powerful inference speed provided by YOLOv4-Tiny, the detection speed of the improved network is the fastest compared with other similar networks, achieving the purpose of quickly dividing the region of interest. Although the detection accuracy of the improved network is slightly lower than that of the other two network series, sacrificing a small part of the accuracy to bring an improvement in the inference speed is beneficial to the construction of the overall algorithm.
[0086] Figure 9 This is the effect diagram of the regional pre-search algorithm of the present invention for dividing the region of interest of the track and the reconstruction effect diagram of the semi-distortion algorithm of the present invention. It can be seen from the experimental results that for the track regions of different sections, the method of the present invention can accurately divide the region of interest of the track on which the train is running, and after reconstruction by the semi-distortion algorithm, the image only retains the region of interest, reducing the search burden of the small target detection network in the next stage.
[0087] S2. Reconstruction stage of the semi-distortion scaling algorithm:
[0088] S2-1: To obtain the relationship between the small target size and the cropping area, first randomly sample 130 groups of data for the calibrated small target box size in the dataset and the region size extracted by the YOLOv4-ASNet network respectively, as Figure 10Most of the small target sizes in the dataset shown are within 90x90 pixels, and the sizes of the extracted regions of interest are mostly around 500x350 pixels. Then, the k-mean clustering algorithm is used to generate two groups of representative data, namely the representative size of small targets (44, 37) and the representative size of the cropped area (543, 337). From these two groups of data, it can be found that there is a certain proportional relationship between the size of small target objects on the cropped area. Based on the above proportional relationship, the critical distortion size for expansion is defined as twice the size of small target objects, and a semi-distortion algorithm is designed based on this.
[0089] S2-2: Similarly, select the orbital regions in different scenarios to test the effectiveness of the semi-distortion scaling algorithm. The effect diagrams are as Figure 4 shown.
[0090] S3. Small target obstacle detection stage:
[0091] S3-1: Randomly divide the made orbital small target obstacle dataset into a training set and a test set according to a ratio of 0.9:0.1. Then, divide the training set into a training set and a validation set according to a ratio of 0.9:0.1.
[0092] S3-2: For the backbone network of the YOLOX-S network, use the Dilated_Block residual structure to improve the first, second, and third extraction structures, and remove the fourth residual extraction structure. The receptive fields of each layer provided by the backbone network before and after improvement are shown in Table 3.
[0093] Table 3 Comparison of receptive fields before and after improving the backbone network
[0094]
[0095] S3-3: Use the ASFF-CBAM structure (as Figure 7 shown) to improve the feature fusion part of the YOLOX-S network on the basis of S3-1. Set the training parameters for the finally improved YOLOX-STDNet network based on YOLOX-S for training. Use the Mosaic+mixup method to increase the number of small target samples during training, and adopt the training strategy of iterating 300 times with the Adam+SGD optimization method. Among them, the momentum of Adam is 0.5 in the first 200 iterations, and the momentum of SGD is 0.9 in the last 100 iterations. The initial learning rate is 0.001.
[0096] S3-4: Conduct experimental verification and result analysis on the network before and after improvement. To test the effectiveness of the improved small target detection network of the present invention, it is compared with the train obstacle detection methods of the YOLOX-S network, YOLOX-L network, FE-YOLO network, and YOLOv4-SE network before improvement under the same experimental conditions. The comparison results of the training accuracy rise curves of different algorithms trained 300 times in the same experimental environment are as Figure 11 shown.
[0097] As Figure 8 can be seen: The YOLOX series networks and their improved networks perform the best in terms of training convergence speed. Followed by the FE-YOLO network. Among them, the YOLOX-L network has the highest average detection accuracy, followed by the improved network of the present invention and the FE-YOLO network. The YOLOX-L network is one of the large detection networks in the YOLOX series, and the number of network parameters is 6 times that of the YOLOX-S network. As can be seen from Table 4, the highest average detection accuracy of this network reaches 82.81%, and the small target detection accuracy also reaches 24.6%, but the detection speed is only 26.5 FPS. The FE-YOLO network is an improved track obstacle detection algorithm based on the YOLOv4 network. According to the literature description, we use K-mean clustering and SE-Net attention module to improve the network. The average detection accuracy of this network on the track small target dataset reaches 80.10%, and the small target detection accuracy reaches 21.3%. The YOLOv4-SE network uses group convolution to form the backbone network and the attention module built by SSE-Net to construct a track obstacle detection network. Based on the computational advantage of group convolution, this network reaches a detection speed of 51.8 FPS, but the average detection accuracy is only 77.80%.
[0098] Before the improvement of the present invention, the average detection accuracy of the YOLOX-S network reached 79.61%, but the small target detection accuracy was only 20.3%. After the improvement of the present invention, the YOLOX-STDNet network benefits from the fact that the improved network can provide a rich receptive field on the small target feature map and subsequent feature fusion optimization on the small target feature map, which improves the small target detection accuracy to 25.0%. At the same time, the reduction of the residual edge of the residual structure reduces the influence and reduces the network training convergence speed while also increasing the detection speed to 47.8 FPS.
[0099] Table 3-1 Comparison of performance indicators between the improved network of the present invention and similar detection networks
[0100]
[0101] The detection results of intrusion objects in different distance track scenarios by different algorithms are as Figure 12 and Figure 13As shown in the figure. It can be seen from the experimental results that for the orbital obstacles within a short distance, all five algorithms can correctly detect them and achieve the basic detection function. However, for the obstacles at medium and long distances, except for the YOLOX-L network and the improved network, the other networks all have missed detection cases. The YOLOX-S network fails to detect pedestrians at the turnout intersection and cars in the distance. In the fifth image, the orbital pedestrians at the farthest distance are not detected. The detection effect of the FE-YOLO network is slightly better than that of the YOLOX-S network, but for the scene where a car and a pedestrian overlap in the fourth image, only the car is detected. The YOLOv4-SE network has the fastest detection speed in the comparative experiment and can detect orbital pedestrians of large and medium sizes, but there are missed detections for distant objects and objects in overlapping scenes. The detection effects of the YOLOX-L network and the improved network of the present invention are similar. Whether it is the detection of small target objects at a long distance or objects in overlapping scenes, there are no missed detection cases. However, the improved algorithm of the present invention achieves a faster detection speed while maintaining a detection accuracy similar to that of the YOLOX-L network.
[0102] After analyzing the comparative performance of different detection networks, to verify the effectiveness of each module of the improved YOLOX-STDNet network of the present invention, based on the YOLOX-S network, a replacement comparative experiment was carried out. The experimental results are shown in Table 4. The experimental process and result analysis are as follows:
[0103] 1) Replace the residual structure in the backbone network of the YOLOX-S network with the Dilated_Block residual module with dilated convolution. After replacement, the MAP of the small target detection accuracy of the network small increased by 1.4%, and the average detection accuracy MAP index increased by 1.47%;
[0104] 2) Replace the PAFPN feature fusion structure in the YOLOX-S network with the ASFF adaptive feature fusion structure. At this time, the network can generate weight parameters biased towards small target detection according to the number of small target samples. The small target detection accuracy of the network increased by 0.9%. Since the number of parameters of ASFF is less than that of the PAFPN structure, the detection speed increased slightly;
[0105] 3) On the basis of 2), replace the ASFF structure with the ASFF-CBAM feature fusion structure. At this time, the small target detection accuracy of the network is improved by more than 2.0% compared with the YOLOX-S, which is stronger than the ASFF structure. This shows that strengthening the intensity of feature fusion and improving the data distribution during fusion can improve the detection effect of the network. At the same time, the increased number of parameters of the ASFF-CBAM structure also reduces the detection speed of the network;
[0106] 4) The YOLOX-S network adopts the construction method of 1) + 2), allowing the backbone network to output high-resolution feature maps with large receptive fields, and uses an adaptive feature fusion structure ASFF. Experimental results show that feature maps with large receptive fields are conducive to the feature fusion structure to fuse efficient output feature maps. The average detection accuracy of the network at this time reaches 81.59%, and the small target detection accuracy reaches 23.5%, which is close to the YOLOX-L network;
[0107] 5) The YOLOX-S network adopts the construction method of 1) + 3). Compared with 4), the more powerful feature fusion also improves the small target detection accuracy to 25.0%;
[0108] Table 4 Control experiment design
[0109]
[0110] In order to further verify the effectiveness of the method proposed in this invention in real-world scenarios, a video shot with a telephoto camera was selected for testing. Figure 14 When the camera focal length is small (such as Figure 14 (a)) The track area divided by the algorithm is larger. As the camera focal length increases (e.g. Figure 14 (d)) and the visual distance becomes longer, the track area divided by the network gradually focuses on the area of interest centered on the railroad track, and all intrusive objects in this area can be detected. For example, the pedestrian crossing the track appears between frames 885 and 1,125 of the video. When the pedestrian is not in the area of interest of the track, the object is not detected due to the inhibitory effect of the algorithm of the present invention. When the pedestrian is crossing the track, the algorithm detects that the object is in the area of interest of the track and can detect it correctly. Similarly, based on the experiment, the actual inference time of each image is 0.0221s, and the calculated inference speed of each image is about 45.1FPS, which is higher than the detection speed of 40.5FPS per second of YOLOX-S, and can meet the application test in actual scenarios.
[0111] In the above scheme, the present invention provides a method for detecting small track obstacles with regional pre-search, which improves the residual structure of the YOLOv4-Tiny network so that it can perform regional pre-search tasks more quickly, and extracts the searched area of interest and suppresses the non-track area according to the constructed image semi-distortion scaling algorithm, and finally realizes the detection of small track obstacles through the improved YOLOX-STDNet network for small target detection. The invention can effectively detect small target obstacles at a distance that appear in a real track scene, and can avoid false detection caused by complex backgrounds, and can detect obstacles in real time with high precision, and has good application value.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for detecting small targets of orbital obstacles by regional pre-search, characterized in that, It includes the following steps: S1. Region pre-search and ROI division stage: Determination of the region of interest: The relevant area centered on the track in the train's forward direction is used as the region of interest. The labelimage software is used to construct the region of interest, and the final track region of interest dataset is formed. The target detection algorithm is used to train and infer on this dataset to obtain the track region of interest; S2. ROI algorithm reconstruction stage: Semi-distortion scaling reconstruction algorithm: First, calculate the difference between the existing size and the expected size of the image and take half of it as the upsampling expansion size. At the same time, calculate the critical distortion size through the representative size of small targets generated by using K-Means clustering. The critical distortion size depends on the proportional relationship between the representative small target size and the cropping size. If the upsampling size is smaller than the critical distortion size, then first use bilinear interpolation for distortion expansion and then use the Letterbox algorithm strategy; otherwise, directly use the Letterbox algorithm; S3. ROI detection of small targets stage: Selection and improvement of the detection network: The detection network selects the lightweight detection network YOLOX-S with the highest accuracy in the YOLO series as the basis, and improves the network to improve the detection accuracy of small target objects; Improvement of the backbone network: The Dilated_Block structure improves all the residual structures in the YOLOX-S backbone network. There are 4 layers of residual structures in the YOLOX-S network. Since the Dilated_Block residual structure can provide sufficient receptive fields, the last layer of the residual structure is removed in the improved network. At the same time, the improved network relies on a separate convolution for downsampling, and the residual structure does not change the width and height of the input feature map; Adaptation between algorithms: For the region pre-search task, an image with an input size of 416×416 is determined as the input. The small input size can quickly perform network inference. Considering the aspect ratio of the track distribution, the semi-distortion scaling algorithm is used to reconstruct the region of interest size to 480×640. Finally, the image with a reconstructed size of 480×640 is input into the small target detection network for obstacle detection. The entire network is finally named RPSNet.
2. The method for detecting small targets of track obstacles with regional pre-search according to claim 1, characterized in that In S1, to enable the detection network to more quickly infer the track region of interest, a highly efficient fast residual structure composed of depthwise separable convolutions is proposed. This structure uses an upsampling structure on the backbone side to convolve the feature map to a high-dimensional channel and uses depthwise separable convolutions for feature extraction. On the residual side, after using a 1×1 convolution to adjust the channels, it is fused with the backbone side. First, use this structure to replace the residual structure in the YOLOv4-Tiny backbone network. Second, improve the decoding prediction of YOLOv4-Tiny to AnchorsFree decoding prediction. The improved network changes the output image to output four parameters, that is, the position coordinates of the detected target are redefined as (x, y, w, h), laying a foundation for subsequent image reconstruction.
3. The method for detecting small targets of track obstacles with regional pre-search according to claim 2, wherein According to the requirements of region of interest (ROI) division, a dataset of track ROIs was constructed using Labimage software, and ROI training division was performed on this basis. A high-efficiency residual feature extraction structure was proposed based on depthwise separable convolution. This structure achieved high-efficiency feature extraction by convolving the feature map into high-dimensional channels and using depthwise separable convolution for feature extraction, and an improvement to the YOLOv4-Tiny backbone network was realized based on this structure. The input feature map and the output feature map are related as expressed by the following formula: Among them represents a standard convolution block with a convolution kernel size of 1, represents a depthwise separable convolution block with a convolution kernel size of 3, represents the addition operation of channel dimensions, represents the max pooling operation; Output adaptation is performed using AnchorsFree decoding prediction, and the network outputs the coordinates of the region of interest to form the YOLOv4-ASNet network.
4. The small target detection method for track obstacles with regional pre-search according to claim 1, wherein, A concept of image semi-distortion is introduced. This algorithm adds semi-distortion augmentation on the basis of the Letterbox algorithm. Compared with the original algorithm, it can augment more image detail information while maintaining the original data distribution of the picture. The calculation formula of this algorithm is as follows: Among them, the function f(x) represents the clustering algorithm.
5. The small target detection method for track obstacles with regional pre-search according to claim 1, characterized in that, YOLOX-STDNet, a detection network for small targets improved based on YOLOX-S, realizes a residual feature extraction structure Dilated_Block that can greatly expand the receptive field of the feature map by designing repeatedly stacked dilated convolutions. In this structure, the HDC design principle is used, and the dilated rates of the repeatedly stacked dilated convolutions are 1, 2, and 4 respectively to avoid the sawtooth effect caused by repeatedly stacking dilated convolutions. The spatial channel self-attention mechanism CBAM is introduced in the same rescaling stage of the adaptive feature fusion structure ASFF, and the number of CBAM channels is set to {256, 512, 1024} respectively, realizing an adaptive self-attention feature fusion structure ASFF-CBAM. The above two structures are used to improve the backbone network and the feature fusion structure of the YOLOX-S network respectively: the Dilated_Block residual structure is used to replace the residual structure in the original algorithm, and the last layer of the residual structure is removed to realize the improvement of the backbone network. For the feature extraction network, the FPN structure is removed and the ASFF-CBAM structure is directly used, and its channel dimensions are set to {256, 512, 1024} respectively to realize the improvement of the feature extraction network.