High-speed rail external environment anomaly detection method based on combination of Wavelet Pool, SA and YOLOv8
By combining WaveletPool, SA and YOLOv8 methods, the YOLOv8 neural network structure is improved, solving the problems of missed detection of small targets and missed detection of similar shapes in the external environment of high-speed rail, and achieving efficient and accurate abnormal detection effects.
Patent Information
- Application Number
- CN202411944496.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
The existing video security systems have problems such as missing targets, missed targets in the external environment of high-speed rail, missed detection performance caused by occlusion, and insufficient robustness in dynamic environments.
Using the high-speed rail external environment abnormality detection method based on WaveletPool, SA and YOLOv8, the YOLOv8 neural network structure is improved, and a network model that can effectively detect these targets is obtained by constructing a picture data set of pedestrians, animals, transportation, etc.
It improves the accuracy of small object detection, reduces the error detection rate of targets similar in shape, enhances the robustness of complex environments, and achieves efficient and accurate abnormal detection.
Smart Images

Figure CN120014562A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video security, and in particular relates to a method for detecting anomalies in the external environment of a high-speed railway based on the combination of WaveletPool, SA and YOLOv8. Background Art
[0002] Existing anomaly detection systems usually monitor the target area in real time by installing devices such as cameras or monitors.
[0003] When an abnormal situation occurs, the anomaly detection system can detect and record it in time. The anomaly detection system also uses advanced target detection algorithms, such as traditional Haar feature classifiers and detection algorithms based on convolutional neural networks. These algorithms can help the system identify target objects more accurately and improve the accuracy and efficiency of detection. However, traditional vision-based target detection algorithms have problems such as poor region selection strategies, poor robustness of manual feature extraction, poor generalization ability, and susceptibility to environmental interference.
[0004] Although the target detection algorithm based on convolutional neural network can automatically extract target features and has the advantages of strong generalization ability and fast processing speed, it still has problems such as missed detection of small targets due to feature map aliasing and false detection of targets with similar shapes in the external environment of high-speed rail where there are occlusions, some targets are too small, and there are many targets with similar shapes. That is, there are problems such as missed detection of small targets such as pedestrians and animals, and false detection between vehicles and objects with similar shapes.
[0005] In summary, existing technologies have problems such as missed detection of small targets, false detection of targets with similar shapes, decreased detection performance due to occlusion, insufficient robustness in dynamic environments, and limited ability to accurately detect multiple targets. Summary of the invention
[0006] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method for detecting abnormalities in the external environment of high-speed rail based on WaveletPool, SA and YOLOv8. By constructing a picture data set of pedestrians, animals, vehicles, etc. in the external environment of high-speed rail, and improving the neural network structure of YOLOv8, a detection network model that can mainly target pedestrians, animals, and vehicles is obtained through training and verification, which solves the problem of missed detection of small targets such as pedestrians and animals in the external environment of high-speed rail and the problem of false detection between vehicles and objects with similar shapes. It has the characteristics of strong small target detection ability, high ability to distinguish targets with similar shapes, and robustness to adapt to complex environments.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is:
[0008] A high-speed rail external environment anomaly detection method based on WaveletPool, SA and YOLOv8, including the following steps
[0009] Step 1: Obtain a dataset of images of the exterior of the high-speed rail, perform preprocessing, and divide the dataset into a training set and a validation set.
[0010] Step 2: Improve the convolution layer of the YOLOv8 model by combining wavelet pooling (WaveletPool), and improve the head network of the YOLOv8 model based on the ShuffleAttention (SA) mechanism to build an improved YOLOv8 network model;
[0011] Step 3: Use the image training set and image verification set to train and verify the improved YOLOv8 network model to obtain a detection network model that can detect pedestrians, animals, and vehicles;
[0012] Step 4: Obtain images of pedestrians, animals, and vehicles with the high-speed rail external environment as the background as the images to be detected, and use the detection network model to predict the images to be detected to obtain the detection results.
[0013] The data set acquisition in step 1 includes the following steps:
[0014] Step (1): Obtain a dataset of images of pedestrians, different kinds of animals, and different kinds of transportation vehicles in the external environment of the high-speed rail by shooting;
[0015] Step (2): Screen the acquired image data set and select images with complex background environments and clear images;
[0016] Step (3): Randomly divide the obtained data set, with one part as the training set and the other part as the validation set.
[0017] The improved YOLOv8 network model is divided into three parts: backbone network, neck network and head network;
[0018] The backbone network is used to extract features from images and output multi-layer feature maps;
[0019] The neck network is used to fuse multiple layers of feature maps to enhance the detection capability of multi-scale targets (especially small targets);
[0020] The head network is used to perform target detection tasks on the fused multi-layer feature maps and output the final result.
[0021] The backbone network includes a picture input module, a first CBS-W module, a second CBS-W module, a first C2f module, a third CBS-W module, a second C2f module, a fourth CBS-W module, a third C2f module, a fifth CBS-W module, a fourth C2f module, and an SPPF module connected in sequence;
[0022] The output of the second C2f module, the output of the fifth CBS-W module, and the output of the SPPF module are all connected to the neck network.
[0023] The image input module is responsible for preprocessing the input image, including resolution adjustment and normalization operations, to ensure that the input data is compatible with the network architecture;
[0024] The processed image is input as a tensor into the CBS-W module, and the local features are extracted through the convolution layer. The features are then normalized through BN, and the features are gradually extracted. The size of the feature map is compressed by downsampling. The C2f module divides the input features into two parts, one of which is directly passed, and the other part is fused after multiple convolution calculations to extract deep features. Then it enters the SPPF module to extract global context information using multi-scale pooling operations.
[0025] The neck network includes a first upsampling module, a first splicing module, a fifth C2f module, a second upsampling module, a second splicing module, a sixth C2f module, a sixth CBS-W module, a third splicing module, a seventh C2f module, a seventh CBS-W module, a fourth splicing module, and an eighth CBS-W module, which are sequentially connected;
[0026] The output end of the SPPF module is connected to the input end of the first upsampling module, the output end of the fifth CBS-W module is connected to the input end of the first splicing module, the output end of the second C2f module is connected to the input end of the second splicing module; the output end of the sixth C2f module is connected to the head network; the output end of the seventh C2f module is connected to the head network; the output end of the eighth C2f module is connected to the head network;
[0027] The upsampling module captures more detail information and restores the feature map output by the SPPF module from low resolution to high resolution. The splicing module combines the upsampled features and the feature map of the shallow layer of the backbone network, further fuses the multi-scale feature map, and splices the feature maps from different resolutions. The fifth C2f module processes the output of the first splicing module, extracts multi-scale features and enhances the feature expression capability. The sixth C2f module processes the output of the second splicing module to further refine the feature representation. The seventh C2f module processes the high-level splicing feature map to extract the final deep-level feature information. The sixth CBS-W module refines the upsampled and spliced feature maps to enhance the core features required for detection. The eighth CBS-W module further purifies the output feature maps to provide high-quality features for the detection head network.
[0028] The head network includes an SA attention module, a first prediction head, a second prediction head and a third prediction head;
[0029] The output end of the SA attention module is connected to the input end of the third prediction head; the input end of the eighth C2f module is connected to the input end of the SA attention module; the output end of the seventh C2f module is connected to the input end of the second prediction head; the output end of the sixth C2f module is connected to the input end of the first prediction head;
[0030] The first prediction head processes the feature map output by the sixth C2f module and predicts the category, location and confidence of the small object;
[0031] The second prediction head processes the feature map output by the seventh C2f module and predicts the category, location and confidence of the medium target;
[0032] The third prediction head processes the feature map output by the SA attention module and predicts the category, location and confidence of large objects.
[0033] The CBS module consists of three parts, including convolution (Conv), batch normalization (BatchNorm), and SiLU activation function; the CBS-W adds WaveletPool operation on the basis of CBS;
[0034] The improvement of CBS includes the following steps:
[0035] Step (1): Based on CBS, that is, after SiLU, a WaveletPool operation is added; WaveletPool is implemented by discrete wavelet transform (DWT) and inverse discrete wavelet transform (IDWT). The one-dimensional DWT and IDWT are as follows:
[0036] x low =Lx,x high =Hx
[0037] x * =L T x low +H T x high
[0038] The upper formula is DWT, the lower formula is IDWT, x low is the low-frequency component of the vector after being split by DWT, x high is the high frequency component, x * Represents x low and x high The vector reconstructed by IDWT, L = {l n-2 ...,l n-2k} T , H = {h n-2 ...,h n-2k} T , l j is the low-pass filter of the orthogonal wavelet, and h j is a high-pass filter, n is the vector length;
[0039] The two-dimensional DWT and IDWT are shown below:
[0040] X ll =LXL T ,X lh =HXL T ,X hl =LXH T ,X hh =HXH T
[0041] WaveletPool is implemented through two-dimensional DWT and IDWT;
[0042] WaveletPooling introduces wavelet transform into the neural network framework to calculate gradients. Further, based on the two-dimensional DWT / IDWT wavelet pooling, a two-dimensional DWT is applied first, and then a two-dimensional IDWT is applied. DWT decomposes the image or feature map into high-frequency detail subbands and low-frequency approximate subbands of the wavelet. The high-frequency subband has X lh , X hl , X hh , the low frequency subband is X ll .
[0043] The SA is a shuffle attention mechanism (ShuffleAttention), and the specific operations are:
[0044] The output feature map X of the eighth C2f module is grouped into n groups along the channel dimension, and the number of channels in each group is c / n, then X is divided into [X1,…,X n ], use X m Refers to;
[0045] The number of channels of each sub-feature is c / 2n, one of which is used to generate the channel attention map, and the other is used to generate the spatial attention map. c 、F gp They are as follows:
[0046] F c =Wx+b
[0047]
[0048] Feature X' after channel attention processing m1 for:
[0049] X' m1 =δ(F c (s))·X m1 =δ(W1s+b1)·X m1
[0050] Where W1 and b1 have sizes of c / 2n×1×1. These two parameters are used for scaling and moving. The feature X' processed by the spatial attention mechanism is m2 for:
[0051] X' m2 =δ(W2·GN(X m2 )+b2)·X m2
[0052] Where W2 and b2 are the functions of W1 and b1, and the size is c / 2n×1×1. Finally, the two attention maps are concatenated to get X' m , the number of channels is the same as the input X m Same as c / n.
[0053] The training and verification of the improved YOLOv8 network model in step 3 includes the following steps:
[0054] Step (1): inputting the pedestrian, animal, and vehicle image training set and image verification set selected in step 1 into the image input module of the improved YOLOv8 network model for standardization;
[0055] Step (2): Using the backbone network process, the pedestrians, animals, and vehicles in the training set are subjected to feature extraction through the convolution layer, and then enter the pooling layer to perform downsampling operations and mean pooling to reduce the amount of calculation while extracting the main features and increase the receptive field. The first and second feature maps are then extracted from the output end of the second C2f module, the output end of the fifth CBS-W module, and the output end of the fourth C2f module.
[0056] Step (3): Use the neck network to upsample, concatenate and extract the first part of the feature map, the second part of the feature map and the feature map extracted by the SPPF module, and obtain the first predicted feature map, the second predicted feature map and the feature map to be extracted by the SA attention module from the first part of the feature map, the second part of the feature map and the feature map extracted by the SPPF module in turn;
[0057] Step (4): using the first prediction head in the head network to perform target prediction on the first prediction feature map; using the second prediction head in the head network to perform target prediction on the second prediction feature map; using the SA attention module of the head network to extract features from the feature map for feature extraction by the SA attention module to obtain a third prediction feature map; using the third prediction head of the head network to perform target prediction on the third prediction feature map;
[0058] Step (5): Repeat the stage training times from step (1) to step (4), and use the pedestrian, animal, and vehicle image verification set to verify the network model combining WaveletPool, SA, and YOLOv8 that has completed the stage training;
[0059] Step (6): Repeat step (5) until the optimal pedestrian, animal, and vehicle detection network model is obtained.
[0060] The step 4 is specifically as follows:
[0061] Step (1): Using cameras installed outside the high-speed railway to take photos, actual images of pedestrians, animals, and various types of transportation vehicles in the high-speed railway environment are extracted;
[0062] Step (2): Use the improved YOLOv8 model to detect and verify the high-speed rail external environment picture, obtain the detection and verification result, and add the high-speed rail external environment picture to the data set in step 1, and then repeat the operation in step 3 to train the data set with the high-speed rail external environment picture added;
[0063] Step (3): Repeat the operations from step (1) to step (2) until the optimal network model for detecting pedestrians, animals, and vehicles in the external environment of the high-speed rail is obtained.
[0064] Beneficial effects of the present invention:
[0065] The present invention obtains image data sets of pedestrians, different kinds of animals, and different kinds of vehicles, and improves the CBS module in YOLOv8 in combination with WaveletPool. The new convolutional layer network can effectively reduce the aliasing phenomenon of the feature map after convolution, and improve the detection accuracy of small targets; based on the SA attention module, the head network of the improved YOLOv8 can more effectively utilize the feature information at different scales, and reduce the network model's misdetection of different kinds of objects with similar shapes. The present invention combines WaveletPool and SA to construct an improved YOLOv8 network model, and trains and verifies the improved YOLOv8 network model using the image training set and image verification set divided by the image data sets of pedestrians, animals, and vehicles, and finally obtains a detection network model that can efficiently and accurately identify pedestrians, animals, vehicles and other targets in the external environment of high-speed rail. When faced with one or more interferences such as small size, long distance, partial occlusion, objects with similar shapes, and slow detection speed, it can effectively overcome them, and the detection results are very reliable. At the same time, a high-speed rail external environment anomaly detection system was built. Using the above-mentioned network model and the detection system, abnormal intrusions of pedestrians, animals, and various types of transportation vehicles in the high-speed rail external environment can be detected in a timely and accurate manner.
[0066] Step 1 can further filter the images of pedestrians, animals and vehicles to reduce the interference of noise, enable the network model to learn more intuitive and accurate features of the target, enhance the generalization ability of the model, and improve the training efficiency. By dividing the data set into a training set and a validation set, the overfitting problem of the model can be prevented, the hyperparameters of the model can be adjusted, and the performance of the model can be further improved.
[0067] The step 2 provides an improved YOLOv8 network model, improves the CBS module of the YOLOv8 backbone network through WaveletPool, adds a WaveletPool operation to the module of the CBS module, and changes the CBS module into a CBS-W module, thereby improving the anti-aliasing capability of the network model feature map and improving the small target detection capability of the network model. The SA attention mechanism is added before the third detection head of the head network, so that the network model can focus on the characteristics of the target during detection, thereby improving the detection accuracy of the model for objects with similar shapes.
[0068] In step 3, based on pedestrians, different types of animals, and different types of vehicles, the network model based on the combination of WaveletPool, SA and YOLOv8 is trained and verified, the optimal network model is determined, and the robustness of the detection network model for pedestrians, animals, and vehicles is enhanced.
[0069] The step 4 uses the measured pictures of the external environment of the high-speed rail for detection verification and retraining, which can improve the robustness of the model for detecting pedestrians, animals, and vehicles in the external environment of the high-speed rail. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a flowchart of the steps of the high-speed rail external environment anomaly detection system based on WaveletPool, SA and YOLOv8 in an embodiment of the present invention.
[0071] Figure 2 It is a schematic diagram of the structure of the YOLOv8 network model combining WaveletPool and SA in an embodiment of the present invention.
[0072] Figure 3 It is a schematic diagram of the structure of CBS-W in an embodiment of the present invention.
[0073] Figure 4 It is a schematic diagram of the WaveletPool structure in an embodiment of the present invention.
[0074] Figure 5 It is a schematic diagram of the SA structure in an embodiment of the present invention. Figure 6 It is a flow chart of the external environment anomaly detection system of the present invention. DETAILED DESCRIPTION
[0075] The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0076] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides a high-speed rail external environment anomaly detection system based on WaveletPool, SA and YOLOv8, comprising the following steps:
[0077] Step 1: Obtain image datasets of pedestrians, different types of animals, and different means of transportation, and preprocess the image datasets to obtain image training sets and image verification sets;
[0078] The step 1 comprises the following steps:
[0079] Step (1): Obtain a data set of pedestrians, different kinds of animals, and different kinds of transportation vehicles; the pedestrians in the data set should be pedestrians in different postures, such as the front, back, side, and occluded; the animals in the data set should be cats, dogs, chickens, cows, sheep, and other animals raised by humans that may appear outside the high-speed rail; the transportation vehicles in the data set should be common transportation vehicles such as cars, trucks, buses, motorcycles, electric vehicles, etc. of various colors, brands, and shapes;
[0080] Step (2): Screen the acquired data set and select pictures with complex background environments and clear pictures. These pictures will be used for training and verification of the YOLOv8 network model.
[0081] Step (3): The obtained data set is randomly divided into two parts in a ratio of 9:1, where the part with a ratio of 9 is used as the training set and the part with a ratio of 1 is used as the validation set.
[0082] Step 2: Improve the convolution layer of the YOLOv8 model by combining wavelet pooling (WaveletPool) to increase the small target detection accuracy of the YOLOv8 model, and improve the head network of the YOLOv8 model based on the ShuffleAttention (SA) mechanism to reduce the false detection rate of the target and obtain an improved YOLOv8 network model;
[0083] like Figure 2 As shown, the improved YOLOv8 network model is divided into three parts: backbone network, neck network and head network;
[0084] The backbone network is used to provide feature extraction from bottom layer to high layer and output multi-layer feature maps;
[0085] The neck network is used to fuse multi-scale feature maps to enhance the detection capability of multi-scale targets (especially small targets);
[0086] The head network is used to perform target detection tasks on the fused features and output the final result.
[0087] The backbone network includes an image input module, a first CBS-W module, a second CBS-W module, a first C2f module, a third CBS-W module, a second C2f module, a fourth CBS-W module, a third C2f module, a fifth CBS-W module, a fourth C2f module, and an SPPF module; the image input module serves as the image input end of the backbone network; the output end of the second C2f module, the output end of the fifth CBS-W module, and the output end of the SPPF module of the backbone network are all connected to the neck network;
[0088] The image input module is responsible for preprocessing the input image, including resolution adjustment and normalization operations, to ensure that the input data is compatible with the network architecture;
[0089] The processed image is input as a tensor into the CBS-W module, local features are extracted through the convolution layer, and then the features are normalized through BN to stabilize the training process; the features are extracted step by step, and the size of the feature map is compressed by downsampling. The C2f module is based on the CrossStagePartial (CSP) idea, which divides the input features into two parts, one of which is directly passed, and the other is fused after multiple convolution calculations, reducing the number of parameters, enhancing gradient flow, extracting deep features, reducing redundant calculations, and maintaining a high level of expression; then entering the SPPF module to extract global context information using multi-scale pooling operations.
[0090] The neck network includes a first upsampling module, a first splicing module, a fifth C2f module, a second upsampling module, a second splicing module, a sixth C2f module, a sixth CBS-W module, a third splicing module, a seventh C2f module, a seventh CBS-W module, a fourth splicing module, and an eighth CBS-W module;
[0091] The output end of the SPPF module is connected to the input end of the first upsampling module, the output end of the fifth CBS-W module is connected to the input end of the first splicing module, the output end of the second C2f module is connected to the input end of the second splicing module; the output end of the sixth C2f module is connected to the head network; the output end of the seventh C2f module is connected to the head network; the output end of the eighth C2f module is connected to the head network;
[0092] The upsampling module captures more detail information and restores the feature map from low resolution to high resolution. The splicing module combines the upsampled features and the shallow feature maps of the backbone network to further fuse the multi-scale feature maps, so that the medium-scale and small-scale targets are strengthened at the same time, and continues to combine the deep global features with the local features, supplement the spatial information and contextual information, and splice the feature maps from different resolutions; the fifth C2f module processes the output of the first splicing module, extracts multi-scale features and enhances the feature expression capability. The sixth C2f module processes the output of the second splicing module to further refine the feature representation. The seventh C2f module processes the high-level splicing feature map and extracts the final deep-level feature information; the sixth CBS-W module: refines the upsampled and spliced feature maps, filters redundant information, and enhances the core features required for detection. The eighth CBS-W module: as the final feature processing module, further purifies the output feature map to provide high-quality features for the detection head network.
[0093] The head network includes an SA attention module, a first prediction head, a second prediction head and a third prediction head; the output end of the SA attention module is connected to the input end of the third prediction head; the input end of the eighth C2f module is connected to the input end of the SA attention module; the output end of the seventh C2f module is connected to the input end of the second prediction head; the output end of the sixth C2f module is connected to the input end of the first prediction head;
[0094] The SA attention module is used to increase the weight of important areas in the feature map and enhance the ability to detect small targets, especially in the complex environment outside the high-speed rail, to reduce missed detection problems. It can also dynamically adjust the feature expression of each pixel based on the relationship between pixels within the feature map, so that the network can more accurately capture the key features of small targets and occluded targets.
[0095] The first prediction head processes the feature map output by the sixth C2f module and predicts the category, location and confidence of the small object;
[0096] The second prediction head processes the feature map output by the seventh C2f module and predicts the category, location and confidence of the medium target;
[0097] The third prediction head processes the feature map output by the SA attention module to predict the category, location and confidence of large targets. At the same time, due to the processing by the SA attention module, the suppression of background interference is also enhanced.
[0098] like Figure 3 As shown in the figure, the CBS-W module is an improvement made on the basis of the CBS module using WaveletPool; the CBS module consists of three parts, including convolution (Conv), batch normalization (BatchNorm), and SiLU activation function; the CBS-W adds the WaveletPool operation on the basis of CBS;
[0099] Features are extracted through convolution (Conv) operations and local receptive field information is aggregated. Batch normalization (BatchNorm) is used to standardize feature maps to accelerate network training, improve the generalization ability of the model, and stabilize the training process. The SiLU activation function introduces nonlinear capabilities to the convolutional features, retains negative information, and improves the model's sensitivity to small feature changes. The feature map processed by Conv+BatchNorm+SiLU retains key low-frequency information through WaveletPool, further compressing the feature map size.
[0100] like Figure 4 The figure shows the structure of WaveletPool, which is implemented by discrete wavelet transform (DWT) and inverse discrete wavelet transform (IDWT). The one-dimensional DWT and IDWT are shown below:
[0101] x low =Lx,x high =Hx
[0102] x * =L T x low +H T x high
[0103] The upper formula is DWT, the lower formula is IDWT, x low is the low-frequency component of the vector after being split by DWT, x high is the high frequency component, x * Represents x low and x high The vector reconstructed by IDWT, L = {l n-2 ...,l n-2k} T , H = {h n-2 ...,h n-2k} T , l j is the low-pass filter of the orthogonal wavelet, and h j is a high-pass filter and n is the length of the vector.
[0104] The two-dimensional DWT and IDWT are shown below:
[0105] X ll =LXL T ,X lh =HXL T ,X hl =LXH T ,X hh =HXH T
[0106] The two-dimensional DWT and IDWT can be implemented because the one-dimensional DWT and IDWT are performed on rows and columns.
[0107] This process is based on the wavelet pooling of two-dimensional DWT / IDWT. First, a two-dimensional DWT is applied, and then a two-dimensional IDWT is applied. DWT decomposes the image or feature map into high-frequency detail subbands and low-frequency approximate subbands of the wavelet. The high-frequency subband is X lh , X hl , X hh , the low frequency subband is X ll IDWT is only used in low-frequency subbands and not in high-frequency subbands. This operation will reduce the resolution of the image to half of the original resolution and filter out the high-frequency information without aliasing, such as Figure 4 .
[0108] like Figure 5 As shown, the SA attention module is a shuffle attention mechanism (ShuffleAttention), which is an attention mechanism that includes both channel attention and spatial attention;
[0109] in, is the dot product operation, C is the concatenation operation, S is the channel shuffle operation, δ(·) is the Sigmoid activation function, X is the input feature map of the output of the eighth C2f module, w, h, c are the length, height, and number of channels of the input data, GN is the normalization operation, Group represents grouping the feature map into n groups along the channel latitude, and the number of channels in each group is c / n, then X is divided into [X1,…,X n ], use X m Refers to; in the figure, each group of sub-features is split into two parts X m1 , X m2 The number of channels of each sub-feature is c / 2n, one of which is used to generate a channel attention map, and the other is used to generate a spatial attention map. c 、F gp They are as follows:
[0110] F c =Wx+b
[0111]
[0112] Feature X' after channel attention processing m1 for:
[0113] X' m1 =δ(F c (s))·X m1 =δ(W1s+b1)·X m1
[0114] The size of W1 and b1 is c / 2n×1×1, and these two parameters are used for scaling and shifting. The feature X' after the spatial attention mechanism m2 for:
[0115] X' m2 =δ(W2·GN(X m2 )+b2)·X m2
[0116] Where W2 and b2 are the functions of W1 and b1, and the size is c / 2n×1×1. Finally, the two attention maps are concatenated to get X' m , the number of channels is the same as the input X m Same as c / n.
[0117] Step 3: Use the image training set and the image verification set to train and verify the improved YOLOv8 network model to obtain a detection network model that can detect pedestrians, animals, and vehicles; in this embodiment, the training parameters are set as follows: the training batch size is 8, the training rounds are 100, and the input image size is 640*640.
[0118] Step (1): input the pedestrian, animal, and vehicle image training set and image verification set selected in step 1 into the image input module of the improved YOLOv8 network model, perform standardization, size adjustment, and channel conversion to adapt to the computing requirements of the deep learning framework;
[0119] Step (2): Use the backbone network process to extract features of pedestrians, animals, and vehicles in the training set through the convolution layer, and enter the pooling layer to perform downsampling operations and mean pooling to reduce the amount of calculation while extracting the main features and increase the receptive field. Then, extract the first part of the feature map from the output end of the second C2f module, the output end of the fifth CBS-W module, and the output end of the fourth C2f module in turn. The feature map has rich spatial detail information and mainly contains low-level local information such as edges and textures. It mainly captures small target features; the second part of the feature map contains both local information and certain global contextual semantics. It mainly captures medium target features. The target feature map to be extracted by the SPPF module has reduced spatial details and contains more abstract semantic features, mainly capturing large target features;
[0120] Step (3): Use the neck network to upsample, concatenate and extract the first part of the feature map, the second part of the feature map and the feature map extracted by the SPPF module, and obtain the first predicted feature map, the second predicted feature map and the feature map to be extracted by the SA attention module from the first part of the feature map, the second part of the feature map and the feature map extracted by the SPPF module in turn;
[0121] Step (4): using the first prediction head in the head network to perform target prediction on the first prediction feature map; using the second prediction head in the head network to perform target prediction on the second prediction feature map; using the SA attention module of the head network to extract features from the feature map for feature extraction by the SA attention module to obtain a third prediction feature map; using the third prediction head of the head network to perform target prediction on the third prediction feature map;
[0122] The first prediction head is used to predict the target of the first prediction feature map. Solve the problem of missed detection of small targets. Pedestrians and animals in the external environment of high-speed rail are often small in size and easily blocked by the background. Use the spatial details of shallow features to enhance the model's perception of small targets. The second prediction head is used to predict the target of the second prediction feature map, taking into account the detection requirements of small and large targets. Predict on the medium-resolution feature map to balance detection accuracy and computational efficiency. Improve the model's ability to detect medium-scale targets in complex scenes. The SA attention module is used to extract features from the feature map. Improve the robustness of scenes with complex backgrounds and easily confused targets. Enhance the model's ability to distinguish occluded targets and targets with similar shapes. The third prediction head is used to predict the target of the third prediction feature map. Solve the problem of false detection of targets with similar shapes (such as vehicles and objects with similar shapes in the background). Improve the detection accuracy of large targets while reducing sensitivity to background noise. Use the features enhanced by the SA attention mechanism to optimize the detection performance in complex scenes.
[0123] Step (5): Repeat the stage training times from step (1) to step (4), and use the pedestrian, animal, and vehicle image verification set to verify the network model combining WaveletPool, SA, and YOLOv8 that has completed the stage training;
[0124] Step (6): Repeat step (5) until the optimal pedestrian, animal, and vehicle detection network model is obtained.
[0125] In this embodiment, some pedestrian, animal, and vehicle image data sets are used as test sets, and the network model is tested based on the test sets. During the test, the YOLOv8 model and the reverse comparison test are compared. The test results are shown in Table 1:
[0126] Table 1
[0127] Model Name Parameter quantity GFLOPS mAP50(%) YOLOv8 11137906 28.7 69.7% Improved YOLOv8 11138098 43.5 73.3%
[0128] In Table 1, mAP50 is the average precision when the intersection-over-union ratio is 0.5, and GFLOPS is the amount of floating-point operations per second. As can be seen from Table 1, the improved YOLOv8 has increased parameters and GFLOPS compared to the YOLOv8 model, indicating that the model is more complex. At the same time, the accuracy has increased and the stability of detection has improved.
[0129] Step 4: Obtain images of pedestrians, animals, and vehicles with the high-speed rail external environment as the background, and use the detection network model to predict the images to be detected to obtain the detection results of the images to be detected;
[0130] In this embodiment, the detection result is a detection picture of pedestrians, chickens, cows, sheep, dogs, cats, cars, buses, trucks, motorcycles, etc. with a target box and confidence. The target box is used to describe the position and size of the target in the detected picture. The target box frames the target. The confidence is a measure of the model's prediction accuracy for the target category to which each detected target box belongs. Person, Cat, Dog, Car, Bus, etc. are the target categories of the detected target boxes, and the subsequent numbers, such as 0.95, 0.73, 0.86, etc., are the confidences corresponding to the target boxes. The detection network model provided by the present invention can detect the specified target more accurately and comprehensively.
[0131] like Figure 6 As shown, this embodiment constructs a high-speed rail external environment abnormality detection system, which will actively detect abnormalities in the images of the cameras installed in the high-speed rail external environment;
[0132] The steps include:
[0133] Step 1: Receive camera images;
[0134] Step 2: Perform target detection on the camera image to determine whether the specified target is detected;
[0135] Step 3: After the designated target is detected, the detection result will be displayed.
[0136] The tested network model is applied to the high-speed rail external environment anomaly detection system.
[0137] Build the operating environment required by the system in Linux or Windows system, such as compiling and installing OpenCV, FFmpeg, installing CUDA and other operating environments;
[0138] Ability to build a complete anomaly detection system and make it run smoothly;
[0139] Build the main program and necessary components of the high-speed rail external environment anomaly detection system;
[0140] Test the operation of the main program and necessary components.
[0141] The high-speed rail external environment anomaly detection system includes a tested network model in the detection method. The model trained in Python is in pt format, which is converted to onnx format or engine model. The converted format is called by C++ and runs faster. The above is only a specific implementation of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered within the protection scope of the present invention.
Claims
1. A high-speed rail external environment anomaly detection method based on WaveletPool, SA and YOLOv8, characterized in that: The following steps are included Step 1: Obtain a dataset of images of the exterior of the high-speed rail, perform preprocessing, and divide the dataset into a training set and a validation set. Step 2: Combine WaveletPool to improve the convolutional layer of the YOLOv8 model, and based on the SA mechanism, improve the head network of the YOLOv8 model to build an improved YOLOv8 network model; Step 3: Use the image training set and image verification set to train and verify the improved YOLOv8 network model to obtain a detection network model that can detect pedestrians, animals, and vehicles; Step 4: Obtain images of pedestrians, animals, and vehicles with the high-speed rail external environment as the background as the images to be detected, and use the detection network model to predict the images to be detected to obtain the detection results.
2. According to claim 1, a high-speed rail external environment anomaly detection method based on WaveletPool, SA and YOLOv8 is characterized in that: The data set acquisition in step 1 includes the following steps: Step (1): Obtain a dataset of images of pedestrians, different kinds of animals, and different kinds of transportation vehicles in the external environment of the high-speed rail by shooting; Step (2): Screen the acquired image data set and select images with complex background environments and clear images; Step (3): Randomly divide the obtained data set, with one part as the training set and the other part as the validation set.
3. According to claim 1, a high-speed rail external environment anomaly detection method based on WaveletPool, SA and YOLOv8 is characterized in that: The improved YOLOv8 network model is divided into three parts: backbone network, neck network and head network; The backbone network is used to extract features from images and output multi-layer feature maps; The neck network is used to fuse multiple layers of feature maps; The head network is used to perform target detection tasks on the fused multi-layer feature maps and output the final result.
4. According to claim 3, a high-speed rail external environment anomaly detection method based on the combination of WaveletPool, SA and YOLOv8 is characterized in that: The backbone network includes a picture input module, a first CBS-W module, a second CBS-W module, a first C2f module, a third CBS-W module, a second C2f module, a fourth CBS-W module, a third C2f module, a fifth CBS-W module, a fourth C2f module, and an SPPF module connected in sequence; The output of the second C2f module, the output of the fifth CBS-W module, and the output of the SPPF module are all connected to the neck network; The image input module is responsible for preprocessing the input image, including resolution adjustment and normalization operations, to ensure that the input data is compatible with the network architecture; The processed image is input as a tensor into the CBS-W module, local features are extracted through the convolution layer, and then the features are normalized through BN, features are gradually extracted, and the size of the feature map is compressed by downsampling. The C2f module divides the input features into two parts, one of which is directly passed, and the other part is fused after multiple convolution calculations to extract deep features, and then enters the SPPF module to extract global context information using multi-scale pooling operations.
5. According to claim 4, a high-speed rail external environment anomaly detection method based on the combination of WaveletPool, SA and YOLOv8 is characterized in that: The neck network includes a first upsampling module, a first splicing module, a fifth C2f module, a second upsampling module, a second splicing module, a sixth C2f module, a sixth CBS-W module, a third splicing module, a seventh C2f module, a seventh CBS-W module, a fourth splicing module, and an eighth CBS-W module, which are sequentially connected; The output end of the SPPF module is connected to the input end of the first up-sampling module, the output end of the fifth CBS-W module is connected to the input end of the first splicing module, the output end of the second C2f module is connected to the input end of the second splicing module; the output end of the sixth C2f module is connected to the head network; the output end of the seventh C2f module is connected to the head network; the output end of the eighth C2f module is connected to the head network; The upsampling module captures more detailed information and restores the feature map output by the SPPF module from low resolution to high resolution. The splicing module combines the upsampling features with the feature map of the shallow layer of the backbone network, further fuses the multi-scale feature maps, and splices the feature maps from different resolutions. The fifth C2f module processes the output of the first splicing module, extracts multi-scale features and enhances the feature expression capability. The sixth C2f module processes the output of the second splicing module to further refine the feature representation. The seventh C2f module processes the high-level splicing feature map to extract the final deep-level feature information. The sixth CBS-W module refines the upsampled and spliced feature map to enhance the core features required for detection. The eighth CBS-W module further purifies the output feature map to provide high-quality features for the detection head network.
6. According to claim 5, a high-speed rail external environment anomaly detection method based on the combination of WaveletPool, SA and YOLOv8 is characterized in that: The head network includes an SA attention module, a first prediction head, a second prediction head and a third prediction head; The output end of the SA attention module is connected to the input end of the third prediction head; the input end of the eighth C2f module is connected to the input end of the SA attention module; the output end of the seventh C2f module is connected to the input end of the second prediction head; the output end of the sixth C2f module is connected to the input end of the first prediction head; The first prediction head processes the feature map output by the sixth C2f module and predicts the category, location and confidence of the small object; The second prediction head processes the feature map output by the seventh C2f module and predicts the category, location and confidence of the medium target; The third prediction head processes the feature map output by the SA attention module and predicts the category, location and confidence of large objects.
7. According to claim 6, a high-speed rail external environment anomaly detection method based on the combination of WaveletPool, SA and YOLOv8 is characterized in that: The improvement of CBS includes the following steps: Step (1): Based on CBS, that is, after SiLU, a WaveletPool operation is added; WaveletPool is implemented by discrete wavelet transform and inverse discrete wavelet transform. The one-dimensional DWT and IDWT are shown as follows: x low =Lx,x high =Hx x * =L T x low +H T x high The upper formula is DWT, the lower formula is IDWT, x low is the low-frequency component of the vector after being split by DWT, x high is the high frequency component, x * Represents x low and x high The vector reconstructed by IDWT, L = {l n-2 …,l n-2k } T , H = {h n-2 …,h n-2k } T , l j is the low-pass filter of the orthogonal wavelet, and h j is a high-pass filter, n is the vector length; The two-dimensional DWT and IDWT are shown below: X ll =LXL T ,X lh =HXL T ,X hl =LXH T ,X hh =HXH T WaveletPool is implemented through two-dimensional DWT and IDWT; WaveletPooling introduces wavelet transform into the neural network framework to calculate gradients; Wavelet pooling based on two-dimensional DWT / IDWT first applies a two-dimensional DWT, then a two-dimensional IDWT. DWT decomposes the image or feature map into high-frequency detail subbands and low-frequency approximate subbands of the wavelet. The high-frequency subband has X lh , X hl , X hh , the low frequency subband is X ll .
8. According to claim 6, a high-speed rail external environment anomaly detection method based on the combination of WaveletPool, SA and YOLOv8 is characterized in that: The SA is a shuffle attention mechanism, and the specific operation is: The output feature map X of the eighth C2f module is grouped into n groups along the channel dimension, and the number of channels in each group is c / n, then X is divided into [X1,…,X n ], use X m Refers to; The number of channels of each sub-feature is c / 2n, one of which is used to generate the channel attention map, and the other is used to generate the spatial attention map. c 、F gp They are as follows: F c =Wx+b Feature X' after channel attention processing m1 for: X' m1 =δ(F c (s))·X m1 =δ(W1s+b1)·X m1 Where W1 and b1 have sizes of c / 2n×1×1. These two parameters are used for scaling and moving. The feature X' processed by the spatial attention mechanism is m2 for: X' m2 =δ(W2·GN(X m2 )+b2)·X m2 Where W2 and b2 are the functions of W1 and b1, and the size is c / 2n×1×1. Finally, the two attention maps are concatenated to get X' m , the number of channels is the same as the input X m Same as c / n.
9. According to claim 8, a high-speed rail external environment anomaly detection method based on the combination of WaveletPool, SA and YOLOv8 is characterized in that: The training and verification of the improved YOLOv8 network model in step 3 includes the following steps: Step (1): inputting the pedestrian, animal, and vehicle image training set and image verification set selected in step 1 into the image input module of the improved YOLOv8 network model for standardization; Step (2): Using the backbone network process, the pedestrians, animals, and vehicles in the training set are subjected to feature extraction through the convolution layer, and then enter the pooling layer to perform downsampling operations and mean pooling to reduce the amount of calculation while extracting the main features and increase the receptive field. The first part of the feature map and the second part of the feature map are then extracted from the output end of the second C2f module, the output end of the fifth CBS-W module, and the output end of the fourth C2f module. Step (3): Use the neck network to upsample, concatenate and extract the first part of the feature map, the second part of the feature map and the feature map extracted by the SPPF module, and obtain the first predicted feature map, the second predicted feature map and the feature map to be extracted by the SA attention module from the first part of the feature map, the second part of the feature map and the feature map extracted by the SPPF module in turn; Step (4): using the first prediction head in the head network to perform target prediction on the first prediction feature map; using the second prediction head in the head network to perform target prediction on the second prediction feature map; using the SA attention module of the head network to extract features from the feature map for feature extraction by the SA attention module to obtain a third prediction feature map; using the third prediction head of the head network to perform target prediction on the third prediction feature map; Step (5): Repeat the stage training times from step (1) to step (4), and use the pedestrian, animal, and vehicle image verification set to verify the network model combining WaveletPool, SA, and YOLOv8 that has completed the stage training; Step (6): Repeat step (5) until the optimal pedestrian, animal, and vehicle detection network model is obtained.
10. According to claim 1, a high-speed rail external environment anomaly detection method based on the combination of WaveletPool, SA and YOLOv8 is characterized in that: The step 4 is specifically as follows: Step (1): Using cameras installed outside the high-speed railway to take photos, actual images of pedestrians, animals, and various types of transportation vehicles in the external environment of the high-speed railway are extracted; Step (2): Use the improved YOLOv8 model to detect and verify the high-speed rail external environment picture, obtain the detection and verification result, and add the high-speed rail external environment picture to the data set in step 1, and then repeat the operation in step 3 to train the data set with the high-speed rail external environment picture added; Step (3): Repeat the operations from step (1) to step (2) until the optimal network model for detecting pedestrians, animals, and vehicles in the external environment of the high-speed rail is obtained.