Unsupervised road obstacle detection method based on reverse distillation
Through the improved reverse distillation network, combined with the multi-scale feature fusion module and the adaptive mask generator module, the problem of lack of abnormal samples and feature domain deviation in road obstacle detection is solved, and high accuracy and high generalization detection effects are achieved.
Patent Information
- Application Number
- CN202510082321.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has problems such as lack of abnormal samples for training, feature domain deviation and overgeneralization of student networks in road obstacle detection, resulting in low detection efficiency and poor generalization.
The unsupervised road obstacle detection method based on reverse distillation is adopted to solve the feature domain deviation and generalization problems through an improved reverse distillation network, including a multi-scale feature fusion module and an adaptive mask generator module.
High-precision and high generalization of road obstacle detection, improved detection accuracy and enhanced generalization ability, which can effectively avoid misjudgment caused by excessive generalization of students' networks.
Smart Images

Figure CN120126094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unsupervised road obstacle anomaly detection, and in particular to an unsupervised road obstacle detection method based on reverse distillation. Background Art
[0002] In people's daily travel, some obstacles on the road, such as fallen goods, vehicle wrecks, stones, etc., will affect traffic safety and passing efficiency, and need to be discovered and cleared in time.
[0003] At present, road obstacle monitoring mainly relies on manual observation. The method of manual observation not only has low efficiency, but also is difficult to meet the real-time requirements. In recent years, with the wide application of deep learning theory in the transportation field, many means worthy of reference have been provided for intelligent transportation. Certain achievements have been made in the research on road obstacle detection and recognition based on deep learning, but there are still the following problems:
[0004] 1) Most of the current work uses a supervised method to train the network to identify obstacles. However, abnormal situations are rare, and it is difficult to obtain enough abnormal samples for training. Moreover, the types of road obstacles are diverse, the cost of collecting and annotating all possible obstacles is very high, and there may be a situation of insufficient data.
[0005] 2) In unsupervised anomaly detection algorithms, the algorithms are often only trained for specific data sets. When the scenario changes, the generalization of the algorithms is relatively poor.
[0006] 3) In the anomaly detection algorithm based on reverse distillation, due to the strong generalization of the neural network, the abnormal area will also be reconstructed, resulting in detection errors.
[0007] Based on the above discussion, it has high practical application value to invent an unsupervised road obstacle detection method with good generalization and accuracy. Summary of the Invention
[0008] The purpose of the present invention is to overcome the defects of the prior art, and propose an unsupervised road obstacle detection method based on reverse distillation, which can effectively solve the problems of lack of abnormal samples for training, feature domain deviation and over-generalization of the student network, and at the same time meet the requirements of high generalization while achieving high accuracy.
[0009] To achieve the above object, the technical solution provided by the present invention is: an unsupervised road obstacle detection method based on reverse distillation, which realizes accurate unsupervised road obstacle detection based on an improved reverse distillation network. The improved reverse distillation network improves the multi-scale feature fusion module on the basis of the original reverse distillation network and adds an adaptive mask generator module; among them, the multi-scale feature fusion module is called the MFF module, and the improvement of the MFF module is: an improved CBAM attention module and an MLP layer are added after the MFF module for processing; the improvement of the CBAM attention module is: a dual-path design of dilated convolution and standard convolution is introduced in the spatial attention, and the fusion effect of different scale information is optimized through a balance factor; the MLP layer is used to fuse multi-level features, so as to better learn the data representation that can distinguish whether there are obstacles; the adaptive mask generator module contains a memory bank for storing the features of obstacle-free pictures, generates an adaptive mask by comparing the memory bank with the test pictures, and multiplies the test pictures and the adaptive mask element by element and then sends them into the student network of the reverse distillation network for reconstruction;
[0010] The specific implementation of the unsupervised road obstacle detection method includes the following steps:
[0011] 1) Collect road data through a road monitoring camera, including pictures of obstacle-free and obstacle-present roads collected on sidewalks, roadways, and highways. Among them, the pictures of obstacle-free roads on the road are normal samples, and the pictures of obstacle-present roads are abnormal samples. Then, the collected road data is divided into a training set and a test set. The training set only contains normal samples, and the test set contains normal samples and abnormal samples;
[0012] 2) Send the training set into the improved reverse distillation network for training. During the training process, first, the teacher network in the reverse distillation network extracts multi-scale features. The improved multi-scale feature fusion module can select the most discriminative features for distinguishing whether there are obstacles on the road. Then, the features are input into the student network for reconstruction, using the cosine similarity loss to align the features with the corresponding scale features extracted by the teacher network. At the same time, the features of normal samples are saved in the memory bank;
[0013] 3) Input the test set into the trained reverse distillation network. Multiply the features extracted by the teacher network and the improved multi-scale feature fusion module element by element with the generated adaptive mask and then send them into the student network for reconstruction. Calculate the cosine similarity between the multi-scale features of the student network and the teacher network as the anomaly score to evaluate the test pictures in the test set. Pictures with an anomaly score greater than the threshold are abnormal samples, and pictures with a score less than the threshold are normal samples.
[0014] Further, the step 1) includes the following steps:
[0015] 1.1) Data collection: Obtain the surveillance video data collected by road surveillance cameras, extract frames from the video and save them as image data;
[0016] 1.2) Data preprocessing: After collecting the data, eliminate the images with poor quality and adjust the images to a fixed size;
[0017] 1.3) Image annotation: Use the road images without obstacles as normal samples, the images with obstacles as abnormal samples, and the obstacle areas as abnormal areas;
[0018] 1.4) Dataset division: Divide the training set and the test set according to a certain proportion. The training set only contains normal samples, and the test set contains both normal samples and abnormal samples.
[0019] Furthermore, in step 2), the images in the training set are input into a pre-trained teacher network to extract multi-scale features, and then the extracted multi-scale features are fed into an improved multi-scale feature fusion module to generate feature maps. The improved multi-scale feature fusion module includes an MFF module, an improved CBAM attention module, and an MLP layer. The MFF module is beneficial for extracting image features of different scales to detect obstacles of different scales. The MFF module concatenates the multi-scale feature representations before feature embedding. To achieve alignment during feature concatenation, the MFF module downsamples the shallow features through one or more 3×3 convolutional layers with a stride of 2, then applies batch normalization and the ReLU activation function, and then uses a 1×1 convolutional layer with a stride of 1, combined with batch normalization and ReLU activation to generate compact features; the improved CBAM attention module consists of two sequential sub-modules: channel attention and spatial attention. Through average pooling and two consecutive 1×1 convolutions, the channel attention explores the correlations between feature channels and adds these correlations back to the original feature map to highlight important feature channels. The spatial attention further explores spatial correlations. It first aggregates spatial information by concatenating the average pooling and max pooling results of the feature map, and then discovers the spatial importance of the aggregated features through a path of standard convolution, dilated 3×3 convolution, normalization, and ReLU, and optimizes the fusion effect of different scale information through a factor set to 0.1. Finally, the feature maps with spatial information are obtained by adding the features of the two paths; by adding the feature maps with spatial information to the original feature map, important regions are highlighted; finally, an MLP layer is used to fuse multi-level features, so as to better learn the data representation that can distinguish normal and abnormal; the improved multi-scale feature fusion module can effectively solve the problems of feature domain shift and feature redundancy, and select the features that are most discriminative for distinguishing obstacle areas;
[0020] The adaptive mask generator module is a module that generates a mask by comparing the memory bank with the test image. The memory bank is a set obtained by storing the features of normal samples extracted during training. During the test process, the features of the test image are compared with the features in the memory bank to generate a mask of the same dimension. Among them, the mask value corresponding to the feature close to the feature in the memory bank is 1, and the mask value corresponding to the feature far from the feature in the memory bank is 0. Before the features are sent to the student network for reconstruction, the features are multiplied element by element with the generated adaptive mask, that is, the obstacle area is masked and then sent to the student network for reconstruction, so as to avoid misjudgment caused by the strong generalization ability of the student network reconstructing the obstacle area. The features reconstructed by the student network are aligned with the features of the corresponding scale extracted by the teacher network, and the area with large difference is the obstacle area.
[0021] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0022] 1. The data of the present invention are real data collected by road monitoring cameras, and there are multiple different road scenarios, which provide data support for road obstacle detection.
[0023] 2. The present invention processes by adding an improved CBAM attention module and an MLP layer after the MFF module of the reverse distillation network. The improvement of the CBAM attention module is to introduce a dual-path design of dilated convolution and standard convolution in the spatial attention, and optimize the fusion effect of different scale information through a balance factor, and use the MLP layer to fuse multi-level features, so as to better learn the data representation that can distinguish normal and abnormal; the optimized multi-scale feature fusion module effectively reduces the feature dimension and solves the domain bias problem of the features extracted by the pre-trained teacher network, making the retained features more suitable for the current data set.
[0024] 3. In the adaptive mask generator module, an adaptive mask is generated by comparing the memory bank of normal samples with the features of the test image, and the obstacle area is masked and then sent to the student network for reconstruction, which can effectively avoid the over-generalization problem of abnormal areas caused by the strong generalization ability of the student network.
[0025] 4. Taking AUROC as the evaluation index, the classification and localization accuracies of the present invention reach 92.2% and 98.1% respectively. Compared with the original reverse distillation network, the present invention has higher detection accuracy and better generalization ability.
[0026] 5. The present invention realizes the accurate detection and positioning of road obstacles, can monitor road obstacles in real time, so that the staff can discover and handle road anomalies in time, and avoid traffic accidents, which is of great significance to traffic safety and passing efficiency. In addition, the present invention can also be applied to other anomaly detection fields, such as industrial product surface defect detection, medical image anomaly detection, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is the overall architecture diagram of the method of the present invention; in the figure, Teacher is the teacher network, Student is the student network, MFF is the multi-scale feature fusion module, DCAM is the improved CBAM attention module, MLP is a fully connected layer, the memory bank is a set of normal sample features saved during training, LOSS is the loss function of training, that is, the cosine similarity between the teacher network and the student network.
[0028] Figure 2 is the structural schematic diagram of DCAM; in the figure, f is the input feature of DCAM, Avgpool is the average pooling, Maxpool is the maximum pooling, Conv1×1 is the 1×1 convolution, Conv3×3 is the 3×3 convolution, Relu is the activation function, Norm is the normalization, Dilated is the dilated convolution, and f’ is the output feature of DCAM.
[0029] Figure 3 is the comparison diagram of the experimental results of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0031] As Figures 1 to 3 shown, this embodiment discloses an unsupervised road obstacle detection method based on reverse distillation. This method realizes accurate unsupervised road obstacle detection based on an improved reverse distillation network. The improved reverse distillation network improves the multi-scale feature fusion module and adds an adaptive mask generator module on the basis of the original reverse distillation network; among them, the multi-scale feature fusion module is called the MFF module, and the improvement of the MFF module is: an improved CBAM attention module and an MLP layer are added after the MFF module for processing, and the improvement of the CBAM attention module is as Figure 2As shown in the figure, a dual-path design of dilated convolution (Dilated) and standard convolution (Conv3×3) is introduced for spatial attention, and the fusion effect of different-scale information is optimized through a balance factor. An MLP layer is used to fuse multi-level features, so as to better learn the data representation that can distinguish the presence or absence of obstacles; the adaptive mask generator module contains a memory bank that stores the features of obstacle-free pictures, and an adaptive mask is generated by comparing the memory bank with the test pictures. After multiplying the test pictures and the adaptive mask element by element, they are sent to the student network of the reverse distillation network for reconstruction.
[0032] The specific implementation of this unsupervised road obstacle detection method includes the following steps:
[0033] 1) Collect road data through a road camera, including pictures of obstacle-free and obstacle-present roads collected on sidewalks, roadways, and highways. Among them, the pictures of obstacle-free roads are normal samples, and the pictures of obstacle-present roads are abnormal samples; then divide the data set into a training set and a test set. The training set only contains normal samples, and the test set contains normal samples and abnormal samples; specifically, it includes the following steps:
[0034] 1.1) Collect data: From March 1st to August 20th, 2023, video data in different scenarios were collected through road monitoring cameras, including scenarios such as sidewalks, roadways, tunnels, and highways. The obstacles on the road are some common items, such as fallen goods, vehicle wrecks, stones, etc.;
[0035] 1.2) Image preprocessing: Since the monitoring camera collects video data, we extract frames from the video and save them as image data, and some low-quality data are removed. In addition, we uniformly adjust the image size to 1280*720 and 960*1280;
[0036] 1.3) Image annotation: The road pictures without obstacles are used as normal samples, and the pictures with obstacles are used as abnormal samples. The obstacle area is the abnormal area;
[0037] 1.4) Divide the data set: After image annotation, the normal sample pictures are divided into a training set and a test set according to a certain proportion, and the abnormal samples are put into the test set.
[0038] 2) Feed the training set into the improved reverse distillation network for training. During the training process, the picture is input into the pre-trained teacher network to extract multi-scale features, and then the extracted features are sent into the improved multi-scale feature fusion module to generate feature maps. The improved multi-scale feature fusion module includes an MFF module, an improved CBAM attention module, and an MLP layer. Among them, the MFF module is beneficial to extracting picture features of different scales to detect obstacles of different scales. The MFF module concatenates the multi-scale feature representations before feature embedding. To achieve alignment during feature concatenation, the MFF module downsamples the shallow features through one or more 3×3 convolutional layers with a stride of 2, then applies batch normalization and the ReLU activation function, and then uses a 1×1 convolutional layer with a stride of 1, combined with batch normalization and ReLU activation to generate compact features; the improved CBAM attention module consists of two sequential sub-modules: channel attention and spatial attention. Through average pooling and two consecutive 1×1 convolutions, the channel attention explores the correlations between feature channels and adds these correlations back to the original feature map to highlight important feature channels. The spatial attention further explores spatial correlations. It first aggregates spatial information by concatenating the average pooling and max pooling results of the feature map, and then discovers the spatial importance of the aggregated features through the path of standard convolution, dilated 3×3 convolution, normalization, and ReLU, and optimizes the fusion effect of different scale information through a factor set to 0.1. Finally, the feature maps with spatial information are obtained by adding the features of the two paths; by adding the feature maps with spatial information to the original feature map, important regions are highlighted; finally, an MLP layer is used to fuse multi-level features, so as to better learn the data representation that can distinguish normal and abnormal; the improved multi-scale feature fusion module can effectively solve the problems of feature domain shift and feature redundancy, and select the features that are most discriminative for distinguishing obstacle regions.
[0039] 3) Input the test set into the trained reverse distillation network. Multiply the features extracted by the teacher network and the improved multi-scale feature fusion module element-wise with the generated adaptive mask, and then send them into the student network for reconstruction. Calculate the cosine similarity between the multi-scale features of the student network and the teacher network as the anomaly score to evaluate the test images in the test set. Images with an anomaly score greater than the threshold are anomaly samples, and images with a score less than the threshold are normal samples. During the test, compare the features of the test images with the features in the memory bank to generate a mask of the same dimension. Among them, the mask value is 1 for the features closer to those in the memory bank and 0 for those farther away. Before the features are sent into the student network for reconstruction, multiply the features element-wise with the generated adaptive mask, that is, shield the obstacle area and then send it into the student network for reconstruction to avoid misjudgment caused by the strong generalization ability of the student network to reconstruct the obstacles. Align the features reconstructed by the student network with the corresponding scale features extracted by the teacher network, and the area with a large difference is the obstacle area.
[0040] The experimental results of this experiment are described in detail below:
[0041] According to the final detection results of the network, the standard metric "area under the receiver operating characteristic curve" (AUROC) is used to measure the performance. For obstacle anomaly detection, Img-AUROC is used; for pixel-level obstacle anomaly localization, Pix-AUROC is used.
[0042] The comparison results with the original reverse distillation network in the form of ablation experiments are shown in Table 1 below.
[0043] Table 1
[0044]
[0045] The results in the above table show that the improved reverse distillation network has increased by 12.6% to 95.1% in Img-AUROC and by 0.9% to 96.1% in Pix-AUROC compared with the traditional reverse distillation network.
[0046] The comparison results of the improved reverse distillation network with other algorithms are shown in Table 2 below.
[0047] Table 2
[0048]
[0049] The results in the above table show that the improved reverse distillation network has achieved relatively large improvements in both image-level classification and pixel-level localization compared with other anomaly detection algorithms.
[0050] The comparison results between the traditional reverse distillation network and the improved reverse distillation network are as Figure 3As shown in the figure, (a) in the figure is the test image and the corresponding annotation, (b) in the figure is the detection effect of the original reverse distillation network, and (c) in the figure is the detection effect of the improved reverse distillation network. It can be seen from the figure that the improved reverse distillation network can achieve a more accurate positioning effect and avoid misdetecting pedestrians as anomalies.
[0051] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An unsupervised road obstacle detection method based on reverse distillation, characterized in that: The method realizes accurate unsupervised road obstacle detection based on an improved reverse distillation network. The improved reverse distillation network improves the multi-scale feature fusion module and adds an adaptive mask generator module on the basis of the original reverse distillation network. The multi-scale feature fusion module is called the MFF module. The improvement of the MFF module is that an improved CBAM attention module and an MLP layer are added after the MFF module for processing. The improvement of the CBAM attention module is that a dual-path design of dilated convolution and standard convolution is introduced in spatial attention, and the fusion effect of information of different scales is optimized by a balance factor. The MLP layer is used to fuse multi-level features, so as to better learn data representations that can distinguish whether there are obstacles or not. The adaptive mask generator module includes a memory bank for storing features of obstacle-free images, and generates an adaptive mask by comparing the memory bank with a test image. The test image is multiplied element by element with the adaptive mask and then sent to the student network of the reverse distillation network for reconstruction. The specific implementation of the unsupervised road obstacle detection method includes the following steps: 1) Collect road data through road monitoring cameras, including pictures of roads with and without obstacles collected on sidewalks, roadways, and highways. The pictures of roads without obstacles are normal samples, and the pictures of roads with obstacles are abnormal samples. Then divide the collected road data into training sets and test sets. The training set only contains normal samples, and the test set contains normal samples and abnormal samples. 2) The training set is sent to the improved reverse distillation network for training. During the training process, the teacher network in the reverse distillation network first extracts multi-scale features. The improved multi-scale feature fusion module can select the most discriminative features for distinguishing whether there are obstacles on the road. The features are then input into the student network for reconstruction. The cosine similarity loss is used to align the features with the features of the corresponding scale extracted by the teacher network. At the same time, the features of normal samples are saved in the memory bank. 3) The test set is input into the trained reverse distillation network. The features extracted by the teacher network and the improved multi-scale feature fusion module are multiplied element-by-element with the generated adaptive mask and then sent to the student network for reconstruction. The cosine similarity between the multi-scale features of the student network and the teacher network is calculated as the anomaly score to evaluate the test images in the test set. Images with anomaly scores greater than the threshold are abnormal samples, and images with anomaly scores less than the threshold are normal samples.
2. The unsupervised road obstacle detection method based on reverse distillation according to claim 1, characterized in that: The step 1) comprises the following steps: 1.1) Data collection: Obtain surveillance video data collected by road surveillance cameras, extract frames from the video and save them as image data; 1.2) Data preprocessing: After collecting data, remove the collected images with poor quality and adjust the images to a fixed size; 1.3) Image annotation: Road images without obstacles are regarded as normal samples, images with obstacles are regarded as abnormal samples, and obstacle areas are regarded as abnormal areas; 1.4) Divide the data set: Divide the training set and test set into a certain proportion. The training set contains only normal samples, while the test set contains both normal and abnormal samples.
3. The unsupervised road obstacle detection method based on reverse distillation according to claim 1, characterized in that: In step 2), the images in the training set are input into the pre-trained teacher network to extract multi-scale features, and then the extracted multi-scale features are sent to the improved multi-scale feature fusion module to generate feature maps. The improved multi-scale feature fusion module includes an MFF module, an improved CBAM attention module and an MLP layer, wherein the MFF module is conducive to extracting image features of different scales to detect obstacles of different scales. The MFF module splices the multi-scale feature representations before feature embedding. In order to achieve alignment during feature splicing, the MFF module downsamples the shallow features through one or more 3×3 convolutional layers with a stride of 2, and then applies batch normalization and ReLU activation functions, and then uses a 1×1 convolutional layer with a stride of 1, combined with batch normalization and ReLU activation to generate compact features; The improved CBAM attention module consists of two sequential submodules: channel attention and spatial attention. Through average pooling and two consecutive 1×1 convolutions, channel attention explores the correlation between feature channels and adds these correlations back to the original feature map to highlight important feature channels. Spatial attention further explores spatial correlation. It first aggregates spatial information by splicing the average pooling and maximum pooling results of the feature map, and then discovers the spatial importance of aggregated features through standard convolution and dilated 3×3 convolution, normalization, and ReLU paths. The fusion effect of information at different scales is optimized by a factor set to 0.
1. Finally, a feature map with spatial information is obtained by adding the features of the two paths. By adding the feature map with spatial information to the original feature map, important areas are highlighted. Finally, an MLP layer is used to fuse multi-level features to better learn data representations that can distinguish between normal and abnormal data. The improved multi-scale feature fusion module can effectively solve the problem of feature domain offset and feature redundancy, and select the most discriminative features for distinguishing obstacle areas. The adaptive mask generator module is a module that generates a mask by comparing a memory bank with a test image, wherein the memory bank is a set obtained by storing the features of normal samples extracted during the training process; during the test process, the features of the test image are compared with the features in the memory bank to generate a mask of the same dimension, wherein the mask value that is close to the features in the memory bank is 1, and the mask value that is far away is 0; before the features are sent to the student network for reconstruction, the features are multiplied element by element with the generated adaptive mask, that is, the obstacle area is shielded and sent to the student network for reconstruction, so as to avoid misjudgment caused by the strong generalization ability of the student network reconstructing the obstacle area; the features reconstructed by the student network are aligned with the features of the corresponding scale extracted by the teacher network, and the areas with large differences are obstacle areas.
Citation Information
Cited By
Bogie defect detection method and device
CN121527035A