Road abnormal obstacle detection method and system based on multi-modal semantic consistency
Patent Information
- Application Number
- CN202610961457.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
立体视觉方法虽能提供几何信息,但在复杂光照、纹理缺失等场景下稳定性不足;语义分割技术在像素级识别上取得显著进展,但在小目标、未知类别障碍物的检测上仍存在局限;不确定性估计为异常检测提供了可靠性衡量工具,但其在实际系统中的集成与实时性仍有待优化;基于图像重建的方法通过生成与比较策略增强了对异常区域的敏感性,但重建质量与计算效率之间的平衡尚需进一步探索
本发明以多模态一致性准则为理论核心,通过协同分析视觉语义、外观纹理与空间几何的一致性偏差来准确定位异常区域。系统首先利用改进的DeepLabV3+网络提取RGB语义特征,该网络通过集成轻量化MobileNetV2主干、ECA注意力机制及自适应特征融合策略,在提升分割性能的同时提取反映模型不确定性的SML分数与Softmax熵。为了引入外观模态约束,系统利用条件生成对抗网络重构伪正常场景图像,并通过预训练的VGG网络计算原始图像与重构图像间的感知差异,从而在纹理层面捕捉目标的异质性。与此同时,深度语义分支则采用双分支ResNet-18结合注意力特征互补模块提取几何空间下的DepthLogits,据此计算跨模态语义一致性差异,用以量化视觉与几何预测之间的概率散度。最终,将Logits矩阵、
差异、不确定性特征及感知差异图等多维异质信息进行通道级联,输入轻量化差异融合网络进行深度关联学习,通过识别一致性偏差所表征的联合异常特征,输出像素级的异常概率预测,实现了对开放环境下未知道路障碍物的高精度检测。
Smart Images

Figure CN122799400A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a method and system for detecting abnormal road obstacles based on multimodal semantic consistency. Background Technology
[0002] Against the backdrop of a surge in motor vehicle ownership and rapid socio-economic development, while transportation convenience has improved, safety hazards such as frequent traffic accidents have also become increasingly prominent, posing a serious threat to public travel safety. Research shows that traffic accidents are influenced by a combination of driver condition, road conditions, and environmental factors. In the early stages of automotive industry development, the industry primarily relied on active safety technologies such as Anti-lock Braking System (ABS), Electronic Stability Program (ESP), and Electronic Control Brake Assist System (EBA), as well as passive protective measures such as seat belts and airbags. With breakthroughs in hardware computing power, automotive safety technology has accelerated its evolution towards intelligence, ushering in a golden age for Advanced Driving Assistance Systems (ADAS). As the core of intelligent driving, ADAS uses sensors such as cameras and radar to perceive the surrounding environment in real time, enabling the identification, prediction, and intervention of traffic participants. However, most accidents stem from drivers' lack of perception of obstacles; achieving accurate identification and rapid response remains a major bottleneck that the ADAS field urgently needs to overcome.
[0003] Among road obstacles, unusual targets such as scattered goods, roadblocks, or small animals pose significant risks. These objects typically exhibit irregular geometric shapes and small spatial scales, and because their features often exceed the coverage of standard training sets, data-driven models suffer from perceptual blind spots in such scenarios, resulting in persistently high false negative rates. Therefore, exploring a method to improve the accuracy and efficiency of abnormal obstacle detection is crucial.
[0004] In the field of environmental perception, machine vision solutions, with their highly biomimetic characteristics, rich information capacity, and low cost, demonstrate superior development potential compared to radar sensing. Research shows that approximately 90% of driving decisions originate from visual feedback. Currently, multi-view geometry and deep learning are the mainstream approaches to visual perception. Although deep learning, with its powerful feature extraction capabilities of Convolutional Neural Networks (CNNs), outperforms geometric algorithms in detection accuracy and semantic richness, its generalization ability remains insufficient when facing unlabeled, unknown targets.
[0005] Road obstacle detection is a key task in autonomous driving environmental perception, aiming to promptly identify unexpected obstacles such as fallen objects and debris to ensure driving safety. With the rapid development of automotive intelligence and autonomous driving technologies, road obstacle detection has become a key research focus for scholars both domestically and internationally. Currently, road obstacle detection methods are mainly divided into two categories: one is based on stereo vision, using binocular cameras to acquire depth information to detect protruding objects on the road surface. Its advantage lies in its lower cost, but its performance tends to degrade in low-light or sparsely textured road conditions. The other is based on deep learning, using neural networks to identify abnormal objects in images. This method has higher detection accuracy and stronger adaptability to complex scenes, but still carries the risk of missing small targets or unseen objects. Overall, road obstacle detection is evolving from traditional visual methods towards deep learning and multimodal fusion, but challenges remain in areas such as long-distance small target detection and adaptability to harsh environments. Specifically: 1) Road anomaly obstacle detection based on stereo vision: Stereo vision obstacle detection typically employs a binocular camera architecture, mimicking the stereoscopic vision mechanism of human eyes. This system recovers the three-dimensional geometric information of objects through the principle of parallax, reconstructing their outlines and spatial positions, and has wide applications in the field of machine vision.
[0006] Existing methods for detecting road anomalies based on stereo vision mainly fall into two categories: one is based on depth maps or disparity maps for obstacle detection; the other directly utilizes geometric information from images to achieve obstacle recognition. Existing solutions propose detection schemes based on randomly occupied grids, transforming free space determination into a dynamic programming problem. The probability of obstacle occupancy at the grid level is calculated using depth information, providing a quantitative basis for subsequent detection. With technological advancements, scene representation optimization has become crucial for improving detection efficiency. The Mercedes-Benz Image Understanding Group proposed the Stixel representation method, which merges pixels with the same depth and semantics into rectangular boxes based on the disparity map, introducing prior knowledge of the street scene to simplify analysis and significantly reduce computational complexity. Existing solutions further optimize the detection logic by proposing the Planar Hypothesis Testing (PHT) algorithm. This algorithm directly performs generalized likelihood ratio testing and maximum likelihood estimation on the local plane in the left and right original images. After distinguishing between free space and obstacle regions, it uses point clouds to achieve a visual representation of the obstacle region. For the need to detect unknown obstacles on streets, FPHT (Quantitatively Optimized PHT) and the Cluster-Stixels method are proposed based on PHT. The former improves the computational efficiency of planar hypothesis testing, while the latter generates columnar pixels through "rectangular region clustering-parallax depth splitting". Ultimately, both achieve clear representation of free space and obstacles through columnar pixels or point clouds.
[0007] 2) Road anomaly obstacle detection based on semantic segmentation: In recent years, deep learning has made groundbreaking progress in the field of object detection. Convolutional Neural Networks (CNNs), trained on large-scale data, can effectively improve feature extraction capabilities and detection performance, thereby achieving accurate object classification. As a powerful feature learning method, deep learning has been widely applied to road anomaly obstacle detection tasks. Especially in the iterative development of semantic segmentation technology, the research community has proposed a series of targeted optimization schemes to address core challenges such as the sparsity and morphological diversity of anomalies.
[0008] Semantic segmentation technology, through dense prediction and pixel-by-pixel category labeling of image pixels, has demonstrated significant advantages in road anomaly obstacle detection tasks. As a pioneering work in semantic segmentation, Fully Convolutional Networks (FCNs) achieved end-to-end pixel-level segmentation for the first time. This model effectively solved the problem of limited input image size by replacing the fully connected layers of traditional CNNs with deconvolutional layers and introducing skip connections. The subsequently proposed SegNet, as a typical improvement model of FCN, uses the first 13 convolutional layers of VGG16 as the encoder and designs a symmetrical decoder structure. This model generates pixel-level probability maps through softmax classification of the decoder output, but still has shortcomings in the detection accuracy of small-sized anomalies. Meanwhile, the existing ENet solution focuses on real-time requirements. This network improves its running speed by significantly reducing the number of parameters while maintaining a certain level of segmentation accuracy, enabling its deployment in embedded devices and providing possibilities for real-time applications in in-vehicle systems. The existing DeepLabV3+ model further optimizes multi-scale segmentation performance. This model integrates an encoder-decoder structure with an Atrous Spatial Pyramid Pooling (ASPP) module and utilizes dilated convolutions to expand the receptive field, significantly improving its segmentation capability for targets of different sizes, making it one of the benchmark models in this field. Addressing the core contradiction between "multi-scale adaptability" and "lightweight deployment," an improved DeepLabV3+ model is proposed. This model, by employing a lightweight backbone network and optimizing multi-scale feature fusion strategies, demonstrates an excellent balance between accuracy and deployment ease on the London Found dataset, providing a new technical path for road anomaly detection. Furthermore, existing technologies propose the YOLOv8-DGAF model based on the YOLOv8n framework. This research, by improving the feature extraction network, introducing an attention mechanism, optimizing multi-scale feature fusion, and designing a novel loss function, and training with a dedicated road anomaly dataset, ultimately surpasses several mainstream models in detection performance, providing a complementary solution for the collaborative optimization of object detection and semantic segmentation in this field.
[0009] 3) Road anomaly obstacle detection based on uncertainty estimation: In recent years, significant progress has been made globally in the research of road anomaly obstacle detection based on uncertainty estimation. Scholars both domestically and internationally have integrated advanced technologies such as deep learning and Bayesian theory to continuously improve the robustness and accuracy of detection systems.
[0010] In road scene semantic segmentation, uncertainty estimation has become a key technology for improving the reliability of anomaly obstacle detection. This field has laid the theoretical foundation for uncertainty quantification in deep neural networks by interpreting Dropout as an approximate Bayesian inference of a deep Gaussian process. Inspired by this, existing solutions have proposed the Bayesian SegNet framework. This framework utilizes Monte Carlo (MC) Dropout to generate pixel-level posterior distributions, achieving for the first time accurate quantification of model uncertainty in semantic segmentation, providing an important theoretical tool for road anomaly detection. Subsequently, existing solutions have proposed a simple baseline method based on Softmax distribution probability. This method effectively detects classification errors and outliers by analyzing the distribution characteristics of output probabilities, laying the foundation for the application of uncertainty estimation in anomaly detection. Building on this, existing solutions have proposed a semantic segmentation method based on model uncertainty. This method quantifies the uncertainty of the segmentation results, identifying high-uncertainty regions as unknown objects, thereby effectively reducing misclassification of unknown categories and improving the reliability of the segmentation results. Furthermore, the existing technical solution proposes an innovative method that integrates a meta-classifier and an aggregated dispersion measure to detect prediction errors and outlier regions in semantic segmentation. By comprehensively analyzing the geometric properties and statistical characteristics of the predicted segments, it significantly enhances the system's ability and reliability in identifying anomalous regions.
[0011] 4) Detection of road anomalies and obstacles based on reconstructed images: In recent years, road anomaly and obstacle detection based on reconstructed images has become a research hotspot in the field of autonomous driving safety. Its core lies in using generative models to reconstruct road scene images and identifying abnormal obstacles by analyzing the differences between the reconstructed and original images. In early research, autoencoders were widely used as a fundamental tool for image reconstruction. However, limited by model capacity and training data, their reconstruction quality is limited, making it difficult to accurately capture subtle anomalies in complex road scenes. To improve reconstruction accuracy and anomaly detection performance, existing technologies have proposed an innovative reconstruction module that focuses on the identification and reconstruction of road surface regions. By calculating the reconstruction error and fusing it with semantic segmentation results, it effectively integrates prior knowledge of the road surface, thereby generating a more accurate anomaly score map. This method not only improves the sensitivity of anomaly detection but also deepens the model's understanding of changes in road structure. With the development of Generative Adversarial Networks (GANs), GAN-based image synthesis methods have shown great potential in anomaly detection. These methods use semantic feature maps generated by semantic segmentation models as conditional inputs to synthesize new images that are semantically consistent with the original image but may differ in details. By comparing the differences between the original image and the synthesized image at the pixel level or feature level, abnormal obstacles can be located more accurately.
[0012] In summary, current research on road anomaly detection has gradually evolved from traditional stereo vision methods to a multi-technology fusion approach centered on deep learning. While stereo vision methods can provide geometric information, they suffer from insufficient stability in scenarios with complex lighting and missing textures. Semantic segmentation technology has made significant progress in pixel-level recognition, but it still has limitations in detecting small targets and obstacles of unknown categories. Uncertainty estimation provides a reliability measurement tool for anomaly detection, but its integration and real-time performance in practical systems still need optimization. Image reconstruction-based methods enhance sensitivity to anomaly regions through generation and comparison strategies, but the balance between reconstruction quality and computational efficiency requires further exploration.
[0013] Overall, most existing methods rely on a single modality or limited priors, and have limited ability to detect abnormal obstacles of unknown categories, small scale, and severe occlusion. Furthermore, there is still room for improvement in terms of multimodal information fusion, real-time performance, and system generalization ability.
[0014] Therefore, there is an urgent need for a road anomaly obstacle detection method based on multimodal semantic consistency. Starting from multimodal semantic consistency, this method integrates various information such as visual, geometric, and generative models to construct a lightweight, efficient, and highly generalizable anomaly obstacle detection framework to address the complex and ever-changing detection challenges in open road scenarios. Summary of the Invention
[0015] To address the problems existing in the prior art, the present invention aims to propose a road anomaly obstacle detection method and system based on multimodal semantic consistency. This method compares the pixel-level semantic distributions of the RGB and Depth modalities, quantifies the estimated divergence between the two, and achieves unsupervised semantic characterization of unknown objects. Furthermore, it integrates RGB Logits... Standardized Max Logits (SML), entropy, and perceptual differences are uniformly mapped to an 8-channel feature space. With extremely low parameter overhead, modal complementarity is utilized to significantly enhance the sensitivity to abnormal regions. At the same time, a lightweight MobileNetV2 backbone, an efficient channel attention (ECA) mechanism, adaptive feature fusion, and a multi-branch classification head are introduced to improve segmentation accuracy while maintaining the real-time requirements of the vehicle end. Furthermore, a full-process multimodal consistency perception framework is established. This framework does not rely on prior information of abnormal categories and achieves detection based on the properties of consistency deviation, making it more suitable for complex perception tasks in open-world environments.
[0016] To achieve the above objectives, the present invention provides the following solution: Road anomaly obstacle detection methods based on multimodal semantic consistency include: Acquire RGB images and their corresponding depth maps, input the RGB images and the depth maps into a pre-trained obstacle detection model, and obtain the detection results; The semantic segmentation module of the obstacle detection model uses an improved DeepLabV3+ network to extract RGB features, obtain semantic graphs and original RGB scores, and extract the multidimensional statistical uncertainty features of the network. The improved DeepLabV3+ network is obtained by introducing a lightweight MobileNetV2 backbone, ECA attention mechanism, adaptive feature fusion strategy and multi-branch classification head into the DeepLabV3+ network. Based on the synthesis module, the pseudo-normal scene image is reconstructed using the semantic map and perceptually compared with the RGB image to generate a perceptual difference map; The depth module is used to extract the hierarchical fusion features of the RGB image and the depth map, obtain the original depth score of each pixel in the spatial depth dimension, and calculate the original RGB score and the original depth score to obtain the cross-modal semantic consistency difference. The difference module is used to perform multidimensional feature concatenation and fusion of the RGB image, the original RGB score, the multidimensional statistical uncertainty feature, the original depth score, the perceptual difference map, and the cross-modal semantic consistency difference, and outputs the final pixel-level anomaly probability map.
[0017] Optionally, the semantic segmentation module includes: An image prediction unit is used to extract RGB features of the RGB image using an improved DeepLabV3+ network, obtain prediction results, and generate the semantic map and the original RGB score. A semantic segmentation unit is used to extract multidimensional statistical uncertainty features of the network from the RGB image; the multidimensional statistical uncertainty features include: Standardized SML score: ; in, The standardized SML score, The predicted category for this point. and These are the preset statistical parameters for the corresponding categories; Softmax entropy: ; in, For pixels Softmax entropy at the location, The current pixel belongs to the category The Softmax probability.
[0018] Optionally, the improved DeepLabV3+ network includes: The traditional deep residual network in the DeepLabV3+ network is replaced with the MobileNetV2 lightweight backbone network, and a linear bottleneck layer is introduced in the low-dimensional projection layer. The single DeepLabV3+ network The convolution is replaced by a multi-scale convolution block constructed using convolution kernels of different scales distributed in parallel. An ECA module is introduced to generate channel weights through cross-channel interaction without dimensionality reduction, which are used to calibrate the output features of the multi-scale convolution block. The residual connection based on the identity mapping fuses the calibrated output features with the original input features at the element level to obtain the low-level output features. At the same time, the ECA module is embedded at the end of each dilated convolution branch of the ASPP module to obtain the high-level output features. The adaptive feature fusion module replaces the direct concatenation of high-level semantic features and shallow spatial features in the DeepLabV3+ network architecture. The high-level semantic features are upsampled by bilinear interpolation until they reach the resolution of the corresponding shallow spatial features for initial concatenation along the channel axis. Initial spatial alignment of multi-scale features is achieved by using target convolution to obtain initial multi-scale features. At the same time, a dynamic weight generator is used to extract the global context vector through global average pooling, and a multilayer perceptron with a two-layer fully connected structure is used to capture the correlation between channels. The generated dynamic weight vector is used to dynamically recalibrate the initial spatial alignment of the initial multi-scale features to obtain weighted features.
[0019] Optionally, the improved DeepLabV3+ network further includes: The spatial information of the weighted features is aggregated by global average pooling and global max pooling through a convolutional block attention module. Channel attention masks are generated through a shared multilayer perceptron network. The channel attention masks are then matrix-multiplied with the weighted features to obtain the mapped and enhanced weighted features. Average pooling and max pooling operations are then performed along the channel axes. The two are then concatenated and input into a convolutional layer of the target size to calculate and generate a spatial saliency mask. The saliency mask is then matrix-multiplied with the mapped and enhanced weighted features to obtain refined features after spatial location calibration. Simultaneously, the refined features are smoothed based on two target convolutional layers and element-wise added to the preliminary multi-scale features to obtain the final weighted features. The output layer of the DeepLabV3+ network is replaced with a parallel multi-scale classification branch. Three sets of convolutional layers with different kernel sizes are used to process the fused enhanced feature map in parallel to extract residual semantic information under different perceptual dimensions, thereby obtaining the primary prediction mask of each branch. At the same time, the preset parameter vector is normalized by the Softmax function to generate adaptive weight coefficients, which are then weighted and fused with the primary prediction mask to obtain the prediction result.
[0020] Optionally, the synthesis module includes: The synthesis unit is used to reconstruct a pseudo-normal scene image from the semantic map using a conditional generative adversarial network, extract target features from the pseudo-normal scene image using a pre-trained VGG network as a feature extractor, measure the difference between the pseudo-normal scene image and the RGB image, and normalize it to the target interval to obtain the perceptual difference map.
[0021] Optionally, the depth module includes: The depth unit is used to input the RGB image and the depth map into the dual-branch ResNet-18 backbone network to obtain the hierarchical fusion features of appearance modality and geometric modality; The attention feature complementarity unit is used to adaptively adjust the modal differences of the hierarchical fusion features through a channel attention mechanism to obtain a fusion feature map; The ASPP unit is used to obtain the multi-scale feature mapping relationship of the fused feature map, and gradually restore it to the original image resolution through multi-level upsampling branches. It obtains the fused semantic map and the original depth score of each pixel in each depth interval, and calculates the cross-modal semantic consistency difference between the original RGB score and the original depth score.
[0022] Optionally, obtaining the fused feature map includes: ; in, The fused feature map This represents element-wise multiplication. Represents global average pooling and Cascaded convolution operations It is the Sigmoid activation function. These are the feature maps for the RGB branch and the depth branch, respectively.
[0023] Optionally, calculating the cross-modal semantic consistency difference between the original RGB score and the original depth score includes: ; in, For cross-modal semantic consistency, Predict probabilities for visual modalities. For deep mode prediction probability, For pixel space coordinates, for Norm.
[0024] Optionally, the difference module includes: The difference unit is used to map the RGB image, the original RGB score, the multidimensional statistical uncertainty feature, the perceptual difference map, the original depth score, and the cross-modal semantic consistency difference to a unified 8-channel feature space using multiple parallel target convolutional layers, and to initially suppress noise through linear projection. The mapped features are concatenated along the channel dimension to form a comprehensive difference feature. The comprehensive difference feature is then subjected to deep aggregation of cross-modal differences based on a lightweight difference CNN network. At the same time, the semantic information in the potential abnormal regions extracted using the cross-modal semantic consistency differences is enhanced to obtain deep aggregated features. The deep aggregated features are mapped to a single channel and then subjected to pixel-level anomaly probability prediction via a Sigmoid activation function to generate the pixel-level anomaly probability map, which serves as the detection result.
[0025] To achieve the above objectives, the present invention also provides a road anomaly obstacle detection system based on multimodal semantic consistency, comprising: The data acquisition subsystem is used to acquire RGB images and their corresponding depth maps. An obstacle detection subsystem is used to input the RGB image and the depth map into a pre-trained obstacle detection model to obtain detection results; the semantic segmentation module of the obstacle detection model uses an improved DeepLabV3+ network to extract RGB features, obtain a semantic map and original RGB scores, and extract the multidimensional statistical uncertainty features of the network; the improved DeepLabV3+ network is obtained by introducing a lightweight MobileNetV2 backbone, ECA attention mechanism, adaptive feature fusion strategy and multi-branch classification head into the DeepLabV3+ network; Based on the synthesis module, the pseudo-normal scene image is reconstructed using the semantic map and perceptually compared with the RGB image to generate a perceptual difference map; The depth module is used to extract the hierarchical fusion features of the RGB image and the depth map, obtain the original depth score of each pixel in the spatial depth dimension, and calculate the cross-modal semantic consistency difference between the original RGB score and the original depth score. The difference module is used to perform multidimensional feature concatenation and fusion of the RGB image, the original RGB score, the multidimensional statistical uncertainty feature, the original depth score, the perceptual difference map, and the cross-modal semantic consistency difference, and outputs the final pixel-level anomaly probability map.
[0026] The beneficial effects of this invention are as follows: This invention uses the multimodal consistency criterion as its theoretical core, accurately locating abnormal regions by collaboratively analyzing the consistency deviations of visual semantics, appearance texture, and spatial geometry. The system first utilizes an improved DeepLabV3+ network to extract RGB semantic features. This network integrates a lightweight MobileNetV2 backbone, ECA attention mechanism, and adaptive feature fusion strategy, improving segmentation performance while extracting SML scores and Softmax entropy that reflect model uncertainty. To introduce appearance modality constraints, the system uses a conditional generative adversarial network to reconstruct pseudo-normal scene images and calculates the perceptual differences between the original and reconstructed images using a pre-trained VGG network, thereby capturing the heterogeneity of the target at the texture level. Simultaneously, the deep semantic branch employs a dual-branch ResNet-18 combined with an attention feature complementarity module to extract DepthLogits in the geometric space, thereby calculating cross-modal semantic consistency differences. This is used to quantify the probability divergence between visual and geometric predictions. Finally, the Logits matrix, Multidimensional heterogeneous information such as differences, uncertainty features, and perceptual difference maps are channel-cascaded and input into a lightweight difference fusion network for deep association learning. By identifying joint abnormal features represented by consistency deviations, pixel-level anomaly probability predictions are output, achieving high-precision detection of unknown road obstacles in open environments. Attached Figure Description
[0027] To more clearly illustrate the embodiments or technical solutions of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a diagram illustrating the overall architecture of the road anomaly obstacle detection method based on multimodal semantic consistency according to an embodiment of the present invention. Figure 2 This is a diagram of the improved Deeplabv3+ network structure according to an embodiment of the present invention; Figure 3 This is a diagram of the deep module network structure according to an embodiment of the present invention; Figure 4 The following are the anomaly prediction results of the semantic module in this embodiment of the invention: (a) is the original RGB image, (b) is the anomaly prediction semantic map, (c) is the SML prediction map, and (d) is the Softmax entropy map. Figure 5 This is a diagram illustrating the anomaly prediction effect of the synthesis module in an embodiment of the present invention. Figure 6 This is a diagram illustrating the anomaly prediction effect of the depth module in an embodiment of the present invention. Figure 7 This is a network structure diagram of the difference module in an embodiment of the present invention; Figure 8 This is a diagram illustrating the anomaly prediction effect of the difference module in an embodiment of the present invention. Figure 9 This is a schematic diagram of a road anomaly obstacle detection system based on multimodal semantic consistency according to an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] like Figure 1 As shown, this embodiment discloses a road anomaly obstacle detection method based on multimodal semantic consistency, including: acquiring RGB images and their corresponding depth maps; inputting the RGB images and depth maps into a pre-trained obstacle detection model to obtain detection results; extracting RGB features through the semantic segmentation module of the obstacle detection model using an improved DeepLabV3+ network to obtain semantic maps and original RGB scores, and extracting the multidimensional statistical uncertainty features of the network; the improved DeepLabV3+ network utilizes a lightweight MobileNetV2 backbone, ECA attention mechanism, and adaptive... The system utilizes a feature fusion strategy and a multi-branch classification head to obtain the image. Based on the synthesis module, it reconstructs a pseudo-normal scene image using a semantic map and performs a perceptual comparison with the RGB image to generate a perceptual difference map. The depth module extracts hierarchical fusion features from the RGB image and the depth map, obtains the original depth score of each pixel in the spatial depth dimension, and calculates the original RGB score and the original depth score to obtain the cross-modal semantic consistency difference. The difference module performs multi-dimensional feature concatenation and fusion of the RGB image, the original RGB score, multi-dimensional statistical uncertainty features, the original depth score, the perceptual difference map, and the cross-modal semantic consistency difference to output the final pixel-level anomaly probability map.
[0032] Specifically, this invention is based on the multimodal consistency criterion, and locates abnormal regions by collaboratively analyzing the consistency deviations of visual semantics, appearance texture and spatial geometry.
[0033] This method consists of four core branches: RGB semantic segmentation, scene image reconstruction, deep semantic perception, and multi-source differential fusion. First, multimodal alignment preprocessing is performed on the Cityscapes dataset to provide visual and geometric input for subsequent models. Then, in the RGB semantic branch, an improved DeepLabV3+ network integrating ECA attention and adaptive feature fusion is used to extract features. This optimizes the segmentation accuracy for complex scenes while outputting key classification logits, semantic maps, and metrics based on uncertainty estimation (including standardized maximum logits and softmax entropy).
[0034] To further introduce appearance modality constraints, a conditional generative adversarial network (cGAN) is used to reconstruct pseudo-normal scene images based on semantic maps, and these images are perceptually compared with the original images to capture the heterogeneity of abnormal targets at the texture level. Secondly, the deep semantic branch integrates a dual-branch ResNet-18 backbone network with an AttentionFeature Complementary (AFC) module, obtaining Depth Logits and a deep semantic map in geometric space through multi-level feature extraction. Based on the obtained features, the system calculates cross-modal semantic consistency differences to quantify the probability divergence between visual and geometric predictions. Finally, the Logits matrix, uncertainty features, perceptual difference maps, and semantic consistency differences are combined... A lightweight fusion network is constructed by cascading channels and inputting them into the network. This network mines the joint anomaly features characterized by multimodal consistency deviations and outputs pixel-level anomaly probabilities, achieving accurate detection of unknown obstacles in complex open environments.
[0035] Improved DeepLabV3+ Model: Addressing the challenges of varying road obstacle scales, blurred edges, and limited computational resources in complex traffic scenarios, this invention provides in-depth optimizations based on the standard DeepLabV3+ architecture. For example... Figure 2 As shown, this network implements a multi-stage refinement strategy, including lightweighting the backbone network, introducing multi-scale feature processing, an adaptive decoder fusion strategy, and a multi-branch classification head to achieve high-precision semantic segmentation. More specifically: Lightweight Backbone Network Design: To address the stringent real-time requirements of road obstacle detection tasks, this model abandons traditional deep residual networks (such as ResNet-101) with their large parameter count and computational complexity, instead adopting a lightweight backbone network, MobileNetV2, optimized for mobile devices. The core of MobileNetV2 lies in replacing standard convolutions with depthwise separable convolutions (DSCs). Standard convolutions perform convolution operations simultaneously in both spatial and channel dimensions, resulting in high computational costs. The formula is as follows: (1); DSC decomposes the computation into two parts: depthwise convolution and pointwise convolution. Its total computational cost is... The formula is as follows: (2); In the above two formulas, For feature map size, and Number of input and output channels, Let be the kernel size. The formula for the computational ratio between the two is as follows: (3); In normal use In the case of convolution kernel, .because Typically larger, DSC's computational cost is only one-third that of standard convolution. This order-of-magnitude optimization of computational complexity enables DeepLabV3+ to achieve high-frequency real-time inference on computationally limited automotive embedded platforms.
[0036] Beyond a significant reduction in computational cost, MobileNetV2 ensures robust feature extraction through deep architectural innovations. Its core employs an inverse residual structure, which, unlike the traditional "compression-convolution-dilation" pattern of residual structures, follows a "dilation-convolution-compression" strategy. By first performing convolutional extraction in a high-dimensional feature space, it maximizes the preservation of complex semantic information in road scenes and effectively alleviates the gradient vanishing problem in deep networks. Simultaneously, the model introduces a linear bottleneck layer in the low-dimensional projection layer. By abandoning nonlinear activation functions (such as ReLU) at narrow channels, it prevents feature collapse caused by nonlinear transformations, thus ensuring the integrity of key details such as obstacle edges during inter-layer transmission. Through these structural optimizations, the model maintains excellent feature representation accuracy while significantly reducing the number of parameters and computational cost (FLOPs), laying a solid foundation for efficient multi-scale feature fusion.
[0037] Enhanced Multi-Scale Feature Processing: Addressing the challenges of varying road obstacle scales and complex backgrounds in autonomous driving scenarios, this invention proposes a collaborative optimization strategy. By reconstructing low-level spatial pathways and high-level semantic pathways, it systematically improves the model's perception accuracy for multi-scale features. In the low-level feature pathway, this invention designs and implements a multi-scale processor integrating an efficient channel attention (ECA) module, aiming to overcome the limitations of the standard DeepLabV3+ module's single-scale processing capability. The problem of spatial geometric information loss due to convolutional projection is addressed. This processor core employs parallel-distributed convolutional kernels of different scales to construct multi-scale convolutional blocks, whose output features… The formula is as follows: (4); In the formula, Indicates the kernel size as Convolution operation, This is the input shallow feature map. To further enhance key features, the model introduces an ECA module for channel calibration. Let... For global average pooling, the calculation formula is as follows: (5); The ECA module generates channel weights through cross-channel interaction without dimensionality reduction. The formula is as follows: (6); In the formula, It is the Sigmoid activation function. Indicates the kernel size as One-dimensional convolution. The final calibrated features are To further ensure shallow spatial information To ensure integrity during the transmission process, the model introduces residual connections based on identity mapping, which element-wise fuse attention-weighted features with the original input features to obtain the final low-level output features. The calculation formula is as follows: (7); Among them, the initial input features In residual connections, information is directly preserved through an identity mapping mechanism. Simultaneously, in high-level feature paths, this invention constructs an EnhancedASPP module to refine multi-scale contextual information. Unlike traditional parallel sampling methods, this invention embeds an independent ECA module after each dilated convolution branch of the ASPP. Let the... The dilated convolution operation for each branch is as follows: (its expansion rate) If ), then the output of that branch. The calculation formula is as follows: (8); in, These are high-level semantic features, obtained by deep extraction of RGB images through the MobileNetV2 backbone network and then processing by the improved Enhanced ASPP module.
[0038] This design enables the network to perform fine-grained channel-level refinement of semantic features acquired with different porosity, effectively filtering out key semantic components in the receptive field at each scale and suppressing background noise.
[0039] Adaptive decoder fusion strategy: In the standard DeepLabV3+ architecture, high-level semantic features and shallow spatial features are usually fused statically by direct concatenation. However, due to the significant feature mismatch between the two types of features in terms of abstraction level and receptive field scale, simple linear stacking often ignores the differences in the contribution of different features to the final segmentation result, which can easily lead to blurred target boundaries or semantic submersion of small targets.
[0040] To address this, the present invention designs an Adaptive Feature Fusion (AFF) module. This module achieves deep interaction and collaborative calibration of multi-level features through dynamic weight adjustment and spatial-channel dual attention constraints.
[0041] 1) Dynamic alignment and weight generation mechanism: To achieve effective integration of semantic features and spatial details, the module first processes high-level features. Perform bilinear interpolation upsampling to adjust its resolution to match that of low-level features. Maintain consistency. Subsequently, the high and low layer features are initially stitched together along the channel axis, and then... Convolution achieves preliminary spatial alignment of multi-scale features, and its calculation formula is as follows: (9); In the formula, These are the blended features after initial alignment.
[0042] To address the competition between different feature pathways, the AFF module introduces a dynamic weight generator. This generator extracts the global context vector through global average pooling and utilizes a multilayer perceptron with a two-layer fully connected structure to capture the correlation between channels, generating a dynamic weight vector. The calculation formula is as follows: (10); In the formula, and The weight matrix is a learnable matrix. The ReLU activation function is used. The generator outputs two complementary scalar weights. and This is used to balance the contributions of multi-level features. The mixed features are dynamically recalibrated through weighting operations, resulting in the final weighted features. The calculation formula is as follows: (11); This mechanism enables the model to adaptively adjust the weights of attention to spatial details or semantic context based on the content of the current scene.
[0043] 2) Refined feature processing based on the CBA module: After initial weighted fusion, to further suppress noise in the mixed features and highlight salient targets, this invention introduces a Convolutional Block Attention (CBA) module for refined feature processing. This module achieves dual calibration of the feature map in terms of dimensionality distribution and spatial location by concatenating channel attention and spatial attention mechanisms. In the channel attention stage, the model simultaneously utilizes global average pooling and global max pooling to aggregate spatial information, and generates a channel attention mask through a shared multilayer perceptron network. The formula is as follows: (12); This enables enhanced mapping of important semantic channels, namely... Following this, in the spatial attention phase, the model performs average pooling and max pooling operations along the channel axis, and utilizes... Large convolutional kernels extract saliency masks for spatial dimensions. The calculation formula is as follows: (13); This yields refined features after spatial location calibration. To ensure the stability of deep network training and optimize feature representation, the adaptive feature fusion module ultimately introduces a residual refinement structure, through two layers... The convolutional layer smooths the features and adds them element-wise to the initially aligned features, ultimately generating a fused feature map. The calculation formula is as follows: (14); Through the aforementioned multi-level attention constraints and residual refinement, the final generated fused feature map possesses both deep semantic category information and sharp obstacle edge details. This lays a solid foundation for feature representation in verifying the performance of the proposed improved model in subsequent complex road scenarios and in achieving high-precision pixel-level semantic segmentation.
[0044] Multi-branch classification head: In the network output layer, to further refine the final semantic segmentation mask, this invention designs and implements a multi-branch adaptive classification head to replace the traditional single classifier. The standard DeepLabV3+ architecture typically uses only a single classifier. While convolutional pixel classification has low computational cost, it often struggles to capture the complex geometric features of obstacle edges due to its fixed and limited receptive field. This is especially true when dealing with extremely small road targets, which can easily result in segmentation holes or blurred edges.
[0045] To this end, this module enhances the robustness of the prediction results by constructing parallel multi-scale classification branches. Specifically, the model utilizes three sets of convolutional layers with different kernel sizes to enhance the fused feature maps. Parallel processing is performed to extract residual semantic information from different perceptual dimensions, resulting in primary prediction masks for each branch. The calculation formula is as follows: (15); In the formula, ( (Total number of categories) represents the primary prediction mask for the output of each branch.
[0046] To achieve dynamic optimization of features at different scales, this module introduces a set of learnable parameter vectors. During model training, the vector is normalized using the Softmax function, thereby generating a set of mutually constrained adaptive weight coefficients that sum to 1. The calculation formula is as follows: (16); Final semantic segmentation output The predicted values from each scale branch are weighted and aggregated. To achieve effective reorganization of multi-scale information, the final weighted fusion formula for the prediction results is as follows: (17); This multi-branch adaptive fusion mechanism empowers the classification head to autonomously adjust the scale of attention based on the scene content. For example, when dealing with small obstacles at a distance, the model can learn to increase the weight of larger kernel branches to incorporate more contextual information; while when dealing with clear edges in the foreground, it tends to use smaller kernel branches to preserve details.
[0047] The final prediction result output by the above multi-branch adaptive fusion mechanism (i.e., the output of Equation 17), mathematically represents the network's original classification confidence score before probability normalization, which is defined as RGBLogits in this invention. Subsequently, the semantic segmentation module performs dual processing on these RGB Logits to generate the key inputs required for subsequent difference fusion. First, to obtain discrete category labels in the spatial dimension, the module applies the Argmax operator in the channel dimension to calculate the maximum index of the RGB Logits, thereby generating a pixel-level semantic map (semantic map) used to define environmental priors and guide the reconstruction of the synthesis module (cGAN). Second, the module retains the undiscretized RGB Logits themselves and further extracts multidimensional statistical uncertainty features such as Standardized Maximum Logits (SML) and Softmax entropy. Finally, the generated RGB Logits and semantic map are respectively fed to the difference module and the synthesis module as the basic anchor points for cross-modal consistency comparison.
[0048] Semantic Segmentation Module: As the core of the anomaly detection architecture of this invention, the semantic segmentation module is responsible for mapping the input RGB image to a high-level semantic space. This module not only generates pixel-level category labels but also extracts multidimensional statistical uncertainty features that reflect the model, including RGB Logits, Standardized MaxLogits (SML), and Softmax entropy. These features collectively constitute the criteria for subsequent differential fusion networks to capture anomalous targets.
[0049] Logits characteristics and statistical distribution modeling: Logits refer to the raw output vector of the last layer (before the Softmax layer) of a neural network, and their magnitude directly reflects the model's confidence in a specific class. Since the numerical distribution of Logits differs significantly between different classes, this invention introduces Standardized Max Logits (SML) to enhance comparability between classes.
[0050] First, this invention establishes a Logits distribution model for each known category by statistically analyzing the training set. For each category... The mean of its Logits With variance The calculation formula is as follows: (18); (19); In the formula, The total number of training samples, Spatial location coordinates, For an indicator function, if and only if the predicted category is... equal The value is 1 at time. This represents the corresponding maximum Logit value.
[0051] In the reasoning phase, regarding position The original maximum Logit value at the location Its standardized SML score The calculation formula is as follows: (20); In the formula, The predicted category for this point. and These are the preset statistical parameters for the corresponding category. Typically, Negative values are taken as anomaly scores. The lower the score, the lower the probability that the pixel belongs to the known category, that is, the higher the probability of anomaly.
[0052] Softmax entropy and prediction uncertainty: To quantify the degree of disorder in the probability distribution, this invention introduces Softmax entropy as a key indicator to measure the uncertainty of model predictions. Softmax entropy essentially reflects the degree of hesitation the model exhibits during classification. For models with... Each category of problem, pixels Softmax entropy at the location The calculation formula is as follows: (twenty one); In the formula, Does this pixel belong to a category? The entropy value represents the softmax probability. A higher entropy value indicates greater uncertainty in the model's classification decision at that pixel. When facing unknown road obstacles, because their features do not belong to any predefined training categories, the probability distribution output by the model is usually relatively flat, leading to a significant increase in entropy. This metric can complement SML features, enhancing the system's ability to perceive abnormal targets from a probabilistic perspective.
[0053] To more intuitively demonstrate the extraction process of the above-mentioned multidimensional statistical uncertainty features, Figure 4 This is a schematic diagram illustrating the processing effect of the semantic segmentation module in an embodiment of the present invention. For example... Figure 4As shown in (a)-(d), after inputting the original RGB image, the module accurately outputs the corresponding anomaly prediction semantic map. In particular, for the unknown road obstacle area in the image, the uncertainty features extracted by this invention show a significant spatial indicative role: the SML prediction map exhibits obvious low-score clustering characteristics in this anomaly area (i.e., significantly deviating from the Logits distribution of known categories), while the Softmax entropy map shows a strong high-entropy response in the corresponding area. This visualization result intuitively verifies that this module successfully achieves high-sensitivity capture of the uncertainty of undefined targets by forming a physical complementarity between the SML prediction map and the Softmax entropy map.
[0054] Synthesis Module: The synthesis module aims to achieve cross-domain reconstruction from the semantic domain to the image domain. Its core logic lies in using Conditional Generative Adversarial Nets (cGANs) to generate pseudo-normal reconstructed images based on the semantic segmentation map. If there are unknown obstacles in the input image, the semantic map will misclassify them as background (such as road surface), and the generated reconstructed image will present "clean" background features. This allows for the capture of abnormal targets by comparing the differences between the original image and the reconstructed image.
[0055] Unlike traditional GANs, cGANs introduce semantic graphs as auxiliary constraint information, giving the generation process a clear direction. Their goal is to make the distribution of generated images as consistent as possible with the distribution of real images given semantic conditions. Although generative adversarial techniques can produce highly realistic urban scenes, semantic graphs themselves only contain category logic and lack fine-grained color and appearance priors, resulting in pixel-level deviations in the accuracy of the generated images. This deviation limits the effectiveness of direct pixel-by-pixel comparison in image space.
[0056] To overcome the aforementioned limitations, this invention, inspired by perceptual loss theory, characterizes anomalies by calculating the perceptual difference between the original and reconstructed images in the feature space. Unlike traditional metrics that focus on pixel-level color distribution, perceptual difference emphasizes the comparison of image content and spatial structure, enabling more accurate capture of semantic inconsistencies caused by classification errors or unknown obstacles. Specifically, a VGG network pre-trained on the ImageNet dataset is used as the feature extractor. If no anomalies are detected in the scene or the classification is correct, the feature representations of the synthesized and original images should be consistent; conversely, if anomalies are present, the feature representations of the two will show significant deviations. For the original input image... and its corresponding reconstructed image Its perceived loss The calculation formula is as follows: (twenty two); In the formula, Indicates the VGG network's... The output of the layer feature map, represent Norm. Regarding the selection of feature layers, since deeper feature layers lose more detailed texture information about anomalous objects, this invention selects the first four layers of the VGG network for difference measurement. To ensure the uniformity of feature scale, the system normalizes the discrete difference metrics calculated from each layer to a norm. The resulting perceptual difference map will serve as a key feature input to a subsequent lightweight difference fusion network, providing strong modal constraints for the detection of unknown road obstacles from the perspectives of texture heterogeneity and appearance features. The perceptual loss is used to filter out invalid noise and accurately capture semantic conflicts; it is essentially equivalent to the difference measure.
[0057] Figure 5 This diagram illustrates the abnormal target capture effect after introducing the synthesis module in this embodiment of the invention. Internally, this module uses a conditional generative adversarial network (cGAN) to reconstruct pseudo-normal scenes and calculate perceptual differences, serving as the core method for capturing heterogeneity in appearance and texture. Figure 5 As shown, thanks to the strong activation response prior of the synthesis module to abnormal regions, the system is finally able to clearly and accurately locate and mark the position of unknown obstacles, effectively overcoming the problem of missed detection in complex scenarios in traditional single-modality systems.
[0058] Depth Module: The depth module takes the original RGB image and the depth map as input, such as... Figure 3 As shown, a hierarchical fusion feature extraction method is used to extract appearance modality and geometric modality features respectively using a dual-branch ResNet-18 backbone network. To fully exploit the complementary information between different modalities, the system introduces an Attention Feature Complementary (AFC) module for the output features of each layer of ResNet-18 to achieve dynamic deep fusion of dual-branch features. This module adaptively adjusts modal differences through a channel attention mechanism. First, global average pooling is used to generate channel descriptors, and then channel descriptors with consistent channel numbers are used... Convolutional layers enable cross-channel interaction, and finally, the sigmoid function is used to map the weight matrix to... Interval. Let the feature maps of the RGB branch and the depth branch be respectively... The final fused feature map The calculation formula is as follows: (twenty three); In the formula, This indicates element-wise (by channel) multiplication. Represents global average pooling and Cascaded convolution operations The activation function is Sigmoid. At the end of the encoding stage, the depth module further integrates the ASPP module to obtain multi-scale feature maps of the fused feature map, and gradually restores it to the original image resolution through multi-level upsampling branches. Finally, this module outputs the fused semantic map. Simultaneously, the depth module utilizes 3D geometric information to calculate the raw score (Depth Logits) of each pixel in each predefined semantic category (such as road surface, vehicle, tree, etc.), thereby quantifying the semantic consistency between visual and physical predictions in the geometric space dimension, providing core discriminative input for obstacle detection in subsequent complex road scenes. This fused semantic map is a purely discrete category label map obtained after applying the Argmax operator to the Depth Logits in the channel dimension. It represents the scene segmentation result made by the model purely based on 3D geometry and depth information, and is only used for auxiliary supervision or visualization.
[0059] Figure 6 This is a schematic diagram illustrating the feature extraction effect of the depth module in an embodiment of the present invention. Figure 6 As shown, the deep network, combining the input geometric modality (Depth graph) information, outputs significant DepthLogits fluctuations at locations where obstacles have 3D spatial protrusions; the generated fused semantic graph further strips away obstacles that are disguised in the 2D RGB appearance in the physical geometric dimension, providing a basis for subsequent calculation of cross-modal semantic consistency differences. It provides a precise spatial geometric basis.
[0060] Difference Module: As the final decision-making layer of the detection architecture of this invention, the difference fusion module's core task is to effectively integrate heterogeneous features from different modalities and evaluation dimensions, achieving accurate target localization by mining the significant deviations between normal and abnormal regions in multimodal representation. For example... Figure 7 As shown, this module uses the original RGB image, RGB Logits and uncertainty metrics (SML and Softmax entropy) output by the semantic segmentation module, Depth Logits generated by the depth module, and cross-modal semantic consistency differences. The system uses the perceptual difference map generated by the synthesis module as joint input. In the feature preprocessing stage, the system first utilizes multiple parallel... The convolutional layer maps the seven sets of multi-source heterogeneous features to an 8-channel feature space. This design not only achieves alignment of heterogeneous feature dimensions but also provides initial noise suppression through linear projection. Subsequently, the six mapped features are concatenated along the channel dimension to form a feature space with dimension . The comprehensive differences in characteristics.
[0061] Before performing cross-modal consistency comparison, the system first uses the Softmax normalized exponential function to map the RGB Logits and Depth Logits output by the preceding module from the original depth score space to the probability distribution space, thereby obtaining the visual modality prediction probabilities respectively. With depth mode prediction probability .
[0062] Cross-modal semantic consistency This is the core metric in this module for measuring conflicts between appearance and geometric logic. It is calculated by determining the probability of visual modality prediction. With depth mode prediction probability The degree of offset between them, quantifying the deviation in semantic estimation, is expressed by the following formula: (twenty four); In the formula, For pixel space coordinates, for Norm. When When the value is large, it indicates a significant semantic contradiction between the visual appearance and spatial geometric features at the same pixel, which usually suggests the presence of unknown or unusual obstacles in the region.
[0063] The cascaded combined features are then fed into a lightweight differential CNN network. In this network, the cross-modal semantic consistency differences calculated in the preceding steps are processed. Furthermore, the perceptual difference map serves as an explicit spatial indication prior, guiding the network to focus on local regions where modal features exhibit logical conflicts. Specifically, this lightweight network performs nonlinear mapping of cascaded features within their local receptive fields using multi-layer convolutional operators, and adaptively learns inter-channel correlation weights under training-driven conditions; when a conflict is detected... When surges and high predictive uncertainties (such as high Softmax entropy) highly overlap spatially, the network's convolutional kernels produce strong activation responses. Through this nonlinear feature aggregation and spatial filtering mechanism, the network effectively suppresses response noise in low-bias background regions, while essentially transforming multi-source quantization bias into a significant enhancement of the representation of potential anomalous targets. Finally, the output head utilizes... Convolution maps features to a single channel and uses the Sigmoid activation function to predict pixel-level anomaly probabilities. The resulting pixel-level anomaly probability map is calculated using the following formula: (25); The difference fusion module has a simple structure and a very low number of parameters, yet it can jointly utilize multi-dimensional clues such as semantic confidence, image reconstruction perception differences, geometric consistency, and model prediction uncertainty to characterize the cross-modal deviation of abnormal targets from multiple physical dimensions. This ensures that the improved model has extremely high anomaly detection sensitivity and robustness in complex and dynamic open scenarios.
[0064] Figure 8 This is a schematic diagram illustrating the anomaly probability prediction effect of the difference module in an embodiment of the present invention, i.e., the final detection result of the present invention. Figure 8 As shown, after inputting multidimensional features into a lightweight differential CNN network and performing deep aggregation, the system successfully suppressed response noise in low-bias background regions. In the final pixel-level anomaly probability map, the true unknown anomaly obstacle regions are accurately enhanced and located with high confidence, achieving robust detection under the multimodal semantic consistency criterion.
[0065] like Figure 9 As shown, this embodiment also provides a road anomaly obstacle detection system based on multimodal semantic consistency, including: The data acquisition subsystem is used to acquire RGB images and their corresponding depth maps. An obstacle detection subsystem is used to input the RGB image and the depth map into a pre-trained obstacle detection model to obtain detection results; the semantic segmentation module of the obstacle detection model uses an improved DeepLabV3+ network to extract RGB features, obtain a semantic map and original RGB scores, and extract the multidimensional statistical uncertainty features of the network; the improved DeepLabV3+ network is obtained by introducing a lightweight MobileNetV2 backbone, ECA attention mechanism, adaptive feature fusion strategy and multi-branch classification head into the DeepLabV3+ network; Based on the synthesis module, the pseudo-normal scene image is reconstructed using the semantic map and perceptually compared with the RGB image to generate a perceptual difference map; The depth module is used to extract the hierarchical fusion features of the RGB image and the depth map, obtain the original depth score of each pixel in the spatial depth dimension, and calculate the cross-modal semantic consistency difference between the original RGB score and the original depth score. The difference module is used to perform multi-dimensional feature concatenation and fusion of the original RGB image, the original RGB score, the multi-dimensional statistical uncertainty feature, the original depth score, the perceptual difference map, and the cross-modal semantic consistency difference, and outputs the final pixel-level anomaly probability map.
[0066] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A road anomaly obstacle detection method based on multimodal semantic consistency, characterized in that, include: Acquire RGB images and their corresponding depth maps, input the RGB images and the depth maps into a pre-trained obstacle detection model, and obtain the detection results; The semantic segmentation module of the obstacle detection model uses an improved DeepLabV3+ network to extract RGB features, obtain semantic graphs and original RGB scores, and extract the multidimensional statistical uncertainty features of the network. The improved DeepLabV3+ network utilizes a lightweight MobileNetV2 backbone, ECA attention mechanism, adaptive feature fusion strategy, and multi-branch classification head to obtain the features. Based on the synthesis module, the pseudo-normal scene image is reconstructed using the semantic map and perceptually compared with the RGB image to generate a perceptual difference map; The depth module is used to extract the hierarchical fusion features of the RGB image and the depth map, obtain the original depth score of each pixel in the spatial depth dimension, and calculate the original RGB score and the original depth score to obtain the cross-modal semantic consistency difference. The difference module is used to perform multidimensional feature concatenation and fusion of the RGB image, the original RGB score, the multidimensional statistical uncertainty feature, the original depth score, the perceptual difference map, and the cross-modal semantic consistency difference, and outputs the final pixel-level anomaly probability map.
2. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 1, characterized in that, The semantic segmentation module includes: An image prediction unit is used to extract RGB features of the RGB image using an improved DeepLabV3+ network, obtain prediction results, and generate the semantic map and the original RGB score. A semantic segmentation unit is used to extract multidimensional statistical uncertainty features of the network from the RGB image; the multidimensional statistical uncertainty features include: Standardized SML score: ; in, The standardized SML score, The predicted category for this point. and These are the preset statistical parameters for the corresponding categories; Softmax entropy: ; in, For pixels Softmax entropy at the location, The current pixel belongs to the category The Softmax probability.
3. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 2, characterized in that, The improved DeepLabV3+ network includes: The traditional deep residual network in the DeepLabV3+ network is replaced with the MobileNetV2 lightweight backbone network, and a linear bottleneck layer is introduced in the low-dimensional projection layer. The single DeepLabV3+ network The convolution is replaced by a multi-scale convolution block constructed using convolution kernels of different scales distributed in parallel. An ECA module is introduced to generate channel weights through cross-channel interaction without dimensionality reduction, which are used to calibrate the output features of the multi-scale convolution block. The residual connection based on the identity mapping fuses the calibrated output features with the original input features at the element level to obtain the low-level output features. At the same time, the ECA module is embedded at the end of each dilated convolution branch of the ASPP module to obtain the high-level output features. The adaptive feature fusion module replaces the direct concatenation of high-level semantic features and shallow spatial features in the DeepLabV3+ network architecture. The high-level semantic features are upsampled by bilinear interpolation until they reach the resolution of the corresponding shallow spatial features for initial concatenation along the channel axis. Initial spatial alignment of multi-scale features is achieved by using target convolution to obtain initial multi-scale features. At the same time, a dynamic weight generator is used to extract the global context vector through global average pooling, and a multilayer perceptron with a two-layer fully connected structure is used to capture the correlation between channels. The generated dynamic weight vector is used to dynamically recalibrate the initial spatial alignment of the initial multi-scale features to obtain weighted features.
4. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 3, characterized in that, The improved DeepLabV3+ network also includes: The spatial information of the weighted features is aggregated by global average pooling and global max pooling through a convolutional block attention module. Channel attention masks are generated through a shared multilayer perceptron network. The channel attention masks are then matrix-multiplied with the weighted features to obtain the mapped and enhanced weighted features. Average pooling and max pooling operations are then performed along the channel axes. The two are then concatenated and input into a convolutional layer of the target size to calculate and generate a spatial saliency mask. The saliency mask is then matrix-multiplied with the mapped and enhanced weighted features to obtain refined features after spatial location calibration. Simultaneously, the refined features are smoothed based on two target convolutional layers and element-wise added to the preliminary multi-scale features to obtain the final weighted features. The output layer of the DeepLabV3+ network is replaced with a parallel multi-scale classification branch. Three sets of convolutional layers with different kernel sizes are used to process the fused enhanced feature map in parallel to extract residual semantic information under different perceptual dimensions, thereby obtaining the primary prediction mask of each branch. At the same time, the preset parameter vector is normalized by the Softmax function to generate adaptive weight coefficients, which are then weighted and fused with the primary prediction mask to obtain the prediction result.
5. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 1, characterized in that, The synthesis module includes: The synthesis unit is used to reconstruct a pseudo-normal scene image from the semantic map using a conditional generative adversarial network, extract target features from the pseudo-normal scene image using a pre-trained VGG network as a feature extractor, measure the difference between the pseudo-normal scene image and the RGB image, and normalize it to the target interval to obtain the perceptual difference map.
6. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 1, characterized in that, The depth module includes: The depth unit is used to input the RGB image and the depth map into the dual-branch ResNet-18 backbone network to obtain the hierarchical fusion features of appearance modality and geometric modality; The attention feature complementarity unit is used to adaptively adjust the modal differences of the hierarchical fusion features through a channel attention mechanism to obtain a fusion feature map; The ASPP unit is used to obtain the multi-scale feature mapping relationship of the fused feature map, and gradually restore it to the original image resolution through multi-level upsampling branches. It obtains the fused semantic map and the original depth score of each pixel in each depth interval, and calculates the cross-modal semantic consistency difference between the original RGB score and the original depth score.
7. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 6, characterized in that, Obtaining the fused feature map includes: ; in, The fused feature map This represents element-wise multiplication. Represents global average pooling and Cascaded convolution operations It is the Sigmoid activation function. These are the feature maps for the RGB branch and the depth branch, respectively.
8. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 6, characterized in that, Calculating the cross-modal semantic consistency difference between the raw RGB score and the raw depth score includes: ; in, For cross-modal semantic consistency, Predict probabilities for visual modalities. For deep mode prediction probability, For pixel space coordinates, for Norm.
9. The road anomaly obstacle detection method based on multimodal semantic consistency according to claim 1, characterized in that, The difference module includes: The difference unit is used to map the RGB image, the original RGB score, the multidimensional statistical uncertainty feature, the perceptual difference map, the original depth score, and the cross-modal semantic consistency difference to a unified 8-channel feature space using multiple parallel target convolutional layers, and to initially suppress noise through linear projection. The mapped features are concatenated along the channel dimension to form a comprehensive difference feature. The comprehensive difference feature is then subjected to deep aggregation of cross-modal differences based on a lightweight difference CNN network. At the same time, the semantic information in the potential abnormal regions extracted using the cross-modal semantic consistency differences is enhanced to obtain deep aggregated features. The deep aggregated features are mapped to a single channel and then subjected to pixel-level anomaly probability prediction via a Sigmoid activation function to generate the pixel-level anomaly probability map, which serves as the detection result.
10. A road anomaly obstacle detection system based on multimodal semantic consistency implemented according to any one of claims 1-9, characterized in that, include: The data acquisition subsystem is used to acquire RGB images and their corresponding depth maps. An obstacle detection subsystem is used to input the RGB image and the depth map into a pre-trained obstacle detection model to obtain detection results; the semantic segmentation module of the obstacle detection model uses an improved DeepLabV3+ network to extract RGB features, obtain a semantic map and the original RGB score, and extract the multidimensional statistical uncertainty features of the network. The improved DeepLabV3+ network utilizes a lightweight MobileNetV2 backbone, ECA attention mechanism, adaptive feature fusion strategy, and multi-branch classification head to obtain the features. Based on the synthesis module, the pseudo-normal scene image is reconstructed using the semantic map and perceptually compared with the RGB image to generate a perceptual difference map; The depth module is used to extract the hierarchical fusion features of the RGB image and the depth map, obtain the original depth score of each pixel in the spatial depth dimension, and calculate the cross-modal semantic consistency difference between the original RGB score and the original depth score. The difference module is used to perform multidimensional feature concatenation and fusion of the RGB image, the original RGB score, the multidimensional statistical uncertainty feature, the original depth score, the perceptual difference map, and the cross-modal semantic consistency difference, and outputs the final pixel-level anomaly probability map.