Vehicle anomaly detection method, device and system based on cloud collaboration

Through the cloud-coordinated vehicle anomaly detection method, the Internet of Vehicles system and RGB-D semantic segmentation model are used to solve the problems of the limitations of a single source and the insufficient detection capabilities of open world in traffic scenarios, and the effective fusion and real-time response of multimodal data are achieved, which improves detection accuracy and security.

CN120510583APending Publication Date: 2025-08-19DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510517107.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Vehicle anomaly detection methods in existing traffic scenarios have limitations in a single source, insufficient detection capabilities of unknown anomaly detection in the open world, insufficient utilization of multimodal data, and real-time problems, resulting in incomplete detection, delay and safety hazards.

Method used

The vehicle anomaly detection method based on cloud-end collaboration is adopted to obtain multi-view data through the Internet of Vehicles system, combine the RGB-D semantic segmentation model and residual learning mechanism, and use cloud-edge collaborative knowledge distillation technology to achieve multimodal data fusion and real-time decision-making.

Benefits of technology

It improves the detection accuracy and generalization ability of unknown anomalies, reduces decision-making delays, enhances the robustness and security of traffic scenarios, and realizes rapid response of multi-source data and global environment perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510583A_ABST
    Figure CN120510583A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle anomaly detection method, device and system based on cloud collaboration. The method comprises the following steps: acquiring RGB image data and depth image data of a vehicle in real time by adopting a vehicle-mounted camera and a radar; after receiving the data, the cloud end matches a corresponding scene image and abnormal event text description data, and carries out abnormal event early warning on a side end vehicle; the cloud end trains an RGB-D semantic segmentation model based on residual learning based on a large amount of image data; performing dynamic knowledge distillation on the trained RGB-D semantic segmentation model to obtain a new RGB-D semantic segmentation model, and transmitting weight parameters of the new RGB-D semantic segmentation model to the vehicle end; and the vehicle end performs reasoning according to the data collected in real time, detects an abnormal condition and makes a decision, broadcasts an abnormal event to the surrounding vehicle end, and uploads the abnormal event to the cloud end to assist in scheduling decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a vehicle anomaly detection method, device and system based on cloud collaboration. Background Art

[0002] With the accelerated development of modern urbanization, intelligent transportation systems have become a key means of addressing urban traffic congestion, improving traffic management efficiency, and reducing traffic accidents. Intelligent transportation utilizes advanced information technology, communications technology, sensing technology, and artificial intelligence to enable real-time monitoring, analysis, and optimal decision-making of road traffic conditions. However, the current traffic environment is rife with complexity and unpredictability. Many abnormal events fall into the so-called "long tail" category, such as wild animals, vehicle wreckage, and garbage. These events are extremely rare in the data but have a significant potential impact on safe driving. Currently, no systematic solution has been developed for this "long tail problem of rare events." Existing anomaly detection methods using data collected by camera and radar sensors achieve scene perception and understanding by leveraging semantic segmentation models. These methods typically do not learn normal representations from training data alone (i.e., they cannot model anomalies based solely on normal data). Instead, they require a semantic segmentation stage and fine-tuning using out-of-distribution (OOD) data. This means that anomaly detection methods rely on additional data and prior knowledge, especially for rare anomaly data. Semantic segmentation is the task of assigning semantic labels (such as lane lines, cars, pedestrians, and road surfaces) to each pixel in an image, providing high-resolution scene understanding. This plays a key role in anomaly detection, the core goal of which is to identify and address abnormalities in driving scenarios that fall outside the normal range, ensuring vehicle safety and system robustness. These anomalies can arise from a variety of factors, including environmental anomalies: sudden road obstacles, inclement weather, or road surface problems; and traffic behavior anomalies: such as pedestrians suddenly crossing the road or non-motor vehicles driving against traffic. Therefore, accurately and rapidly detecting these anomalies is crucial for vehicle anomaly detection in current traffic scenarios. Existing traffic anomaly detection methods primarily rely on a single data source. However, the real-world traffic environment is complex and ever-changing. Relying solely on data collected by a single vehicle cannot comprehensively detect anomalies in traffic scenes, such as view occlusions. Furthermore, existing research primarily focuses on unimodal scenarios in image data. However, real-world applications are inherently multimodal, making it crucial to leverage multimodal information to improve the effectiveness of anomaly detection. Furthermore, while existing methods have made some progress in closed-loop test scenarios, they still face significant shortcomings in open-world driving scenarios. Open-world environments mean vehicles will encounter a vast array of unseen scenarios, objects, and behaviors, requiring the system to possess sufficient generalization capabilities to detect unknown anomalies. However, traditional anomaly detection methods are often based on known categories in the training data and have poor adaptability to unseen categories or rare events, severely limiting anomaly detection effectiveness in open environments. Furthermore, a crucial issue facing anomaly detection systems deployed onboard is the speed of anomaly detection. The question of how to effectively leverage the vast amounts of traffic scene data to achieve faster and more robust anomaly detection is also a pressing issue.

[0003] At present, the existing anomaly detection methods mainly include the following: 1) Independent detection work based on a separate anomaly detection system on the vehicle side. This method inevitably has the problem of being unable to quickly process large amounts of data to achieve real-time reasoning, and because the perspective of a single vehicle is too limited, it cannot fully and timely detect abnormal events encountered while driving; 2) Methods based on standard segmentation and detection models: This method deploys an anomaly detection model on the vehicle side system, including standard semantic segmentation, instance segmentation, object detection, video tracking and other methods. The success of this model is attributed to the dataset with a fixed set of semantic categories, without considering the possibility of identifying unknown objects from new categories. The poor adaptability to unseen categories or rare events severely limits the anomaly detection effect in open environments. Secondly, this method cannot utilize rich multimodal data. The model is only trained using a single RGB image. Moreover, this method only faces the problem of method 1, that is, collecting data from a single perspective, and cannot fully and timely detect abnormal events encountered while driving. 3) Method based on retrained out-of-distribution (OOD) semantic segmentation model: This method deploys anomaly detection models on the vehicle-side system. Unlike the standard model-based method, the idea of this method is to use abnormal data, namely OOD data, to retrain the early segmentation model. This retraining improves the pixel-by-pixel anomaly detection accuracy, but unlike the method of freezing the initial segmentation, this method has a high computational cost and will reduce the accuracy of the scene segmentation within the cloud distribution. Moreover, retraining may lead to a decline in the performance of in-distribution segmentation. For example, pixels of minority categories (such as "fence" and "sidewalk") may be misclassified as majority categories (such as "road"). And it also damages the accuracy of closed set segmentation. The problem with this method is that it has high computational cost and cannot meet the real-time and fast inference speed required for vehicle detection of abnormal events in traffic scenarios. This method also faces the same problem as method 2, that is, it cannot utilize rich multimodal data and the problem of a single perspective. In addition, this method faces the problem of being unable to cope with the anomaly detection problem in the open world. When the vehicle enters a new driving environment, the originally trained model cannot generalize well, resulting in the failure of anomaly detection in the current scene.

[0004] The main problems with existing methods include: 1) Limitations of a single signal source: The information collected by a single vehicle cannot fully detect driving anomalies. A common example is that the leading vehicle detects a traffic accident ahead, resulting in an anomaly left on the road (which the following vehicle's system has not learned and identified). In this case, the following vehicle cannot respond quickly due to occlusion, and because it has not learned this anomaly, re-inference cannot promptly identify this anomaly, which will pose a huge safety hazard to traffic and safe driving. 2) Insufficient detection capabilities for unknown anomalies in the open world: When operating in the open world, autonomous vehicles will encounter objects and behaviors that are not included in the training data (such as unknown obstacles and rare traffic incidents). Most existing methods rely on the ability to detect known categories and lack the generalization and adaptability to unknown anomalies, which can easily lead to missed detections or false detections. 3) Insufficient effective utilization of multimodal data: Data in the Internet of Vehicles scenario is often multi-source, heterogeneous, and multimodal data. Vehicles and scenes are usually equipped with multiple sensors (such as cameras, lidar, etc.), which provide the system with rich multimodal data. However, existing methods have significant shortcomings when integrating and utilizing this data, primarily manifested in the following: Strong reliance on a single modality: Because abnormal traffic events can manifest in various forms, such as vehicle congestion, pedestrian violations, and traffic accidents, relying solely on single-modality data makes it difficult to fully perceive the traffic scene. Lack of robustness: Single-modality data is prone to failure in complex scenarios. For example, cameras struggle to operate effectively in low-light or obstructed environments, resulting in a significant decrease in detection performance. Multimodal data fusion can address the shortcomings of single-modality data, but existing fusion techniques are still immature and fail to fully tap the potential of multimodal data. High computational cost: The introduction of multimodal data inevitably increases the computational cost of data processing and inference, necessitating a solution to the conflict between high accuracy and low computational efficiency. 4) Real-time performance: Anomaly detection must be completed quickly to support real-time decision-making by vehicle-side systems. However, many existing methods often sacrifice computational efficiency while improving detection accuracy, resulting in excessively high inference latency and difficulty meeting the real-time requirements of practical applications. This real-time issue becomes even more severe when multimodal data is introduced.

[0005] Therefore, in order to solve the above problems, the present invention proposes an innovative vehicle anomaly detection method, device and system based on cloud collaboration.

[0006] First, the present invention innovatively proposes to obtain rich multi-perspective driving scene data based on the information interaction system of the Internet of Vehicles, which is used to solve a series of problems caused by perspective occlusion and obstacle obstruction caused by a single signal source when the car is driving; at the same time, in order to solve the problem of insufficient open world detection capability faced by the anomaly detection model (semantic segmentation model), unlike the existing anomaly detection model that uses limited data collected in advance to train the model, resulting in insufficient generalization ability for unknown things, the present invention innovatively proposes to match the same scene data of a large amount of traffic scene data transmitted by the Internet of Vehicles through the cloud on the basis of collecting real-time multi-perspective scene data, and obtain targeted contextual scene data, and solve the problem by incorporating a large amount of current scene targeted data into the training; the present invention proposes an RGB-D semantic segmentation model based on the CNN model and the residual learning mechanism to detect unknown and potential anomalies in driving scenes, and improve the ability of traffic participants to recognize unknown anomalies. The difference between the present invention and the existing method is that a new RFL module is designed in the anomaly detection model architecture. Through the first stage of training, the model's recognition ability in known categories is enhanced. Then, through the second stage of training, the original model is frozen to maintain the original known category recognition ability. The RFL module is trained separately to guide the model to focus on identifying abnormal areas. In addition, the model training not only relies on color images (RGB data), but also processes two image data, color images and depth maps (Depth data) when the vehicle is driving based on the CNN model to generate feature maps of the two modal data and use the fused feature maps to perform anomaly detection, thereby realizing the effective utilization of multimodal data. In order to solve the real-time problem of anomaly detection and support real-time decision-making of the vehicle-side system, the present invention innovatively proposes to compress the model based on cloud-edge-end collaborative knowledge distillation technology, and train a small-scale semantic segmentation model by migrating the knowledge of the complex semantic segmentation model to achieve rapid reasoning during vehicle driving, meeting the high real-time and safety requirements of traffic scenarios, and adopting cloud-edge-end collaborative anomaly detection to timely allocate resources and handle accidents for abnormal events based on the information of the cloud data center. Summary of the Invention

[0007] In view of the problems existing in the prior art, the present invention proposes a vehicle anomaly detection method based on cloud collaboration, which specifically includes the following steps:

[0008] The vehicle's RGB image data and depth image data are collected in real time using onboard cameras and radars. The vehicle's information interaction system acquires real-time data about surrounding vehicles and transmits the acquired information to the cloud.

[0009] After receiving the data, the cloud matches the corresponding scene image and abnormal event text description data, and issues an abnormal event warning to the edge vehicle;

[0010] The cloud trains the RGB-D semantic segmentation model based on residual learning based on a large amount of image data;

[0011] Perform dynamic knowledge distillation on the trained RGB-D semantic segmentation model to obtain a new RGB-D semantic segmentation model, and transmit the weight parameters of the new RGB-D semantic segmentation model to the vehicle side;

[0012] The vehicle side makes inferences based on the data collected in real time, detects abnormal situations and makes decisions. At the same time, it broadcasts abnormal events to surrounding vehicles and uploads them to the cloud to assist in scheduling decisions.

[0013] Furthermore, while driving, the end-side vehicle collects RGB image data and depth image data through its on-board camera and radar, and uses the information interaction system based on the Internet of Vehicles to obtain real-time data of surrounding vehicles. It obtains multi-angle real-time data of the current driving scene collected by multiple vehicles and road infrastructure, and uploads all data to the cloud data center; the information interaction system based on the Internet of Vehicles can realize information exchange between vehicles, vehicles and road infrastructure, and vehicles and the cloud.

[0014] Furthermore, data retrieval is performed in the cloud data center based on the data uploaded by the edge to obtain targeted data. A large amount of image data that conforms to the current driving scenario and text descriptions of related abnormal events are prompted to the edge vehicle through the information interaction system, and warnings are issued for abnormal events that often occur in the current driving scenario.

[0015] Furthermore, the training process of the residual learning-based RGB-D semantic segmentation model is divided into two stages: using normal traffic scene image data to train the encoder, pyramid pooling module, and decoder in the CNN model; and using abnormal event data to train the residual learning module separately;

[0016] The first stage of training is performed on the RGB-D semantic segmentation model. The normal scene image X is used to train the dual-branch encoder FCN with a fully convolutional neural network as the backbone, the feature extraction module ASPP and the decoder SEG.

[0017] The dual-branch encoder FCN consists of four layers of convolutional neural networks, each layer of which is used to train X r 、X d Perform convolution operation to obtain feature map And the feature map S is obtained by fusing the two feature maps through the attention mechanism:

[0018]

[0019] After the dual encoder, the original image is converted into a feature map Z:

[0020] f φfcn:X→Z

[0021] The feature map is passed through the ASPP module to extract multi-scale context information:

[0022] f φaspp :Z→K

[0023] The decoder combines the high-level features of the encoder with the low-level detail features through skip connections, thereby restoring the details of the image and obtaining the segmentation result Y:

[0024] f φseg :Z→Y

[0025] The first stage training process can be summarized as follows:

[0026] Y=f φseg (f φaspp (f φfcn (X)))

[0027] The second stage of training is to train the RGB-D semantic segmentation model: the purpose of this stage is to train the residual learning module RPL separately on the basis of not damaging the original segmentation model for normal driving scenes, induce RPL to learn the abnormal parts in the segmentation image, and use abnormal scene images Freeze the encoder, feature extraction module, and decoder, and train the RFL module separately. The second stage training process can be summarized as follows:

[0028]

[0029] Furthermore, with the help of massive information training in cloud data centers, a teacher model (semantic segmentation model) suitable for real-time environments is obtained. This model has the characteristics of high complexity, high computing resource requirements, high precision and robustness.

[0030] With the output of the teacher model as a guide, using dynamic knowledge distillation technology, following the RGB-D semantic segmentation model training strategy based on residual learning, a lightweight, low-computation student model (semantic segmentation model) is trained, and the obtained optimized student model or adjusted weights and parameters are pushed to the edge vehicle system.

[0031] Furthermore, the vehicle collects different real-time data in different driving scenarios, and inputs the collected real-time image data into the updated semantic segmentation model for inference to obtain segmentation results. The vehicle detects abnormal events based on the results and makes decisions. After the vehicle detects an abnormal event, it broadcasts the abnormal event to surrounding vehicles through the information interaction system based on the Internet of Vehicles, and uploads it to the cloud and the dispatch center. After the dispatch center receives the abnormal event report, it generates relevant descriptive information about the abnormal event based on the cloud data center; the relevant descriptive information includes abnormal event description files, accident handling plans, current traffic scene data, and flow information.

[0032] A cloud-coordinated vehicle anomaly detection device includes: a memory and a processor for storing a computer program; the processor is used to implement a cloud-coordinated vehicle anomaly detection method when executing the computer program.

[0033] A cloud-coordinated vehicle anomaly detection system comprising:

[0034] The information interaction module based on the Internet of Vehicles is used to realize information interaction between vehicles, vehicles and road infrastructure, and vehicles and the cloud, and is used for uploading multi-angle real-time scene information, obtaining abnormal event information, and broadcasting abnormal events.

[0035] The intelligent perception end, including cameras and lidar, is used to collect real-time RGB image and depth image data.

[0036] The intelligent cloud platform is used to communicate with the intelligent platform in real time, retrieve and match the uploaded real-time data with a large amount of image data, abnormal event text description data, and traffic resource data in the cloud data center, and carry out model training with a semantic segmentation model based on residual learning.

[0037] Due to the adoption of the above technical solution, the present invention provides a vehicle anomaly detection method based on cloud collaboration, which has the following beneficial effects and advantages:

[0038] 1. Existing anomaly detection methods for traffic scenarios face several challenges. Regarding real-time performance, many methods sacrifice computational efficiency for improved detection accuracy. This is especially true when multimodal data from traffic scenarios is incorporated, significantly increasing model inference latency and making it difficult to meet the real-time decision-making requirements of autonomous driving. Furthermore, due to the limited information collection capabilities of a single vehicle, it is difficult to fully perceive and promptly respond to abnormal driving events (such as road obstacles under occlusion), posing a safety hazard. Furthermore, due to a lack of abnormal image data, existing methods are not adaptable to the open world. Their primary reliance on known categories in training data results in a lack of generalization and adaptability to unknown anomalies, leading to misclassification of normal pixels as anomalies and vice versa. This further limits detection accuracy and system robustness, potentially leading to missed or false detections. Although multimodal data (such as cameras and LiDAR) provides rich information, their fusion technology remains immature, resulting in a strong reliance on a single modality and prone to failure in complex scenarios. For example, depth images, compared to RGB images, contain more positional and contour information, making them a key indicator of objects in real-world driving scenarios. This invention addresses the potential risks of a single source by extracting features from RGB and depth graphics data collected during vehicle operation and combining this with data from surrounding traffic and infrastructure. Furthermore, it queries historical anomaly detection information stored in cloud data centers to obtain rich anomaly information, addressing the poor adaptability of open-world anomaly detection. By matching different data to different driving scenarios, this method queries and predicts more complete and targeted anomaly event information through a collaborative data fusion-driven approach across the cloud, edge, and end. This method improves the precision, accuracy, and generalization capabilities of anomaly detection.

[0039] 2. The present invention proposes an RGB-D semantic segmentation model based on a CNN model and a residual learning mechanism, which effectively fuses RGB images and depth image data in traffic scenes, improving the current method's insufficient utilization of multimodal data, strong dependence on a single modality, lack of robustness, and the problem that single modality data easily fails in complex scenes, resulting in a significant decrease in detection performance and lack of robustness. Secondly, the key to the residual learning mechanism is that because it does not interfere with the original model reasoning process, it can better identify unknown and potential anomalies in driving scene images without affecting the original closed set segmentation performance. This improves the model's adaptability to different driving scene contexts and improves traffic participants' ability to recognize unknown anomalies. It solves the problem of insufficient adaptability of existing methods in the open world.

[0040] 3. The present invention proposes a model compression method based on dynamic knowledge distillation of cloud-edge collaboration. A large amount of real-time data in traffic scenes and historical data stored in the data center are integrated on the cloud side to obtain feature maps containing more comprehensive information. The feature maps are then used to pre-train the complex large-scale teacher model deployed on the cloud side to obtain richer information. The teacher model is then used to guide the training and reasoning of the small model deployed on the edge side, and the trained small model or its weight parameters are transmitted to the vehicle side to update the model. By using knowledge distillation technology, the technology of migrating large model knowledge to train small models is used to significantly reduce model parameters and computing resource consumption while maintaining model performance, avoiding the problem of slow model reasoning speed caused by large amounts of data and large models, thereby greatly improving real-time performance. With the help of this dual detection mechanism, while ensuring the rapid reasoning of the edge side of driving scene anomaly detection, a large model is used to improve the accuracy of detection to ensure the safety of vehicle driving.

[0041] 4. This invention proposes an interactive system based on the Internet of Vehicles (IoV). Through information exchange between vehicles (V2V), between vehicles and road infrastructure (V2I), and between vehicles and the cloud (V2C), it can significantly improve the comprehensiveness and real-time performance of anomaly detection when vehicles are involved in traffic. Through multi-vehicle collaborative perception, following vehicles can quickly respond without re-inference, thereby reducing decision latency. Our method achieves multi-source data fusion. The IoV interactive system integrates multimodal information from multiple vehicles, road infrastructure, and the cloud, providing vehicles with global environmental perception capabilities. For example, vehicles can rely not only on their own sensors but also utilize real-time computational results from roadside cameras or the cloud to improve their understanding of complex traffic scenarios and anomaly detection capabilities. Furthermore, traffic safety is improved by sharing environmental information in real time. The IoV interactive system can effectively avoid the perception blind spots of a single source and quickly respond to potential traffic anomalies, thereby reducing accident risks and improving road safety. Our method also provides enhanced dynamic learning and updating capabilities. The IoV interactive system can support dynamic learning and knowledge sharing in the cloud. Even if a vehicle has not learned a certain type of abnormal situation, it can obtain prior knowledge from other vehicles that have learned such abnormalities or the cloud through the Internet of Vehicles, so as to quickly make the right response. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1It is the overall logical framework diagram for the implementation of the present invention;

[0044] Figure 2 This is a flowchart of the information interaction system based on the Internet of Vehicles;

[0045] Figure 3 This is a flowchart of the model compression process based on cloud-edge collaborative knowledge distillation;

[0046] Figure 4 Processing flow chart for the dispatch center;

[0047] Figure 5 This is the architecture diagram of the RGB-D semantic segmentation model based on residual learning. DETAILED DESCRIPTION

[0048] To make the technical solutions and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention:

[0049] The present invention obtains rich multi-perspective driving scene data based on the information interaction system of the Internet of Vehicles, and at the same time matches the same scene data for a large amount of traffic scene data under the application of the Internet of Vehicles in the cloud to obtain targeted contextual scene data; proposes an RGB-D semantic segmentation model based on the CNN model and the residual learning mechanism to detect unknown and potential anomalies in the driving scene, and improves the ability of traffic participants to recognize unknown anomalies. Based on the CNN model, the color image (RGB data) and depth map (Depth data) of the vehicle are processed when driving, and feature maps of the two modal data are generated and anomaly detection is performed with the help of the fused feature maps; model compression is performed based on cloud-edge collaborative knowledge distillation, and small models are trained by migrating large model knowledge. Taking into account the characteristics of traffic scenes with high real-time and safety requirements, cloud-edge collaborative anomaly detection is adopted, and timely resource allocation and accident handling are performed for abnormal events based on the information of the cloud data center.

[0050] like Figure 1 The figure shows the overall logical framework of the present invention, the core idea of which is:

[0051] ① Information interaction system based on the Internet of Vehicles: An interactive system based on the Internet of Vehicles. First, the vehicle collects RGB and depth data through its onboard sensors while driving, and uses the information interaction system based on the Internet of Vehicles to obtain real-time data of surrounding vehicles. At the same time, all data will be uploaded to the cloud and input into the real-time multimodal anomaly detection model based on residual learning for rapid inference. If an abnormal event is detected, it will be fed back to the original vehicle and reported to the surrounding vehicles, infrastructure and the cloud. Similarly, other vehicles and infrastructure use the same mode to communicate with each other. The specific processing flow is as follows: Figure 2 shown.

[0052] The innovation of this invention lies in that through the information interaction between vehicles (V2V), vehicles and road infrastructure (V2I), and vehicles and the cloud (V2C), the comprehensiveness and real-time performance of traffic anomaly detection can be significantly improved. Through multi-vehicle collaborative perception, the following vehicle can respond quickly without re-reasoning, thereby reducing decision delays. And the difference from existing methods is that our method realizes multi-source data fusion. The vehicle network interaction system integrates multimodal information from multiple vehicles, road infrastructure and the cloud to provide vehicles with global environmental perception capabilities. Providing better dynamic learning and updating capabilities, the vehicle network-based interaction system can support dynamic learning and knowledge sharing in the cloud. Even if a vehicle has not learned a certain type of abnormal situation, it can obtain prior knowledge from other vehicles or the cloud that have learned such anomalies through the vehicle network, so as to quickly make the right response.

[0053] ② Driving scenario information extraction and abnormal event warning: Each vehicle has different uses and faces different common driving scenarios. When driving in different driving scenarios, the vehicle's camera and radar collect real-time RGB image data and depth data to obtain data on the vehicle's current driving environment and upload it to the cloud. The cloud matches the uploaded data with richer scene data and key abnormal event data that may be encountered in the current driving scenario, and uses this more targeted data for subsequent processing.

[0054] Existing methods detect abnormal events by deploying anomaly detection models on the vehicle side. The models used are often trained with large amounts of data. Such models are costly, lack specificity, and suffer from slow inference speed and poor generalization in different contexts. In particular, they are unable to make judgments that are more in line with the current scenario for different commonly used driving scenarios. Therefore, the difference from the current method is that by matching real-time scene data with cloud data, more targeted relevant data can be obtained, which helps reduce training costs and facilitates subsequent processing.

[0055] ③ Cloud data fusion and RGB-D semantic segmentation model training based on residual learning: Cloud data contains relevant scene data of different scenes and abnormal event data, including common non-abnormal RGB images in traffic scenes, depth data, image data of abnormal events, description data, etc. The end-side data refers to the current driving scene data. The cloud data fusion method proposed in the present invention is: first, the common non-abnormal RGB images, depth data, image data of abnormal events, description data stored in the cloud data center are matched with the real-time scene image data collected in real time on the end side, and then targeted data under the same scene is obtained. The large amount of data obtained is then used for the representation learning of abnormal events in the cloud. The representation learning of abnormal events in the cloud is mainly divided into two steps.

[0056] The obtained data is used for training in the cloud. The training of the RGB-D semantic segmentation model based on residual learning is divided into two stages: using normal traffic scene image data to train the encoder, pyramid pooling module, and decoder in the CNN model; and using abnormal event data to train the residual learning module separately.

[0057] First, the first stage of model training uses the normal scene image X (RGB image data X r , and the corresponding depth map data X d ), train the dual-branch encoder FCN with a fully convolutional neural network as the backbone, the feature extraction module ASPP and the decoder SEG.

[0058] The dual-branch encoder consists of four layers of convolutional neural networks, each layer of which is r 、X d Perform convolution operation to obtain feature map And the feature map S is obtained by fusing the two feature maps through the attention mechanism AFC:

[0059]

[0060] After the dual encoder, the original image is converted into a feature map Z:

[0061] f φfcn :X→Z

[0062] The feature map is passed through the ASPP module to extract multi-scale context information:

[0063] f φaspp :Z→K

[0064] The decoder then combines the high-level features (ASPP output) of the encoder with the low-level detail features through skip connections, thereby restoring the details of the image and obtaining the segmentation result Y:

[0065] f φseg :Z→Y

[0066] The first stage training process can be summarized as follows:

[0067] Y=f φseg (f φaspp (f φfcn (X)))

[0068] The second stage is model training. The purpose of this stage is to train the residual learning module RPL separately without damaging the original segmentation model for normal driving scenes, and induce RPL to learn the abnormal parts in the segmentation image; using abnormal scene images (RGB image data Corresponding depth map data ) training, freezing the encoder, feature extraction module, and decoder, and training the RFL module separately; the second stage training process can be summarized as follows:

[0069]

[0070] The difference from existing methods lies in the design of a new convolutional neural network. Unlike the traditional CNN model, this model uses a dual-branch structure to process data from two modalities separately. For multimodal data from the Internet of Vehicles, the convolutional neural network (CNN) is first used to extract features from each modal data (such as RGB images and depth images) to obtain semantic information from low-level to high-level layers. Subsequently, by introducing a learning method based on the attention mechanism, the model can dynamically analyze and filter key information from different modal features. This mechanism can adapt to the asymmetry and redundancy of information in multimodal input and more effectively extract key features from the two modal data from a global perspective. This feature screening process helps the AFC module fully utilize the feature expressions of the RGB and depth branches during the fusion process, enhance the complementarity and synergy between the two, and further improve the model's adaptability to complex scenarios and its ability to capture abnormal events.

[0071] Secondly, it differs from existing methods in that it uses a residual learning mechanism to capture potential anomalies in the scenarios faced by the vehicle while driving. In multimodal feature extraction and fusion, capturing potential anomalies in the open world is a key link, and the introduction of a residual learning mechanism can effectively improve the model's sensitivity and robustness to abnormal features. The key to this mechanism is that it does not interfere with the original model reasoning process and can better identify unknown and potential anomalies in driving scenes without affecting the original closed set segmentation performance. Through residual learning, the model can not only focus on deep-level abnormal features, but also retain low-level original feature information, thereby reducing the risk of information loss. The residual module ensures the effective capture of abnormal features by adding its output to the output of the spatial pyramid pooling module.

[0072] ④ Model compression based on collaborative knowledge distillation between cloud and edge: The difference from the existing methods is that, based on the above model design, it is proposed to compress the large-scale model trained on the cloud side with a large amount of driving scene data by using knowledge distillation technology, reduce its time and space complexity, and then deploy it to the edge side. The basic idea of knowledge distillation technology is to train small models by migrating the knowledge of large models, which is used to significantly reduce model parameters and computing resource consumption while maintaining model performance. Its goal is to allow an efficient small model (student model) to learn knowledge of specific tasks from a complex large model (teacher model), thereby achieving compression effects. The specific process is as follows: Figure 3 shown.

[0073] ⑤ Real-time anomaly detection and dispatch center processing: Different from traditional anomaly detection that only relies on the vehicle-side system for anomaly detection, this method provides a cloud-edge collaborative anomaly detection method, and timely allocates resources and handles accidents for abnormal events based on the information of the cloud data center. First, the real-time data collected on the edge side (vehicle-side and traffic scene infrastructure) is input into its own anomaly detection model for timely detection, and uploaded to the cloud at the same time to perform anomaly detection using a model trained with a large amount of cloud data. When an abnormal event is detected, the abnormal event will first be fed back to the vehicle-side for early warning. The target vehicle will report the abnormal event to nearby vehicles through the information interaction system based on the Internet of Vehicles. Secondly, the abnormal event will be reported to the dispatch center. The dispatch center uses the information of the cloud data center to allocate traffic resources, control traffic flow, and handle subsequent accidents for abnormal events and traffic accidents. Specifically, Figure 4 and Figure 5 shown.

[0074] In summary, based on Figure 1 ①, ②, ③, ④, and ⑤ shown complete the overall implementation process description process of the solution of the present invention.

[0075] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A vehicle anomaly detection method based on cloud collaboration, characterized in that step: The vehicle's RGB image data and depth image data are collected in real time using onboard cameras and radars. The vehicle's information interaction system acquires real-time data about surrounding vehicles and transmits the acquired information to the cloud. After receiving the data, the cloud matches the corresponding scene image and abnormal event text description data, and issues an abnormal event warning to the edge vehicle; The cloud trains the RGB-D semantic segmentation model based on residual learning based on a large amount of image data; Perform dynamic knowledge distillation on the trained RGB-D semantic segmentation model to obtain a new RGB-D semantic segmentation model, and transmit the weight parameters of the new RGB-D semantic segmentation model to the vehicle side; The vehicle side makes inferences based on the data collected in real time, detects abnormal situations and makes decisions. At the same time, it broadcasts abnormal events to surrounding vehicles and uploads them to the cloud to assist in scheduling decisions.

2. The vehicle anomaly detection method based on cloud collaboration according to claim 1 is characterized by: While driving, the end-side vehicle collects RGB image data and depth image data through its on-board camera and radar, and uses the information interaction system based on the Internet of Vehicles to obtain real-time data of surrounding vehicles. It obtains multi-angle real-time data of the current driving scene collected by multiple vehicles and road infrastructure, and uploads all data to the cloud data center; the information interaction system based on the Internet of Vehicles can realize information exchange between vehicles, vehicles and road infrastructure, and vehicles and the cloud.

3. The vehicle anomaly detection method based on cloud collaboration according to claim 2 is characterized by: Based on the data uploaded by the edge, data retrieval is performed in the cloud data center to obtain targeted data. A large amount of image data that conforms to the current driving scenario and text descriptions of related abnormal events are prompted to the edge vehicle through the information interaction system, and warnings are issued for abnormal events that often occur in the current driving scenario.

4. The vehicle anomaly detection method based on cloud collaboration according to claim 3 is characterized by: The training process of the RGB-D semantic segmentation model based on residual learning is divided into two stages: using normal traffic scene image data to train the encoder, pyramid pooling module, and decoder in the CNN model; and using abnormal event data to train the residual learning module separately. The first stage of training is performed on the RGB-D semantic segmentation model. The normal scene image X is used to train the dual-branch encoder FCN with a fully convolutional neural network as the backbone, the feature extraction module ASPP and the decoder SEG. The dual-branch encoder FCN consists of four layers of convolutional neural networks, each layer of which is used to train X r 、X d Perform convolution operation to obtain feature map And the feature maps of the two are fused through the attention mechanism to obtain the feature map S: After the dual encoder, the original image is converted into a feature map Z: f φfcn :X→Z The feature map is passed through the ASPP module to extract multi-scale context information: f φaspp :Z→K The decoder combines the high-level features of the encoder with the low-level detail features through skip connections, thereby restoring the details of the image and obtaining the segmentation result Y: f φseg :Z→Y The first stage training process can be summarized as follows: Y=f φseg (f φaspp (f φfcn (X))) The second stage of training is to train the RGB-D semantic segmentation model: the purpose of this stage is to train the residual learning module RPL separately on the basis of not damaging the original segmentation model for normal driving scenes, induce RPL to learn the abnormal parts in the segmentation image, and use abnormal scene images Freeze the encoder, feature extraction module, and decoder, and train the RFL module separately. The second stage training process can be summarized as follows:

5. The vehicle anomaly detection method based on cloud collaboration according to claim 1 is characterized in that: The teacher model is trained with massive amounts of information in cloud data centers to be suitable for real-time environments. This model has the characteristics of high complexity, high computing resource requirements, high accuracy and robustness. With the output of the teacher model as guidance, using dynamic knowledge distillation technology, following the residual learning-based RGB-D semantic segmentation model training strategy, a lightweight, low-computation student model is trained, and the obtained optimized student model or adjusted weights and parameters are pushed to the edge vehicle system.

6. The vehicle anomaly detection method based on cloud collaboration according to any one of claims 1 to 5, characterized in that: The vehicle collects different real-time data in different driving scenarios, and inputs the collected real-time image data into the updated semantic segmentation model for inference to obtain segmentation results. The vehicle detects abnormal events based on the results and makes decisions. After the vehicle detects an abnormal event, it broadcasts the abnormal event to surrounding vehicles through the information interaction system based on the Internet of Vehicles, and uploads it to the cloud and the dispatch center. After the dispatch center receives the abnormal event report, it generates relevant descriptive information about the abnormal event based on the cloud data center; the relevant descriptive information includes abnormal event description files, accident handling plans, current traffic scene data, and flow information.

7. A cloud-coordinated vehicle anomaly detection device, characterized in that: include: A memory and a processor for storing and computer programs; the processor is used to implement a vehicle anomaly detection method based on cloud collaboration as described in any one of claims 1 to 6 when executing the computer program.

8. A cloud-coordinated vehicle anomaly detection system, characterized in that: include: An information interaction module based on the Internet of Vehicles is used to realize information interaction between vehicles, vehicles and road infrastructure, and vehicles and the cloud, and is used for uploading multi-angle real-time scene information, obtaining abnormal event information, and broadcasting abnormal events; Intelligent sensing end, including cameras and lidar, for collecting real-time RGB image and depth image data; The intelligent cloud platform is used to communicate with the intelligent platform in real time, retrieve and match the uploaded real-time data with a large amount of image data, abnormal event text description data, and traffic resource data in the cloud data center, and carry out model training with a semantic segmentation model based on residual learning.

Citation Information

Cited By

  • Vehicle driving assistance method and system and cloud processor

    CN121180232A

  • Internet of vehicles big data real-time monitoring and user service optimization platform for intelligent traffic system

    CN121542630A