Real-time sensing method for traffic emergencies in campus scene based on multi-source information fusion
Through the decision-making fusion of lidar and camera and multi-source information fusion technology, combined with multi-index collision detection model and open-set learning algorithm, the problem of insufficient accuracy and reliability of single sensors in campus scenarios is solved, and accurate traffic emergency detection and efficient accident response are achieved all-weather and all-time periods.
Patent Information
- Application Number
- CN202510355800.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology has insufficient accuracy and reliability of single sensor sense knowledge in campus scenarios, the existing perception system is poor in applicability, and lacks timely perception and handling of traffic emergencies.
The decision-making level fusion of lidar and camera is adopted, combined with YOLO-World and PointPillars algorithms are used to fusion of multi-source information, the CLOCs network is used to improve detection accuracy, and a multi-index collision detection model and OpenMax's open-set learning algorithm are introduced to realize real-time upload through the MQTT protocol.
It realizes accurate traffic emergency detection around the clock and all-time period, improves detection accuracy and reliability in campus scenarios, and optimizes accident response efficiency.
Smart Images

Figure CN120279250A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of traffic safety, and particularly relates to a real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion. Background Technique
[0002] With the continuous expansion of the scale of higher education, the campuses of universities in China have gradually evolved into comprehensive open spaces integrating teaching and research, life services, and social interactions. As of 2022, the scale of full-time students in ordinary universities across the country has exceeded 40 million, and the campus traffic flow has shown a geometric growth trend. According to the statistics of the Campus Safety Big Data Center of the Ministry of Education, the average annual growth rate of campus traffic accidents during 2020-2022 reached 18.7%, among which the proportion of electric vehicle collision accidents was as high as 63.4%, and the incidence of night-time traffic accidents increased by 42% compared with daytime. These accidents not only cause direct economic losses, but also lead to student casualties, seriously threatening the campus safety ecosystem.
[0003] Regarding the perception of traffic emergencies in campus scenarios, there are three key technical bottlenecks in existing research. 1) The perception and recognition accuracy and reliability of single sensors are limited. Video cameras can obtain richer target semantic information, but are easily affected by lighting conditions; radar detection is not affected by weather conditions, but has a high omission rate for small targets. 2) The differences in traffic flow lead to poor applicability of existing perception systems in campus scenarios. The traffic conditions in campus traffic accidents in the campus have three characteristics: low speed, large flow, and tidal nature. Most existing real-time safety level prediction studies are for signal intersections and highway safety level predictions, and their traffic flow characteristics are quite different from those of campus traffic flow. 3) Most researchers pay more attention to the prediction of campus traffic events, and there is currently no research on the timely perception and processing after accidents. Based on the above existing problems, we designed a real-time perception device for traffic emergencies in campus scenarios. Summary of the Invention
[0004] The purpose of the present invention is to provide a real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion, and its functions include: multi-source information fusion, campus scenario emergencies, and real-time perception.
[0005] The technical solution of the present invention is as follows:
[0006] A real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion includes the following steps:
[0007] First, install a perception device that integrates a camera and a lidar, enabling it to simultaneously acquire two-dimensional image video data and three-dimensional lidar point cloud data, achieving more accurate detection and perception under normal circumstances and remaining effective even in low-light and obstructed situations, providing a hardware device foundation for the lidar-camera decision-level fusion target recognition system. Among them, during the day, decision-level fusion is performed on the three-dimensional point cloud image and two-dimensional video image data of the target area, and the three-dimensional point cloud image is used alone at night or in special weather conditions.
[0008] After obtaining the receipt, based on the YOLO-World algorithm and the PointPillars algorithm, achieve the target information perception of traffic participants in the campus scene dataset under a single sensor, and then based on the decision-level fusion method of the CLOCs network, achieve multi-source information fusion. The specific implementation method is as follows: Perform decision-level fusion on the three-dimensional point cloud image and video image data of the target area, use a low-complexity multi-modal fusion framework that can combine the candidate results of 2D and 3D detectors before non-maximum suppression (NMS), and use geometric and semantic consistency to improve the detection accuracy of the CLOCs network to establish a lidar-camera fusion target recognition system to detect traffic emergencies and ensure that the system can accurately detect traffic emergencies all day and all night. The principle is as follows: First, reduce the confidence of the YOLO World detector, retain as many candidate boxes as possible, and disable the non-maximum suppression of Point Pillar to retain the original candidate boxes; Second, project the 3D detection box onto the image plane and calculate the IoU with all 2D detection boxes; Then construct the tensor channel as follows, project the 3D detection onto the 2D detection plane;
[0009]
[0010] Among them, IoU i,j is the intersection over union of the 3D detection box and the 2D detection box, are the confidences of 2D detection and 3D detection respectively, and d i,j is the distance between the normalized 2D detection box and the 3D detection box
[0011] Finally, calculate the IoU between the combined detection boxes in the image plane again and use non-maximum suppression to obtain the combined target detection result.
[0012] After obtaining the joint detection target results of lidar and camera, it is necessary to perceive the original campus traffic emergencies and derivative campus traffic emergencies respectively. Among them, this system divides campus scene emergencies into two categories: original traffic emergencies and derivative traffic emergencies. The original traffic emergencies mainly include occurred emergencies such as traffic collisions (such as collisions between pedestrians and motor vehicles, non-motor vehicles and motor vehicles), and the derivative traffic emergencies mainly include emergencies such as road debris that are likely to cause traffic accidents.
[0013] For the original campus traffic emergencies, in order to provide more decision-making basis for the school security department, this system needs to record and upload the process of the accident, that is, to predict the occurrence of a collision event and start recording. For predicting collisions, this system has established a collision model (a multi-index collision detection model based on Time to Collision (TTC), Separating Axis Theorem (SAT), and Impulse of Momentum (IMP)), and determines whether a collision occurs by judging whether two objects can come into contact in a two-dimensional image, mainly predicting collisions in two dimensions: on the time axis and in geometric position.
[0014] From the perspective of the time axis angle, this system introduces a measurement index in the time dimension: Time to Collision (TTC), that is, the minimum collision time is calculated based on the relative position and relative speed between objects. If it is lower than a certain threshold, it indicates that a collision is about to occur. In this study, the TTC calculation formula is as follows:
[0015]
[0016] TTC = min{T x ,T y ,T z}
[0017] Where T k is the collision time in the k direction, and k can take x, y, z, d k is the distance of the traffic participant in the k direction, is the speed of the traffic participant in the k direction, and TTC is the minimum time to collision.
[0018] Since only collisions can be predicted in the time dimension, therefore, from the geometric position angle, we use the Separating Axis Theorem (SAT) to judge geometric intersection, calculate the projection intervals of two objects on all possible projection axes, and if there is an overlap on all axes, it is considered that the objects are in contact. When uploading the original campus traffic emergencies, this system will judge its severity along with the event, improving the efficiency of the emergency rescue decision-making for the school security department.
[0019] Since the actual collision situation cannot be accurately judged based on geometric intersection alone, this system uses momentum change rate (IMP) analysis to judge the severity of the collision. We select the speed change of the previous and next 3 frames for analysis to determine that the collision has sufficient impact force, thereby avoiding misjudgment of minor contact. The IMP calculation formula is as follows:
[0020]
[0021] ΔIMP=|IMP t -IMP t-1 |
[0022] IMP n is the momentum change rate of the nth traffic participant, m n is the mass of the nth traffic participant. According to actual statistics, the pedestrian can be estimated to be 60kg, the motor vehicle can be estimated to be 2000kg, and the bicycle can be estimated to be 30kg. is the speed of the nth traffic participant at time t.
[0023] However, due to the tidal characteristics of campus traffic events, the density of human traffic is high in some time periods, which makes it easy to make misjudgments, resulting in a huge amount of "accident" information uploaded in a short period of time, and an imbalance in the allocation of emergency response resources of the school security department. In order to optimize this situation, this system introduces a historical frame detection method, that is, maintaining the collision detection results of the most recent N frames, setting the condition that at least K frames have collisions within the most recent M frames, so as to filter out the interference caused by single-frame misjudgment.
[0024] For derivative campus traffic emergencies, this system mainly focuses on foreign object detection. Foreign objects on campus road scenes will cause changes in the traffic trajectories of people and vehicles, which may lead to traffic accidents. Since the traffic element categories of campus scenes are stable, if objects that are different from conventional traffic elements are perceived and identified, they will be regarded as foreign objects. At the algorithm level, the statistical characteristics of known categories are used to determine the existence of unknown categories.
[0025] Specifically, this system detects foreign objects in traffic in 2D video data based on the YOLO-World algorithm of OpenMax. That is, the OpenMax algorithm is used to further improve the YOLO-World model. The core idea of this algorithm is to use the statistical characteristics of known categories to determine the existence of unknown categories, and to use the advantages of open set learning over closed set learning, that is, to learn and classify the test set without knowing all the labels in the test set in advance. This idea mainly realizes the detection of foreign objects on open world roads through the two stages of "training-inference":
[0026] During the training phase, the system pre-trains the YOLO-World model based on known category data, uses it to extract the visual features of candidate boxes, and generates text embeddings of known categories as prototypes through CLIP. At the same time, the tail probability distribution parameters of each category (Weibull model) are fitted based on the extreme value distribution (EVT) to provide a basis for OpenMax calibration.
[0027] During the inference phase, after the input image generates candidate boxes and visual features through YOLO-World, the OpenMax module calculates the similarity between the features and the prototypes of known categories, calibrates the confidence level in combination with the pre-trained EVT model, and dynamically generates the probability of "unknown category". In the post-processing phase, known / unknown detection boxes are separated by a threshold, and unknown targets are independently marked and output. This unknown target is the road foreign object.
[0028] Finally, when a campus traffic event is successfully detected, the system will upload the original information of the campus emergency, that is, the intuitive video, and the accident information automatically identified by the model, that is, the text information such as the accident location, specific position, specific time, accident type, and accident level, in real time, providing strong decision-making basis for the school security department through dual information and improving the accuracy of its decision-making and the efficiency of emergency response. Specifically: if a collision is detected in the traffic accident perception module, the camera will be triggered to take pictures of the accident point and transmit them to the user side together with the text information such as the accident location, specific position, specific time, accident type, and accident level. After receiving the information, the user side will trigger an emergency response (such as an alarm sound) to remind relevant personnel to organize emergency rescue.
[0029] At the technical level, the system constructs an information transmission platform between the device side and the user side using the MQTT protocol. That is, the device side acts as a publisher or subscriber and communicates with the user side through the MQTT proxy server to achieve efficient, reliable, and low-power bidirectional communication, which is suitable for data transmission and control of large-scale radar-vision integrated machine devices installed at important intersections in the campus scene with high-density traffic flow.
[0030] The beneficial effects of the present invention are as follows:
[0031] 1. Based on the multi-source information fusion technology, the present invention innovatively adopts the decision-level fusion of lidar and camera to achieve all-weather target detection, and combines the open-set learning algorithm and the multi-index collision detection model to improve the real-time perception ability of traffic emergencies. First, by fusing the YOLO-World target detection network and the PointPillar 3D detection technology, campus traffic participants are accurately identified, and the robustness of the perception system in complex traffic environments is optimized. Second, the TT, SAT, and IMP multi-index methods are used to detect collision events, and the detection reliability is improved by combining historical frame analysis.
[0032] 2. The present invention introduces an open-set learning method based on OpenMax to achieve the recognition of traffic foreign objects (such as fallen objects, bicycle obstacles, etc.) not included in the training dataset, making up for the defects of existing algorithms in detecting dynamic unknown targets in open campus scenarios.
[0033] 3. The present invention uses the MQTT protocol to establish a real-time cloud upload and emergency management mechanism, further optimizing the accident response efficiency. When a traffic emergency is detected, information such as the accident location, time, type, and severity is automatically uploaded to the campus security department, achieving a second-level response and reducing the lag time in accident handling. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a diagram of the roadside device of the real-time perception system;
[0035] Figure 2 It is a diagram of the visual display of the data of the perception device;
[0036] Figure 3 It is a diagram introducing the foreign object detection process;
[0037] Figure 4 It is a diagram of the cloud upload interface. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings in the specific embodiments.
[0039] Embodiment 1:
[0040] As Figure 1 shown, the roadside device of this real-time perception system consists of three major parts: a roadside perception unit, a roadside computing unit, and a mobile power supply. The mobile power supply provides stable and sufficient power for the roadside perception unit and the roadside computing unit to ensure their normal operation; the roadside computing unit is a computing device responsible for processing the three-dimensional point cloud data and two-dimensional image and video data obtained by the roadside perception device, preprocessing the data, and performing collision detection and foreign object detection in real time, and uploading the perceived original information and detection results to the school security department in real time; the roadside perception unit consists of a camera and a lidar. The camera has 6 million pixels from Zhejiang Dahua, and the lidar is of the RoboSense M1 model, respectively obtaining the two-dimensional image and video data and three-dimensional radar point cloud data of the same campus traffic scene.
[0041] Embodiment 2:
[0042] As Figure 2As shown in the figure, after acquiring data, this real-time perception system will locate traffic elements in two-dimensional image video data and three-dimensional radar point cloud data respectively based on the YOLO-World algorithm and the PointPillars algorithm, and visualize them with candidate boxes, preparing for collision detection and foreign object detection.
[0043] Embodiment 3
[0044] As Figure 3 shown in the figure, this figure explains the implementation process of foreign object detection, which is mainly divided into three steps: First, by collecting a large number of recorded videos of campus scenes, train the YOLO-World model so that it can accurately perceive and detect common traffic elements in campus scenes; Second, when perceiving foreign object detection, calculate the similarity between features and known class prototypes through the OpenMax module, calibrate the confidence level in combination with the pre-trained EVT model, and dynamically generate the probability of "unknown class"; Finally, output the detected results, where traffic elements that are not common in campus scenes marked during training are defined as foreign objects and fed back and uploaded to the security department.
[0045] Embodiment 4
[0046] As Figure 4 shown in the figure, this interface is the visualization interface after the real-time perception system uploads data, providing a decision-making basis for the school security department to conduct emergency responses. This interface is mainly divided into three parts: The upper left corner is the video playback of traffic emergencies in the campus scene, which is convenient for emergency staff to intuitively understand the situation at the accident scene; The lower left corner is the emergency event information presented after being calculated, judged and processed by this system, mainly including three parts: accident level, accident time and accident location, which is convenient for emergency staff to quickly understand the relevant information of the emergency event; The right side is the electronic map of the accident location and its surroundings, which is an embedded Amap map, used for emergency staff to quickly locate the accident location in order to carry out rescue in a timely manner.
[0047] Matters not covered by this invention are well-known technologies.
[0048] The above embodiments are only used to illustrate the technical concept and features of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion, characterized in that: Including the following steps: S1. Simultaneously acquire two-dimensional image video data and three-dimensional lidar point cloud data through a perception device integrating a camera and a lidar. Based on the YOLO-World algorithm and the PointPillars algorithm, realize the target information perception of traffic participants in a single sensor under the campus scene dataset. Then, based on the decision-level fusion method of the CLOCs network, realize multi-source information fusion to obtain the joint target detection results of the camera and the lidar; S2. Perceive campus traffic emergencies, including primary campus traffic emergencies and derivative campus traffic emergencies; S3. When a campus scene traffic event is successfully perceived, real-time upload the original information of the campus scene emergency and the accident information automatically judged by the model, and provide strong decision-making basis for the school security department through dual information, improving the decision-making accuracy and emergency response efficiency.
2. The real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion according to claim 1, wherein: The multi-source information fusion in step S1 is to perform decision-level fusion on the three-dimensional point cloud image and two-dimensional video image data of the target area during the day, and separately adopt the three-dimensional point cloud image at night or in special weather conditions.
3. The real-time perception method for campus scene traffic emergencies based on multi-source information fusion according to claim 2, characterized in that: The decision-level fusion in step S1 is to use a low-complexity multi-modal fusion framework that can combine the candidate results of 2D and 3D detectors before non-maximum suppression, and use geometric and semantic consistency to improve the detection accuracy of the CLOCs network to establish a lidar-camera fusion target recognition system to detect traffic emergencies; specifically: First, reduce the confidence of the YOLO World detector, retain as many candidate boxes as possible, and disable the non-maximum suppression of Point Pillar to retain the original candidate boxes; Second, project the 3D detection box onto the image plane and calculate the IoU with all 2D detection boxes; Then construct the tensor channel as follows, project the 3D detection onto the 2D detection plane; Among them, IoU i,j is the intersection over union of the 3D detection box and the 2D detection box, are the confidence levels of 2D detection and 3D detection respectively, and d i,j is the distance between the normalized 2D detection box and the 3D detection box; Finally, calculate the IoU between the joint detection boxes in the image plane again, and adopt non-maximum suppression, that is, obtain the joint target detection results.
4. The real-time perception method for campus scene traffic emergencies based on multi-source information fusion according to claim 1, characterized in that: In step S2, the primary traffic emergencies include the occurred emergencies of traffic collisions, and the derivative traffic emergencies include the emergencies of traffic accidents caused by foreign objects on the road surface.
5. The real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion according to claim 4, characterized in that: In step S2, save the relevant information of the perceived primary campus scene emergencies and upload it to the school security department in real time to ensure that the campus scene emergencies can be processed in time.
6. The real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion according to claim 4, characterized in that: The perception method of primary traffic emergencies includes: establishing a multi-index collision detection model based on the time to collision TTC, the separating axis theorem SAT, and the rate of change of momentum IMP, and improving its reliability according to historical frame detection; First, introduce the measurement index in the time dimension: the time to collision TTC, calculate the minimum time to collision through the relative position and relative speed between objects, if it is lower than a certain threshold, it means that a collision is about to occur; The TTC calculation formula is as follows: TTC = min{T x , T y , T z} where T k is the collision time in the k direction, and k can take x, y, z, d k is the distance of the traffic participant in the k direction, is the speed of the traffic participant in the k direction, and TTC is the time to closest approach; The separating axis theorem SAT is used to judge geometric intersection and calculate the projection intervals of two objects on all possible projection axes. If there is overlap on all axes, it is considered that the objects are in contact. In order to measure the severity of the collision, the momentum change rate IMP analysis is used to select the speed change of the previous and next 3 frames for analysis to determine that the collision has sufficient impact force, thereby avoiding misjudgment of slight contact. The IMP calculation formula is as follows: ΔIMP = |IMP t - IMP t-1 | where IMP n is the rate of change of momentum of the nth traffic participant, m n is the mass of the nth traffic participant; is the velocity of the nth traffic participant at time t; In order to further improve the reliability of the model, a historical frame detection method is introduced, that is, the collision detection results of the most recent N frames are maintained, and the condition that at least K frames collide within the most recent M frames is set to filter out the interference caused by single-frame misjudgment.
7. The real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion according to claim 4, characterized in that: The perception methods of derivative traffic emergencies include: detecting traffic foreign objects in 2D video data based on the YOLO-World algorithm of OpenMax; The OpenMax algorithm is used to further improve the YOLO-World model, taking advantage of the open set learning over the closed set learning, that is, learning and classifying the test set without knowing all the labels in the test set in advance; and realizing the open world road foreign object detection through the two stages of "training-inference": In the training phase, the YOLO-World model is pre-trained based on known category data, which is used to extract the visual features of the candidate boxes, and the text embedding of the known categories is generated as a prototype through CLIP. At the same time, the tail probability distribution parameters of each category are fitted based on the extreme value distribution EVT to provide a basis for OpenMax calibration. In the inference stage, after the input image generates candidate boxes and visual features through YOLO-World, the OpenMax module calculates the similarity between the features and the known category prototypes, combines the pre-trained EVT model to calibrate the confidence, and dynamically generates the "unknown category" probability; in the post-processing stage, the known / unknown detection boxes are separated by thresholds, and the unknown targets are independently marked and output; the unknown targets are foreign objects on the road.
8. The real-time perception method for traffic emergencies in campus scenarios based on multi-source information fusion according to claim 1, characterized in that: The real-time uploading in step S3 includes: using the MQTT protocol to build an information transmission platform between the device end and the user end; if a collision is detected in the traffic emergency accident perception module, the camera will be triggered to take a photo of the accident point, and the text information of the accident location, specific location, specific time, accident type, and accident level will be transmitted to the user end. After receiving the information, the user end will trigger an emergency response to remind relevant personnel to organize emergency rescue.
Citation Information
Cited By
Multi-touch-point positioning method and system fusing distributed sound wave and wireless common judgment
CN121508656A