Intelligent identification and evidence collection system for traffic violations based on multi-modal fusion perception

By using a multimodal fusion perception system, which utilizes sensors such as visible light cameras, lidar, and millimeter-wave radar, the system solves the problems of low recognition accuracy and poor robustness of single-vision evidence collection systems in complex scenarios, and achieves high-precision recognition of traffic violations and evidence generation.

CN122290330APending Publication Date: 2026-06-26HUBEI TIANCUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI TIANCUN INFORMATION TECH CO LTD
Filing Date
2026-02-06
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing single-vision forensics systems suffer from low accuracy and poor robustness in complex scenarios such as severe weather, poor lighting, and target occlusion.

Method used

A multimodal fusion perception system is adopted, including sensors such as visible light cameras, lidar and millimeter-wave radar. A unified hardware trigger signal and time reference are provided through a synchronization control module. Combined with a data processing module, cross-modal data association and fusion estimation are performed to generate a fused motion trajectory, which is then compared with a digital road traffic rule model to determine traffic violations.

Benefits of technology

Maintaining high identification accuracy in complex environments, reducing false alarm rates, generating stable illegal evidence packages, and improving the system's adaptability and evidence credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290330A_ABST
    Figure CN122290330A_ABST
Patent Text Reader

Abstract

This application relates to the field of intelligent transportation technology and discloses an intelligent identification and evidence collection system for traffic violations based on multimodal fusion perception. The system includes an edge sensing terminal, which uses a synchronous control module to achieve hardware synchronous acquisition of multimodal sensors. A data processing module performs spatiotemporal alignment of the acquired data and performs cross-modal correlation and fusion tracking of the target to generate a high-precision fused motion trajectory. Finally, the trajectory is compared with a digital road traffic rule model to determine the violation. This invention solves the technical problems of low accuracy and susceptibility to false alarms and missed detections caused by single sensors in complex scenarios such as adverse weather and target occlusion. It can generate highly credible multi-dimensional structured evidence, significantly improving the automation level and credibility of traffic law enforcement.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This invention relates to the field of intelligent transportation technology, specifically to an intelligent identification and evidence collection system for traffic violations based on multimodal fusion perception. Summary of the Invention

[0002] To address the shortcomings of existing technologies, this invention provides an intelligent identification and evidence collection system for traffic violations based on multimodal fusion perception. This system solves the problems of low accuracy and poor robustness of existing single-vision evidence collection systems in complex scenarios such as severe weather, poor lighting, and target occlusion.

[0003] To achieve the above objectives, the present invention provides the following technical solution: an intelligent identification and evidence collection system for traffic violations based on multimodal fusion perception, comprising: An edge-sensing terminal, the edge-sensing terminal comprising: The multimodal sensor module is configured to simultaneously acquire sensor data of multiple different modalities in a traffic scenario; The synchronization control module is configured to provide a unified hardware trigger signal to the multimodal sensor module to achieve the synchronous acquisition and to provide a unified time reference for the acquired sensor data. The data processing module, electrically connected to the multimodal sensor module and the synchronization control module, is configured as follows: Using preset external parameters, the sensor data of the various different modalities are aligned to a unified world coordinate system; The aligned sensor data is processed to achieve cross-modal data association for the same physical target and to assign a globally unique identifier to the target; Based on the results of the cross-modal data association, the motion state of the target is fused and estimated to generate a fused motion trajectory of the target; The fused motion trajectory is compared with a preset digital road traffic rule model to determine whether the target has committed any traffic violations.

[0004] Preferably, the multimodal sensor module includes: at least one visible light camera, a lidar, and a millimeter-wave radar.

[0005] Preferably, the synchronization control module has a built-in satellite positioning and timing unit, and the unified hardware trigger signal is the second pulse signal output by the satellite positioning and timing unit.

[0006] Preferably, the data processing module aligns the sensor data using the preset external parameters by performing coordinate transformation on the sensor data using a pre-calibrated rotation matrix and translation vector that defines the relationship between each sensor coordinate system and the unified world coordinate system.

[0007] Preferably, the data processing module implements the cross-modal data association by: constructing a cost matrix, the elements of which are used to quantify the consistency of sensor data detection results of different modalities in spatial location and motion state; and using an optimization algorithm to solve the cost matrix to obtain the optimal match, thereby realizing the cross-modal data association.

[0008] Preferably, the data processing module performs fusion estimation of the target's motion state by using a Kalman filter, taking observations from different modal sensors as input, and iteratively updating a state vector containing the target's position and velocity information.

[0009] Preferably, the digital road traffic rule model includes: vectorized geographic information of lane lines and stop lines, and real-time status data of traffic lights.

[0010] Preferably, the data processing module is further configured to: Calculate the fusion confidence level of a violation determination based on multiple determination sources; The weighting coefficients of the various determination sources are dynamically adjusted based on the current environmental factors sensed by the multimodal sensor module.

[0011] Preferably, after determining that a traffic violation has occurred, the data processing module is further configured to: Generate a structured evidence package, which includes: video clips recording the illegal process, keyframe images overlaid with illegal information text, and a metadata file.

[0012] The intelligent identification and evidence collection method for traffic violations based on multimodal fusion perception includes the following steps: S1. Responding to a unified hardware trigger signal, it synchronously collects sensor data of multiple different modes in traffic scenarios. S2. Using preset external parameters, align the sensor data of the various different modes to a unified world coordinate system; S3. Process the aligned sensor data to achieve cross-modal data association for the same physical target and assign a globally unique identifier to the target; S4. Based on the results of the cross-modal data association, the motion state of the target is fused and estimated to generate a fused motion trajectory of the target; S5. Compare the fused motion trajectory with a preset digital road traffic rule model to determine whether the target has committed any traffic violations.

[0013] This invention provides an intelligent system for identifying and collecting evidence of traffic violations based on multimodal fusion perception. It has the following beneficial effects: 1. This invention sets up complementary multimodal sensor modules such as visible light cameras, lidar, and millimeter-wave radar, and adopts a multi-level confidence decision mechanism to dynamically adjust the weight of different sensor judgment sources according to the current environmental factors. This enables the system to output stable and reliable perception results even when the performance of a single sensor declines due to severe weather or poor lighting, thereby maintaining a high recognition accuracy. This solves the technical problem of poor adaptability of existing technologies in complex environments.

[0014] 2. This invention generates a continuous and stable fused motion trajectory that does not rely on a single visual information by performing cross-modal correlation and fusion estimation on multimodal data. Even if the target is briefly visually obscured, the system can still maintain target tracking and state prediction based on data from lidar or millimeter-wave radar, thereby effectively solving the problem of missed detection of illegal behavior caused by target occlusion and significantly reducing the false alarm rate caused by trajectory calculation errors.

[0015] 3. This invention utilizes a unified hardware trigger signal provided by a synchronization control module to ensure strict alignment of all sensor data on a time reference. This enables the system to precisely correlate information from different sources on the same timeline when generating evidence. The resulting structured evidence package features multi-dimensional information corroboration, constructing a complete chain of evidence and significantly improving the credibility of evidence of illegality. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the system architecture of the present invention; Figure 2 This is a schematic diagram of the method steps of the present invention. Detailed Implementation

[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see the appendix Figure 1 A traffic violation intelligent identification and evidence collection system based on multimodal fusion perception includes: Edge sensing terminals, including: The multimodal sensor module is configured to simultaneously acquire sensor data of multiple different modalities in a traffic scenario; The multimodal sensor module integrates various sensors with different operating principles through an integrated, common platform design, ensuring that all sensors have a consistent field of view reference and minimal physical deviation. In one specific embodiment, the module includes: a high-resolution visible light camera with wide dynamic range (WDR) capability, which captures color texture images of the scene to provide a basis for obtaining elements of law enforcement documents such as vehicle license plates and body color, and generating visual evidence that conforms to legal norms; a long-wave infrared thermal imager, which images by detecting infrared radiation from the surface of objects, and can reliably detect heat-generating components such as engines and tires in completely dark environments or in low-visibility weather conditions such as rain, snow, fog, and haze, thereby reliably detecting and tracking motor vehicles; a lidar, which actively emits laser pulses and receives reflected signals to generate high-density three-dimensional point cloud images of the environment, providing key information for distinguishing vehicle types, accurately calculating vehicle distances, and separating independent targets in congested scenarios; and a 77GHz band millimeter-wave radar, which has extremely strong penetration capabilities in adverse weather conditions and can directly and accurately measure the radial velocity of targets by analyzing the Doppler frequency shift of the echo signal, providing high-confidence raw data for determining speeding and other violations.

[0019] The synchronization control module is configured to provide a unified hardware trigger signal to the multimodal sensor module to achieve synchronous acquisition and to provide a unified time reference for the acquired sensor data. The synchronization control module provides a unified and accurate spatiotemporal reference for the entire terminal, a prerequisite for effective multimodal data fusion. At its core is a satellite timing and positioning unit integrating Global Positioning System (GPS) and BeiDou Navigation Satellite System (BDS) functions. This unit provides spatial stamps with geographic coordinates for all events and data, and outputs a highly stable pulse-of-seconds (PPS) hardware signal. This PPS signal is used as a global synchronization clock, triggering all sensors within the multimodal sensor module to acquire data at the same physical moment via hardware interrupts. This ensures strict timestamp alignment between data streams, achieving a synchronization accuracy at the nanosecond level.

[0020] The data processing module, electrically connected to the multimodal sensor module and the synchronization control module, is configured to: align sensor data from multiple different modes to a unified world coordinate system using preset external parameters; process the aligned sensor data to achieve cross-modal data association for the same physical target and assign a globally unique identifier to the target; Based on the results of cross-modal data association, the motion state of the target is fused and estimated to generate a fused motion trajectory of the target; the fused motion trajectory is compared with a preset digital road traffic rule model to determine whether the target has committed any traffic violations. The data processing module is the computing core of the edge sensing terminal. Its physical carrier is a system-on-a-chip (SoC) designed specifically for edge computing, integrating a multi-core central processing unit (CPU), a high-performance graphics processing unit (GPU), and a dedicated neural network processing unit (NPU). This heterogeneous computing architecture enables it to support real-time parallel computation of complex deep learning models and multimodal data fusion algorithms. All raw data streams collected from the multimodal sensor modules are processed within this module, transforming massive amounts of unstructured sensor data streams into structured target and event information on-site. The synchronization control module provides a unified and accurate spatiotemporal reference for the entire terminal, a prerequisite for effective multimodal data fusion. At its core is a satellite timing and positioning unit integrating Global Positioning System (GPS) and BeiDou Navigation Satellite System (BDS) functions. This unit provides spatial stamps with geographic coordinates for all events and data, and outputs a highly stable pulse-of-seconds (PPS) hardware signal. This PPS signal is used as a global synchronization clock, triggering all sensors within the multimodal sensor module to acquire data at the same physical moment via hardware interrupts. This ensures strict timestamp alignment between data streams, achieving a synchronization accuracy at the nanosecond level.

[0021] The communication module is responsible for data exchange between the terminal and the outside world. It has a built-in communication chip supporting 5G and Gigabit Ethernet, allowing the terminal to flexibly choose the access method according to the on-site network conditions and supporting dual-link hot backup. This module also executes data transmission protocols, encrypts and encapsulates generated evidence packets of illegal activities, and sets transmission priorities for different types of data through Quality of Service (QoS) policies, prioritizing the transmission of high-level alarm data to ensure the real-time nature and reliability of critical law enforcement information.

[0022] The storage module uses industrial-grade solid-state drives (SSDs) to provide reliable local data storage capabilities for the terminal. It has two main functions: first, in the event of network connectivity issues, it acts as a data buffer, temporarily storing evidence information to be uploaded and resuming transmission once the network is restored to prevent data loss; second, it functions as a looping black box, storing raw sensor data for a short period (e.g., 24 hours) for remote system diagnostics, algorithm iteration optimization, or in-depth tracing analysis of specific events.

[0023] Please see the appendix Figure 2 The intelligent identification and evidence collection method for traffic violations based on multimodal fusion perception includes the following steps: S1. Responding to a unified hardware trigger signal, it synchronously collects sensor data of multiple different modes in traffic scenarios. S2. Using preset external parameters, align sensor data from multiple different modalities to a unified world coordinate system; S3. Process the aligned sensor data to achieve cross-modal data association for the same physical target and assign a globally unique identifier to the target; S4. Based on the results of cross-modal data association, the motion state of the target is fused and estimated to generate a fused motion trajectory of the target; S5. The motion trajectory is compared with the preset digital road traffic rules model to determine whether the target has committed any traffic violations.

[0024] The method first performs data acquisition and spatiotemporal reference alignment. This step aims to provide a multi-source heterogeneous data stream that is strictly synchronized in time and has unified spatial coordinates for subsequent fusion sensing algorithms. Triggered by the PPS signal of the synchronization control module, all sensors in the multimodal sensor module acquire synchronized data frames and attach high-precision global timestamps.

[0025] Based on this unified time reference, the data processing module uses preset external parameters to align sensor data from various modalities to a unified world coordinate system. This spatial alignment process is based on the principle of rigid body transformation, and its core lies in performing coordinate transformation on the data from each sensor:

[0026] Where P worid , represents the three-dimensional coordinates of the target point in the world coordinate system; P sensorThis represents the three-dimensional coordinates of the point in the sensor's own coordinate system; R is a 3x3 rotation matrix; and T is a 3x1 translation vector. Preferably, R and T are precisely calculated through a one-time offline joint calibration process. It should be noted that to ensure the effectiveness of the transformation, the orthogonality of the rotation matrix R must be verified, and its determinant should be +1 to eliminate singularities or coordinate system reflections. The technical purpose of this step is to eliminate spatial deviations introduced by the different physical installation positions and orientations of the sensors, providing a unified spatial reference for subsequent data fusion.

[0027] After all sensor data is mapped to the same space, the data processing module processes the aligned data to achieve cross-modal data association and fusion tracking of the same physical target. This process first uses parallel processing threads to extract features independently from each modal data. For example, YOLO-based neural network models are used to process visible light images to extract the target's 2D bounding box and category; PointPillars and other 3D detection networks are used to process LiDAR point clouds to extract the target's 3D bounding box and pose.

[0028] To attribute these independent detection results to the same physical target, this module constructs a cost matrix. The cost calculation integrates the spatial overlap (e.g., 3D Intersection over Union) and kinematic consistency of different detection results. The matrix is ​​then solved using optimization algorithms such as the Hungarian algorithm or Joint Probabilistic Data Association (JPDA) to find the globally optimal match. For each successfully matched target set, the system assigns a globally unique identifier. Unmatched detection results are treated as candidates for potential new targets, and temporary tracking and observation are initiated.

[0029] Based on the results of cross-modal data association, this module then performs a fusion estimation of the target's motion state to generate a fused motion trajectory for the target. Preferably, this step is implemented using an Extended Kalman Filter (EKF). The target's state vector X t It is defined as a set containing information such as its three-dimensional position, velocity, and acceleration. The filter, through continuous iteration of prediction and update steps, optimally fuses noisy observations from different sensors, outputting the best posterior estimate of the target state. The generated fused motion trajectory is essentially a continuous sequence of the target state changing over time, with accuracy and smoothness superior to the results of any single sensor. The parameters of the process noise covariance matrix Q can be adaptively adjusted based on environmental perception results to cope with changes in the motion model under different road conditions.

[0030] Based on this high-precision trajectory, the final violation determination is achieved. The data processing module compares the fused trajectory with a preset digital road traffic rule model that includes lane line vector data, precise stop line positions, and real-time traffic light status. By analyzing the interaction between the trajectory and rule elements in the spatiotemporal dimensions, it determines whether the target has committed a traffic violation.

[0031] To ensure accuracy and suppress false alarms, the system employs a multi-level confidence decision model. The final violation determination confidence score, Sfusion, is calculated using the following formula:

[0032] Where N is the total number of decision sources involved in the decision-making process; Si is the confidence score output by the i-th decision source (such as visual judgment, trajectory analysis, etc.), and its value ranges from [0,1]; Wi is the weight dynamically assigned to the i-th decision source based on the current environmental factors, and satisfies the following conditions: For example, in low-light conditions at night, the system automatically reduces the weight of visible light image analysis results while increasing the weight of infrared thermal imaging and radar analysis results. Only when Sfusion exceeds a preset threshold τ is the violation finally confirmed. The value of this threshold τ is determined by plotting Receiver Operating Characteristic (ROC) curves on the validation dataset and based on the expected balance between true positive and false positive rates; it is typically between 0.85 and 0.95.

[0033] Once a violation is confirmed, the system immediately initiates the process of generating and uploading evidence. A structured evidence generation module is activated, automatically extracting a video clip from the visible light camera's video cache that fully records the violation and capturing several keyframe images. The system then annotates these images with overlaid text (OSD), indicating enforcement elements such as time, location, license plate number, and type of violation. Finally, these video clips, still images, and a metadata file containing all structured information about the violation are packaged into a unified, standardized evidence package, encrypted, and uploaded to the central business layer.

[0034] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A traffic violation intelligent identification and evidence collection system based on multi-modal fusion perception, characterized in that, The application relates to a multi-modal sensor system for intelligent traffic surveillance, comprising: an edge-aware terminal, comprising: a multi-modal sensor module configured to synchronously collect sensor data of multiple different modalities of a traffic scene; a synchronous control module configured to provide a unified hardware trigger signal to the multi-modal sensor module to realize the synchronous collection and provide a unified time reference for the collected sensor data; a data processing module electrically connected with the multi-modal sensor module and the synchronous control module and configured to: align the sensor data of the multiple different modalities to a unified world coordinate system by using preset external parameters; process the aligned sensor data to realize cross-modality data association of a same physical target and assign a globally unique identity to the target; based on the cross-modality data association result, fuse and estimate a motion state of the target to generate a fused motion trajectory of the target; and compare the fused motion trajectory with a preset digital road traffic rule model to determine whether the target has committed a traffic violation.

2. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, The multi-modal sensor module comprises at least one visible light camera, a laser radar and a millimeter wave radar.

3. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, The synchronous control module is internally provided with a satellite positioning and timing unit, and the unified hardware trigger signal is a second pulse signal output by the satellite positioning and timing unit.

4. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, The data processing module aligns the sensor data by using a rotation matrix and a translation vector that are defined by a relationship between a sensor coordinate system and the unified world coordinate system and are pre-calibrated to perform coordinate transformation on the sensor data.

5. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, The data processing module realizes the cross-modality data association by constructing a cost matrix whose elements are used to quantify the consistency of detection results of different modalities of sensor data in spatial position and motion state and by solving the cost matrix by using an optimization algorithm to obtain an optimal match and thus realize the cross-modality data association.

6. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, The data processing module fuses and estimates the motion state of the target by using a Kalman filter to take observation values from different modalities of sensors as input to iteratively update a state vector containing position and velocity information of the target.

7. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, The digital road traffic rule model comprises vectorized geographic information of lane lines and stop lines and real-time state data of traffic signal lamps.

8. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, The data processing module is further configured to: calculate a fused confidence of a violation judgment based on multiple judgment sources; and wherein weight coefficients of the multiple judgment sources are dynamically adjusted according to current environmental factors perceived by the multi-modal sensor module.

9. The multi-modal fusion perception based intelligent identification and evidence collection system for traffic violations according to claim 1, characterized in that, After determining that a traffic violation exists, the data processing module is further configured to: generate a structured evidence package, the structured evidence package comprising a video clip recording a violation process, a key frame picture with violation information text superimposed and a metadata file.

10. The intelligent identification and evidence collection method for traffic violations based on multi-modal fusion perception, applied to the intelligent identification and evidence collection system for traffic violations based on multi-modal fusion perception according to claims 1-9, characterized in that, The application further relates to a method for intelligent traffic surveillance, comprising the following steps: S1. Responding to a unified hardware trigger signal, it synchronously collects sensor data of multiple different modes in traffic scenarios. S2. Using preset external parameters, align the sensor data of the various different modes to a unified world coordinate system; S3. Process the aligned sensor data to achieve cross-modal data association for the same physical target and assign a globally unique identifier to the target; S4. Based on the results of the cross-modal data association, the motion state of the target is fused and estimated to generate a fused motion trajectory of the target; S5. Compare the fused motion trajectory with a preset digital road traffic rule model to determine whether the target has committed any traffic violations.