Traffic checkpoint holographic recognition method and system based on multi-mode perception and medium
By deploying multimodal sensor arrays at traffic checkpoints for data collection and deep learning model fusion, the problem of insufficient multimodal data fusion in existing traffic monitoring systems is solved, and high-precision and efficient traffic identification and abnormal warning are achieved.
Patent Information
- Application Number
- CN202510829230.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-17
AI Technical Summary
The existing traffic monitoring system is unable to efficiently integrate data from multiple sensors, resulting in low traffic identification accuracy and insufficient real-time performance, making it difficult to meet the combined needs of smart transportation for vehicle holographic portraits and real-time risk warnings.
By deploying multimodal sensor arrays at traffic checkpoints, vehicle images, radar point clouds and infrared thermal imaging data are collected in real time, a unified time base and spatial coordinates are used for spatiotemporal synchronization processing, and a deep learning model is used for data fusion to generate holographic feature vectors for target vehicle identification and abnormal event warnings.
It realizes the spatiotemporal synchronization and deep feature fusion of multimodal data, improves the recognition accuracy and real-time performance of traffic management, and enhances the efficiency and safety of traffic management.
Smart Images

Figure CN120808306A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of traffic management, in particular to a traffic checkpoint holographic identification method and system based on multi-modal perception and a medium. BACKGROUND
[0002] Under the background of accelerated urbanization and rapid evolution of intelligent transportation systems, the traditional traffic checkpoint monitoring system faces multiple challenges: single visual perception method is easily disturbed by environmental factors such as light and shielding, resulting in fluctuations in license plate recognition accuracy; radar and video data lack a spatio-temporal alignment mechanism, making it difficult to construct a three-dimensional vehicle motion trajectory; infrared thermal imaging data has not been fully tapped, and there is a blind area in driver fatigue state monitoring. The existing technical solutions generally have problems such as insufficient depth of multi-modal data fusion, lack of spatio-temporal synchronization accuracy, and single dimension of behavior analysis, which are difficult to meet the complex needs of intelligent transportation for vehicle holographic portrait and real-time risk warning. SUMMARY
[0003] The present application provides a traffic checkpoint holographic identification method and system based on multi-modal perception to solve the technical problems that the existing traffic monitoring system cannot efficiently fuse multiple sensor data for holographic identification, resulting in low traffic identification accuracy and insufficient real-time performance, and achieves the technical effect of improving traffic management efficiency and safety through spatio-temporal synchronization and deep feature fusion of multi-modal data.
[0004] In a first aspect, the present application provides a traffic checkpoint holographic identification method based on multi-modal perception, wherein the traffic checkpoint holographic identification method based on multi-modal perception comprises:
[0005] Through a multi-modal sensor array deployed at the target traffic checkpoint, real-time acquisition of an original perception data set is performed, the original perception data set including vehicle images, radar point clouds, and infrared thermal images; based on a unified time reference and spatial coordinates, spatio-temporal synchronization processing of the original perception data set is performed to generate a multi-modal data frame sequence; a deep learning model is used to perform fusion processing on the multi-modal data frame sequence to generate a holographic feature vector of the target vehicle; based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle are performed, and an identification result is output and an abnormal event warning is performed.
[0006] In a second aspect, the present application further provides a traffic checkpoint holographic identification system based on multi-modal perception, wherein the traffic checkpoint holographic identification system based on multi-modal perception comprises:
[0007] The data acquisition module: through the multi-modal sensor array deployed at the target traffic checkpoint, real-time acquisition of the original perception dataset, the original perception dataset including vehicle images, radar point clouds and infrared thermal imaging; the data synchronization processing module: based on a unified time reference and spatial coordinates, spatio-temporal synchronization processing of the original perception dataset, generating a multi-modal data frame sequence; the data fusion module: using a deep learning model to fuse the multi-modal data frame sequence, generating a holographic feature vector of the target vehicle; the abnormal early warning module: based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle, outputting the recognition result, and abnormal event early warning.
[0008] In a third aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the multi-modal perception based traffic checkpoint holographic identification method provided by the present application.
[0009] The present application discloses a multi-modal perception based traffic checkpoint holographic identification method, system and medium, comprising: through the multi-modal sensor array deployed at the target traffic checkpoint, real-time acquisition of the original perception dataset, the original perception dataset including vehicle images, radar point clouds and infrared thermal imaging; based on a unified time reference and spatial coordinates, spatio-temporal synchronization processing of the original perception dataset, generating a multi-modal data frame sequence; using a deep learning model to fuse the multi-modal data frame sequence, generating a holographic feature vector of the target vehicle; based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle, outputting the recognition result, and abnormal event early warning. The multi-modal perception based traffic checkpoint holographic identification method, system and medium disclosed by the present application solve the technical problems that the existing traffic monitoring system cannot efficiently fuse multiple sensor data for holographic identification, resulting in low traffic identification accuracy and insufficient real-time performance, and achieve the technical effects of improving traffic management efficiency and safety through spatio-temporal synchronization and deep feature fusion of multi-modal data. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 It is a flowchart of the multi-modal perception based traffic checkpoint holographic identification method of the present application.
[0011] Figure 2 It is a structural schematic diagram of the multi-modal perception based traffic checkpoint holographic identification system of the present application.
[0012] Legend of the drawing: data acquisition module 11, data synchronization processing module 12, data fusion module 13, abnormal early warning module 14. DETAILED DESCRIPTION
[0013] The above technical solutions will be described in detail below in combination with the accompanying drawings and specific embodiments, so that the above technical solutions can be better understood. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments for explaining the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application. In addition, it should be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings, not all.
[0014] Embodiment one, as Figure 1 The flowchart of the traffic portal holographic identification method based on multi-modal perception of the present application, wherein the traffic portal holographic identification method based on multi-modal perception comprises:
[0015] Through the multi-modal sensor array deployed at the target traffic portal, the original perception data set is collected in real time, and the original perception data set includes vehicle images, radar point clouds and infrared thermal imaging.
[0016] Specifically, at the target traffic portal, a multi-modal sensor array is deployed to simultaneously collect multiple types of perception data. These sensors include high-definition cameras, millimeter wave radars and infrared thermal imagers, etc., which are responsible for acquiring different types of information. The high-definition camera is used to collect visual images of vehicles, capturing the appearance and license plate information of vehicles; the millimeter wave radar emits electromagnetic waves, receives reflected signals, and generates radar point cloud data to perceive the position, speed and three-dimensional structure of the surrounding environment of the vehicle; the infrared thermal imager can detect and capture the thermal radiation emitted by the vehicle, especially in low visibility or night environment, it can effectively capture the thermal image of the vehicle, further supplementing the temperature information that other sensors cannot obtain. These different types of data are collected synchronously and jointly constitute a comprehensive original perception data set, providing sufficient information basis for subsequent data processing and analysis.
[0017] In some embodiments, through the multi-modal sensor array deployed at the target traffic portal, the original perception data set is collected in real time, including:
[0018] The multi-modal sensor array comprising high-definition cameras, millimeter wave radars and infrared thermal imagers is configured, the sampling frequency and data interface protocol of each sensor in the multi-modal sensor array are defined; the original data collected by the multi-modal sensor array is preprocessed to determine the original perception data set.
[0019] Specifically, high-definition cameras, millimeter-wave radars, and infrared thermal imagers suitable for the target traffic checkpoint environment are selected and deployed reasonably according to actual needs. These sensors will cover the traffic checkpoint area from different angles and positions, ensuring the comprehensive collection of multi-dimensional data of vehicles and the surrounding environment. Then, each sensor is configured to define its sampling frequency, that is, how many times data are collected per second. For example, a high-definition camera may require a high frame rate (such as 30 frames per second) to capture clear vehicle images; a millimeter-wave radar may be set to a lower frequency to reduce data volume, but ensure accurate perception of vehicle motion and position; the sampling frequency of an infrared thermal imager should also be configured according to the rate of change of its perceived target. After that, the data interface protocol of each sensor is determined, that is, the data exchange method between each sensor and the system terminal. These protocols can be standard communication protocols such as USB, Ethernet, CAN bus, etc., to ensure smooth data transmission between sensors and to the system terminal. Then, the configured multi-modal sensor array is used to monitor the target traffic checkpoint in real time, obtaining multiple raw data, including multiple raw vehicle images, multiple raw radar point clouds, and multiple raw infrared thermal images. For multiple raw vehicle images, noise removal (such as Gaussian filtering, median filtering, wavelet transform, etc.) and image enhancement (such as histogram equalization, contrast-limited adaptive histogram equalization, etc.) are performed to ensure image quality under different lighting conditions. If necessary, image cropping, scaling, etc. are also performed to meet subsequent processing requirements. For multiple raw radar point clouds, noise removal is performed to accurately reflect the three-dimensional position and motion state of the vehicle. For multiple raw infrared thermal images, noise removal and enhancement processing are also performed to accurately reflect the spatial distribution of heat sources and temperature gradients. After completing the preprocessing of raw data, the preprocessed raw vehicle images, raw radar point clouds, and raw infrared thermal images are integrated into an original perception data set, providing basic data for subsequent deep learning model fusion.
[0020] Based on a unified time reference and spatial coordinates, the original perception data set is processed in space-time synchronization to generate a multi-modal data frame sequence.
[0021] Specifically, first, in order to ensure that the data from different sensors can be accurately aligned in time and space, a unified time reference will be obtained, and a spatial coordinate system conversion relationship will be established to ensure that different sensor data can be aligned in the same time and physical space. Once the unified time reference and spatial coordinate system conversion relationship are determined, the original data collected by all sensors at the same time will be matched according to the data timestamp of each sensor, ensuring that the image, radar point cloud and thermal imaging data at the same time can be accurately corresponded. In space, different sensor data after coordinate system conversion will be spatially registered, so that the same target (such as a vehicle) data collected from different perspectives will be spatially aligned, ensuring that their positions in the same physical space are consistent. After completing the time and space synchronization processing, each sensor data at each time is combined into a multi-modal data frame, each data frame containing synchronized data from different sensors, such as image data, radar point cloud data and infrared thermal imaging data of a vehicle. Finally, all multi-modal data frames at different times are arranged in time sequence to generate a multi-modal data frame sequence, providing a unified and time-aligned data set for subsequent data fusion and analysis.
[0022] In some embodiments, the spatio-temporal synchronization processing is performed on the original perception data set to generate a multi-modal data frame sequence, including:
[0023] obtaining a unified time reference of the multi-modal sensor array configuration as a unified time reference; establishing a spatial coordinate system conversion relationship of the multi-modal sensor array; performing spatial registration on the original perception data set collected at the same time based on the spatial coordinate system conversion relationship according to the unified time reference; and combining the registered original perception data set into a multi-modal data frame, and arranging the multi-modal data frame sequence in time sequence.
[0024] Specifically, first, a unified time reference is configured as a unified time reference for synchronizing the data of each sensor. Typically, this time reference can be achieved by configuring a unified clock source or using a standard synchronization protocol (e.g., GPS clock or NTP protocol), ensuring that the data of all sensors can be accurately aligned in time. Since multi-modal sensors are usually deployed at different locations and angles, each sensor's collected data has its own independent spatial coordinate system. In order to unify the data from different sensors for processing, a spatial coordinate system conversion relationship is established for each sensor. In this process, the position, direction and angle of view of each sensor are calibrated, and the relative position relationship and installation angle of the sensor are calculated to obtain a unified conversion formula or matrix. Then, through this formula or matrix, the local coordinate system of each sensor is converted to a global coordinate system (e.g., world coordinate system), ensuring that the spatial data from different sensors can be compared and processed in the same coordinate frame. After obtaining the unified time reference and establishing the spatial coordinate system conversion relationship, according to the unified time reference, data collected from multiple sensors at the same time is selected, and based on the previously established spatial coordinate system conversion relationship, the data of different sensors is spatially registered, so that the vehicle image, radar point cloud and infrared thermal imaging are converted to the same global coordinate system, ensuring that their spatial positions are consistent. Once the spatial registration is complete, the data from different sensors at the same time is combined to form a multi-modal data frame, each data frame containing synchronization information from different sensors. Finally, the multi-modal data frames at each time are arranged in chronological order to generate a multi-modal data frame sequence, which is composed of multiple multi-modal data frames sorted by timestamp, reflecting the complete record of data collected from different sensors at each time within a period of time, serving as the basis for subsequent data processing and analysis.
[0025] The multi-modal data frame sequence is fused and processed using a deep learning model to generate a holographic feature vector of the target vehicle.
[0026] Specifically, after obtaining the multi-modal data frame sequence, the multi-modal data frame sequence is transmitted to the feature extraction network and point cloud processing network based on the deep learning neural network as input data for feature extraction of different types of data. After completing the feature extraction, the extracted features are spliced according to the preset vector template, thereby fusing these features into a unified feature representation to form a holographic feature vector. This holographic feature vector is a high-dimensional digital representation that contains all key information of the target vehicle, such as shape, position, motion trajectory, thermal signal, etc. These features can fully describe the comprehensive features of the target vehicle for subsequent identity recognition, behavior analysis and anomaly detection.
[0027] In some embodiments, the multi-modal data frame sequence is fused by using a deep learning model to generate a holographic feature vector of the target vehicle, comprising:
[0028] A pre-trained feature extraction network is used to extract visual features and thermal radiation features from vehicle images and infrared thermal images in the multi-modal data frame sequence respectively; point cloud processing is performed on radar point clouds in the multi-modal data frame sequence to extract three-dimensional structure features and motion features; the visual features, thermal radiation features, three-dimensional structure features and motion features are fused and spliced to output the holographic feature vector of the target vehicle.
[0029] Specifically, after obtaining the sequence of multi-modal data frames, the vehicle image and infrared thermal image are extracted from the sequence of multi-modal data frames, and the vehicle image and infrared thermal image of each multi-modal data frame are input into a pre-trained feature extraction network, which includes an image feature extraction branch and a thermal image feature extraction branch. Both branches are based on a convolutional neural network (CNN). The training data includes historical vehicle images, historical visual features, historical infrared thermal images, and historical thermal radiation features. The training steps include forward propagation, loss calculation, back propagation, and parameter optimization. After the feature extraction network receives the vehicle image and infrared thermal image, the vehicle image and infrared thermal image are assigned to the corresponding branch. Each branch processes the received data using learned knowledge, automatically identifies and extracts visual features and thermal radiation features from the image. Visual features include the shape, color, and outline of the vehicle, which can help identify and distinguish different vehicles. Thermal radiation features include the heat source distribution of the vehicle and the thermal radiation of the vehicle body, which are valuable for vehicle identification in low-visibility environments, especially at night or in bad weather. Subsequently, the radar point cloud in the sequence of multi-modal data frames is spatially segmented using Euclidean clustering, and adjacent points are divided into the same cluster, thereby dividing the point cloud into multiple point cloud clusters, each corresponding to an object or part of an object. For each point cloud cluster, its geometric features such as shape, volume, surface normal, and curvature are calculated to describe the three-dimensional structure of the object, such as the outline of the vehicle and the geometric shape of the vehicle body. The bounding box and minimum enclosing sphere of the point cloud cluster are calculated to further extract the three-dimensional spatial structure of the object, forming three-dimensional structure features including the shape, volume, and spatial relationship with surrounding objects. These structure features help understand the layout of the vehicle in space. In addition, the radar point cloud is input into a motion analysis model based on a 3D convolutional neural network (3D-CNN) to extract dynamic features such as the vehicle's motion trajectory, speed, and direction. These motion features are crucial for determining the vehicle's behavior pattern. After completing the above feature extraction, the visual features extracted from the vehicle image, the thermal radiation features extracted from the infrared thermal image, the three-dimensional structure features extracted from the radar point cloud, and the motion features are fused and spliced. In this process, these features are sequentially spliced according to a pre-set feature template to form a holographic feature vector. This vector is a high-dimensional digital representation that comprehensively describes all important information about the target vehicle, effectively supporting subsequent target identity recognition, behavior analysis, and anomaly detection.
[0030] Based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle are performed, and the recognition result is output, and an abnormal event warning is given.
[0031] Specifically, after obtaining the holographic feature vector, the holographic feature vector will be input into a target classification model to identify the type of target object (such as different models of vehicles). Subsequently, based on the identified target type, the visual features and three-dimensional structural features in the holographic feature vector are combined to identify the vehicle's license plate information. After completing identity recognition, the motion state and cockpit thermal distribution are identified based on the motion features and thermal radiation characteristics in the holographic feature vector to determine the driving state. Finally, the license plate information and driving state are summarized to form the recognition result. If the recognition result shows abnormal behaviors such as speeding, driving in the wrong direction, illegal lane changes, fatigue driving, etc., or if the vehicle is monitored to stay in an illegal area for too long, an abnormal event warning will be triggered. At this time, relevant personnel will be notified in a timely manner through sound and light alarms, text messages or emails. These warning information can help traffic management personnel respond quickly and take necessary measures to ensure traffic safety.
[0032] In some embodiments, based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle are performed, and recognition results are output, including:
[0033] The holographic feature vector is input into a target classification model to identify the target type; based on the target type, the license plate area of the target vehicle is located and the characters are recognized in combination with the holographic feature vector to obtain the license plate information; the traffic behavior recognition model is activated to identify the motion state and cockpit thermal distribution of the target vehicle based on the holographic feature vector to determine the driving state, which includes the vehicle behavior and the driver state; and the license plate information and the driving state are combined to output the recognition result.
[0034] Specifically, first, the holographic feature vector is input into the target classification model, which analyzes multi-dimensional data such as images, radars, infrareds, and three-dimensional structures to identify the type of target (such as cars, motorcycles, trucks, etc.), and the target classification model is constructed in the same way as described above. Subsequently, according to the type identification result of the target vehicle, the general license plate position of the vehicle model will be searched from the pre-established license plate position mapping library, and the license plate position of each vehicle model will be different according to its visual features and body structure, so the target type information will be combined to accurately match the license plate area of the vehicle model. Once the license plate position is located, the visual features and three-dimensional structure features are further analyzed, and by analyzing the details of the license plate area, the specific information of the license plate (such as license plate number, characters, and letters) can be identified. This process is usually completed through license plate recognition technology (such as OCR), so as to obtain the license plate information. Then, the traffic behavior recognition model is activated, and the motion features (such as the speed, acceleration, and motion trajectory of the vehicle) and thermal radiation features (such as the temperature distribution of the vehicle's cabin) in the holographic feature vector are input into the traffic behavior recognition model to analyze the motion state and driver state of the target vehicle. The traffic behavior recognition model will judge the driving speed, acceleration, driving direction, etc. in the motion state according to the learned mapping relationship, and identify whether there are behaviors such as sudden acceleration, sudden braking, or irregular driving. At the same time, the temperature distribution in the cabin is analyzed using thermal radiation features to identify the state of the driver, for example, if the driver maintains a low temperature for a long time or has irregular temperature changes, it can be inferred that the driver is fatigued or engaged in improper behavior such as using a mobile phone. Once the license plate information and driving state (including vehicle behavior and driver state) of the vehicle are identified, these information will be stored in a set to form a complete identification result, which will be used for subsequent abnormal event warning to help management personnel make timely decisions, thereby improving the intelligent level of the traffic monitoring system and ensuring traffic safety.
[0035] In some embodiments, the long short-term memory network is supervised learning using historical identification records, and a traffic behavior recognition model is constructed, wherein the historical identification records at least include historical motion features, historical thermal radiation features, and historical driving states.
[0036] Specifically, historical identification records including historical motion features, historical thermal radiation features and historical driving states are obtained by collecting and organizing historical data. Then, a long short-term memory network (LSTM) is supervised to learn by using the data. The LSTM model continuously learns the data through steps such as forward propagation, loss calculation, back propagation and parameter optimization until the maximum iteration number or loss convergence is reached. After that, the LSTM model at the end of training is verified using verification data. If the accuracy of the model meets the preset accuracy, the current LSTM model is output as a traffic behavior identification model. Otherwise, the learning rate, training batch number and other hyperparameters are adjusted to further improve the identification effect of the LSTM model.
[0037] In summary, the traffic portal holographic identification method based on multi-modal perception provided by the present application has the following technical effects:
[0038] By deploying a multi-modal sensor array at the target traffic portal, raw perception data sets are collected in real time, including vehicle images, radar point clouds and infrared thermal images. Based on a unified time reference and spatial coordinates, the raw perception data sets are processed in time and space synchronization to generate a multi-modal data frame sequence. A deep learning model is used to fuse and process the multi-modal data frame sequence to generate a holographic feature vector of the target vehicle. Based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle are performed, and the identification result is output, and an abnormal event warning is given, thereby achieving the technical effects of time-space synchronization of multi-modal data and deep feature fusion, improving traffic management efficiency and safety.
[0039] Embodiment two, as Figure 2 is a structural schematic diagram of the traffic portal holographic identification system based on multi-modal perception of the present application. For example, Figure 1 The flowchart of the traffic portal holographic identification method based on multi-modal perception of the present application can be implemented by the structure as Figure 2 shown.
[0040] Based on the same idea as the traffic portal holographic identification method based on multi-modal perception in the embodiments, the traffic portal holographic identification system based on multi-modal perception provided by the present application comprises:
[0041] The data acquisition module 11 acquires a raw perception data set in real time through a multi-modal sensor array deployed at a target traffic checkpoint, the raw perception data set including vehicle images, radar point clouds, and infrared thermal images; the data synchronization processing module 12 performs spatio-temporal synchronization processing on the raw perception data set based on a unified time reference and a spatial coordinate system, to generate a multi-modal data frame sequence; the data fusion module 13 performs fusion processing on the multi-modal data frame sequence using a deep learning model, to generate a holographic feature vector of a target vehicle; and the anomaly warning module 14 performs target identity recognition and behavior analysis on the target vehicle based on the holographic feature vector, outputs a recognition result, and performs anomaly event warning.
[0042] In some embodiments, the data acquisition module 11 includes:
[0043] The multi-modal sensor array is configured to include a high-definition camera, a millimeter wave radar, and an infrared thermal imager, and the sampling frequency and data interface protocol of each sensor in the multi-modal sensor array are defined; the raw data collected by the multi-modal sensor array is preprocessed to determine the raw perception data set.
[0044] In some embodiments, the data synchronization processing module 12 includes:
[0045] A unified time reference configured for the multi-modal sensor array is obtained; a spatial coordinate system conversion relationship of the multi-modal sensor array is established; the raw perception data sets collected at the same time are spatially registered based on the spatial coordinate system conversion relationship according to the unified time reference; and the registered raw perception data sets are combined into multi-modal data frames, which are arranged in time sequence to generate the multi-modal data frame sequence.
[0046] In some embodiments, the data fusion module 13 further includes:
[0047] A pre-trained feature extraction network is used to extract visual features and thermal radiation features from the vehicle images and infrared thermal images in the multi-modal data frame sequence, respectively; point cloud processing is performed on the radar point clouds in the multi-modal data frame sequence to extract three-dimensional structure features and motion features; and the visual features, thermal radiation features, three-dimensional structure features, and motion features are fused and spliced to output the holographic feature vector of the target vehicle.
[0048] In some embodiments, the anomaly warning module 14 includes:
[0049] The holographic feature vector is input into a target classification model to identify a target type; according to the target type, a license plate region of a target vehicle is positioned and characters are recognized in combination with the holographic feature vector to obtain license plate information; a traffic behavior recognition model is activated to identify a motion state and a cockpit heat distribution of the target vehicle based on the holographic feature vector to determine a driving state, the driving state including vehicle behavior and a driver state; and the license plate information and the driving state are comprehensively used to output an identification result.
[0050] In some embodiments, the anomaly warning module 14 includes:
[0051] The long short-term memory network is supervised learning using historical identification records, the historical identification records including at least historical motion features, historical heat radiation features and historical driving states.
[0052] In embodiment three, the application further provides a computer readable storage medium, which can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the traffic checkpoint holographic identification method based on multi-modal perception in the embodiments of the application, so as to realize the traffic checkpoint holographic identification method based on multi-modal perception.
[0053] It should be understood that the disclosed embodiments and the above description enable those skilled in the art to implement the application. Meanwhile, the application is not limited to the above-mentioned part of the embodiments, and it should be understood that those skilled in the art can still modify the technical solutions recorded in the above-mentioned embodiments or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the application, and should be included in the protection scope of the application.
Claims
1. A traffic checkpoint holographic recognition method based on multimodal perception is characterized by: Methods include: A multimodal sensor array deployed at the target traffic checkpoint collects raw perception data sets in real time, including vehicle images, radar point clouds, and infrared thermal images. Based on a unified time reference and spatial coordinates, the original perception data set is subjected to spatiotemporal synchronization processing to generate a multimodal data frame sequence; Using a deep learning model to fuse the multimodal data frame sequence to generate a holographic feature vector of the target vehicle; Based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle are performed, the recognition result is output, and an abnormal event warning is issued.
2. The traffic checkpoint holographic recognition method based on multimodal perception according to claim 1 is characterized in that: Through the multimodal sensor array deployed at the target traffic checkpoint, the original perception data set is collected in real time, including: Configuring the multimodal sensor array comprising a high-definition camera, a millimeter-wave radar, and an infrared thermal imager, and defining the sampling frequency and data interface protocol of each sensor in the multimodal sensor array; The original data collected by the multimodal sensor array is preprocessed to determine the original perception data set.
3. The traffic checkpoint holographic recognition method based on multimodal perception according to claim 1 is characterized in that: Based on a unified time reference and spatial coordinates, the original perception data set is subjected to spatiotemporal synchronization processing to generate a multimodal data frame sequence, including: Obtaining a unified time reference configured by the multimodal sensor array as a unified time reference; Establishing a spatial coordinate system transformation relationship of the multimodal sensor array; According to the unified time reference, the original perception data sets collected at the same time are spatially registered based on the spatial coordinate system transformation relationship; The registered original perception data sets are combined into multimodal data frames, and arranged in time series to generate the multimodal data frame sequence.
4. The traffic checkpoint holographic recognition method based on multimodal perception according to claim 1 is characterized in that: The multimodal data frame sequence is fused using a deep learning model to generate a holographic feature vector of the target vehicle, including: Using a pre-trained feature extraction network, extracting visual features and thermal radiation features from the vehicle image and infrared thermal image in the multimodal data frame sequence respectively; performing point cloud processing on the radar point cloud in the multimodal data frame sequence to extract three-dimensional structural features and motion features; The visual features, thermal radiation features, three-dimensional structure features and motion features are fused and spliced to output a holographic feature vector of the target vehicle.
5. The traffic checkpoint holographic recognition method based on multimodal perception according to claim 1 is characterized in that: Based on the holographic feature vector, target identity recognition and behavior analysis of the target vehicle are performed, and recognition results are output, including: Inputting the holographic feature vector into a target classification model to identify the target type; According to the target type, the license plate area of the target vehicle is located and the characters are recognized in combination with the holographic feature vector to obtain the license plate information; activating a traffic behavior recognition model, identifying the motion state and cockpit thermal distribution of the target vehicle based on the holographic feature vector, and determining a driving state, wherein the driving state includes vehicle behavior and driver state; The license plate information and the driving status are combined to output a recognition result.
6. The traffic checkpoint holographic recognition method based on multimodal perception according to claim 5 is characterized in that: A traffic behavior recognition model is constructed by using historical recognition records to perform supervised learning on a long short-term memory network, wherein the historical recognition records include at least historical motion characteristics, historical thermal radiation characteristics, and historical driving status.
7. The traffic checkpoint holographic recognition system based on multimodal perception is characterized by: A system for implementing the traffic checkpoint holographic recognition method based on multimodal perception according to any one of claims 1 to 6 includes: Data acquisition module: This module uses a multimodal sensor array deployed at target traffic checkpoints to collect raw perception data sets in real time. The raw perception data sets include vehicle images, radar point clouds, and infrared thermal images. Data synchronization processing module: Based on a unified time reference and spatial coordinates, the original perception data set is subjected to spatiotemporal synchronization processing to generate a multimodal data frame sequence; Data fusion module: uses a deep learning model to fuse the multimodal data frame sequence to generate a holographic feature vector of the target vehicle; Abnormal warning module: Based on the holographic feature vector, it performs target identity recognition and behavior analysis of the target vehicle, outputs the recognition result, and issues abnormal event warning.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the traffic checkpoint holographic recognition method based on multimodal perception as described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
License plate recognition timing system based on infrared triggering
CN121708759A
Infrared trigger-based license plate recognition timing system
CN121708759B
Unmanned aerial vehicle decoy target classification method and system
CN121935790A