Super-size vehicle identification and monitoring system based on multi-modal large model and implementation method of super-size vehicle identification and monitoring system
The multi-modal large model system integrates cameras, lidar, and radar for precise oversized vehicle detection and tracking, addressing environmental challenges and enhancing traffic management efficiency.
Patent Information
- Application Number
- CN202510787983.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing vehicle size detection methods and road patrol methods have detection accuracy that is greatly affected by ambient light, occlusion, and weather conditions, and lacks the ability to deeply integrate cross-modal features, resulting in low recognition accuracy of over-size vehicles and difficult to adapt to complex road environments.
The multimodal large model is used to combine high-definition cameras, lidar and millimeter wave radar. Through multimodal data acquisition, fusion and feature extraction, accurate identification and monitoring of super-sized vehicles are achieved, and information fusion is introduced by introducing a cross-modal attention mechanism, and the detection accuracy is optimized through parameter fine-tuning of the multimodal large model.
It improves the accuracy and robustness of vehicle detection, realizes real-time monitoring and early warning of oversized vehicles, reduces the work burden of traffic management personnel, and improves the intelligence level of traffic management.
Smart Images

Figure CN120314929A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of wireless communication and intelligent transportation, and specifically relates to a super-sized vehicle recognition and monitoring system based on a multimodal large model and its implementation method. Background Art
[0002] With the rapid development of intelligent transportation systems, vehicle size detection and road inspection play crucial roles in traffic management, road maintenance, and urban planning. Traditional vehicle size detection methods mainly rely on inductive loop detectors, lidar, or single vision sensors. These methods often suffer from the problem that the detection accuracy is greatly affected by factors such as environmental illumination, occlusion, and weather conditions, and it is difficult to meet the requirements of modern traffic supervision and intelligent highways. In addition, existing road inspection methods usually adopt manual inspections or automated inspections based on single sensors, which have limitations such as low efficiency, limited coverage, and insufficient data fusion capabilities.
[0003] In recent years, computer vision and deep learning technologies have been widely applied in the field of traffic monitoring. However, most current intelligent detection systems still rely on single-modal data (such as camera images, lidar point cloud data, or millimeter-wave radar point cloud data), resulting in limited detection accuracy and robustness in complex road environments. For example, vision-based systems are affected by illumination and occlusion, while lidar data is prone to noise in adverse weather conditions. In addition, existing systems usually perform classification or regression based on independent sensor data, lacking a deep fusion mechanism for cross-modal features, resulting in low recognition accuracy for super-sized vehicles and difficulty in adapting to different environmental conditions.
[0004] To address the above problems, multimodal large models have become a new direction for solving intelligent monitoring problems due to their powerful cross-modal feature learning and knowledge transfer capabilities. However, existing research mainly focuses on general object detection tasks and has not been optimized for the fine measurement of super-sized vehicles. How to use multi-sensor data to synergistically optimize multimodal large models to improve the accuracy and stability of the monitoring system is an urgent problem in the current technology. Summary of the Invention
[0005] The purpose of the present invention is to provide a super-sized vehicle recognition and monitoring system based on a multimodal large model and its implementation method to achieve high-precision size measurement of various vehicles on the road, and combine the road inspection task to perform intelligent monitoring and early warning of super-sized vehicles. The system can adapt to complex traffic environments, ensure the accuracy and reliability of data collection, and provide technical support for intelligent traffic management and road safety supervision. To achieve the above purpose, the technical solution adopted by the present invention is as follows: A super-sized vehicle identification and monitoring system based on a multi-modal large model, including a multi-modal sensor data acquisition module, a multi-modal large model processing module, a key vehicle monitoring and tracking module, a super-sized vehicle real-time report generation module, and a user management terminal; The multi-modal data acquisition module collects multi-dimensional and multi-modal vehicle data through a variety of sensors; the collected vehicle data is transmitted to the multi-modal large model processing module; The processing process of the multi-modal large model processing module is divided into a training stage and an inference stage. In the training stage, the parameters of the multi-modal large model processing module are fine-tuned to make the multi-modal large model processing module adapt to the super-sized vehicle identification and monitoring task; in the inference stage, the data collected by the multi-modal data acquisition module is fused and feature-extracted; the multi-modal large model processing module determines super-sized vehicles for the fused and feature-extracted data and generates sensor scheduling information, which is transmitted to the key vehicle monitoring and tracking module; based on the scheduling information, the key vehicle monitoring and tracking module controls the multi-modal sensor data acquisition module to track the trajectory of the super-sized vehicle, and at the same time generates a super-sized vehicle real-time report and submits it to the user management terminal interface.
[0006] Further, the variety of sensors used by the multi-modal data acquisition module include cameras, lidar, and millimeter-wave radars, and the multi-dimensional and multi-modal vehicle data collected includes camera images, lidar point cloud data, and millimeter-wave radar point cloud data.
[0007] Further, the collected vehicle data is transmitted using a parallel wireless and wired transmission method to ensure real-time data transmission to the multi-modal large model processing module.
[0008] The present invention also discloses an implementation method for a super-sized vehicle identification and monitoring system based on a multi-modal large model, including the following steps: Step 1, deploy the multi-modal sensor data acquisition module; Step 2, perform multi-modal sensor data acquisition; various sensors work together to obtain multi-modal data; Step 3: Fusion and processing of data from the multimodal data acquisition module; the processing flow includes two stages: training and inference; in the training stage, the system fine-tunes the parameters of the pre-trained multimodal large model processing module based on the previously collected dataset to make it adapt to the feature recognition of oversize vehicles; in the inference stage, the system inputs the real-time collected multimodal data into the optimized multimodal large model processing module, and the multimodal large model processing module performs feature extraction and fusion, analyzes multi-source data, and generates high-dimensional comprehensive feature vectors; subsequently, the multimodal large model processing module calculates the vehicle size using these feature vectors and determines whether the vehicle is oversize based on a preset threshold; finally, the system generates sensor scheduling information and transmits the determination result to the key vehicle monitoring and tracking module to ensure the accuracy and real-time nature of the recognition result; Step 4: Detection, tracking of key vehicles, and generation of real-time reports; Step 5: Visualization and intelligent management of data.
[0009] Furthermore, in Step 1, the multimodal data acquisition module is installed at highway toll stations, bridges, and tunnel entrances, and mobile multimodal sensors are deployed on inspection vehicles to achieve dynamic monitoring of oversize vehicles.
[0010] Furthermore, the multiple sensors used by the multimodal data acquisition module include cameras, lidars, and millimeter-wave radars. The multi-dimensional and multimodal vehicle data collected includes camera images, lidar point cloud data, and millimeter-wave radar point cloud data. Step 3 specifically includes the following steps: Camera image: represented as , where , , are the height, width, and number of channels of the image respectively; Radar point cloud: represented as , where represents the th point cloud data, the feature contains three-dimensional coordinates and intensity information, is the number of points in the point cloud; The joint input information is represented as: ; where and are the image and radar point cloud encoders respectively, is the size embedding matrix, is the finally fused joint input information 's feature dimension; The multimodal large model processing module introduces a cross-modal attention mechanism for information fusion: ; Among them, is the weight matrix for cross-modal information fusion, , are the query, key, and value matrices, respectively, which are feature representations from different modalities, , and are the weights of the query, key, and value matrices, respectively, is the transpose symbol, is the activation function for normalization, is the scaling factor for the feature dimension, and the weight matrix ; During the fine-tuning process, the goal is to minimize the loss function: ; Among them, and are the weights of the loss function, indicating the importance of the two types of loss functions for the fine-tuning goal. The classification loss uses the cross-entropy loss to optimize the recognition accuracy of over-sized vehicles: ; Among them, is the predicted vehicle size result, is the accurate vehicle size result; The localization loss uses the L1 smooth loss to optimize the detection box position: ; Among them, is the detection box position output by the multi-modal large model processing module, is the accurate target position manually annotated, is the smooth L1 loss function; After completing the fine-tuning, the model parameters of the multi-modal large model processing module will be saved and lightweight processed according to the device requirements for deploying the multi-modal large model; in the inference stage, the system performs detection based on the real-time collected data. For each frame of input, the multi-modal large model processing module performs forward propagation: ; Among them, is the new multi-modal input data, and the output of the multi-modal large model processing module includes: the recognition result of over-sized vehicles, where 1 represents detecting an over-sized vehicle and 0 represents a normal vehicle; the bounding box information of the vehicle; If the vehicle size is close to the threshold, the system will dispatch more high-definition cameras for further appearance recognition to confirm whether it is an oversized vehicle; if the recognition confidence is low, the system will request additional lidar for additional scanning to improve detection accuracy; if the oversized vehicle is in the driving stage, the system will analyze its trajectory based on historical data, and dispatch monitoring equipment along the route in advance through trajectory prediction to improve monitoring coverage and accuracy.
[0011] Furthermore, step four specifically includes the following steps: when the system detects an abnormal oversized vehicle, it will automatically trigger an early warning mechanism; for oversized or illegally modified vehicles, the system will send relevant information to the traffic management platform to intercept or guide the illegal vehicles to the designated detection area; based on the sensor scheduling information, the system controls the multi-sensor data acquisition equipment of the corresponding road section to focus on tracking and monitoring such targets.
[0012] Furthermore, step five specifically includes the following steps: designing an interactive visual user management terminal; viewing vehicle detection status and oversized vehicle traffic records in real time through a PC or mobile terminal; the system supports historical data query and conducts in-depth analysis of oversized vehicle traffic patterns and trends in specific areas.
[0013] Compared with the traditional road monitoring system based on a single sensor, the oversized vehicle identification and monitoring system based on a multimodal large model and its implementation method of the present invention have significant advantages in multimodal data fusion, real-time monitoring and abnormal warning. The system improves the detection accuracy and robustness of the system by integrating high-definition cameras, laser radars, and millimeter-wave radars. The specific performance is as follows: 1. This system achieves comprehensive and accurate detection of vehicles by fusing multimodal data from high-definition cameras, lidar, and millimeter-wave radars, and fine-tuning based on pre-trained multimodal large models. Compared with traditional single-sensor systems, data fusion processing based on multimodal large models greatly improves detection accuracy, can work stably in complex environments, effectively overcomes the impact of environmental changes on single sensor data, and improves overall robustness.
[0014] 2. Compared with the traditional manual detection system, the abnormal warning and recording module of the present invention can automatically identify the abnormal situation of oversized vehicles, and generate warning information and sensor dispatch information in real time, so as to realize the monitoring and tracking of key vehicles. This function not only improves the timeliness and accuracy of warning, but also reduces the workload of traffic management personnel, allowing them to focus more on the handling and decision-making of abnormal situations.
[0015] 3. The system of the present invention adopts mature sensor technologies and data processing algorithms, has a low complexity, and is easy to maintain and expand in the later stage. With the growth of road monitoring requirements, the system can be flexibly expanded to adapt to more complex scenarios and high-density traffic environments, further improving the intelligent level of traffic management.
[0016] 4. The data visualization and management terminal equipped in the system displays real-time detection data, early warning information, and historical records through an interactive interface, enabling traffic management personnel to query data and conduct decision-making analysis more intuitively. Compared with the traditional manual recording and query system, the visualization function provided by the present invention significantly improves work efficiency and provides accurate decision-making support for traffic planning. Brief Description of the Drawings
[0017] Figure 1 is a block diagram of the architecture of a super-sized vehicle recognition and monitoring system based on a multi-modal large model of the present invention; Figure 2 is a schematic diagram of the deployment structure and data transmission of multi-modal sensors; Figure 3 is a flow chart of fine-tuning and inference of a multi-modal large model; Figure 4 is a flow chart of detection and early warning of the system; Figure 5 is a schematic diagram of the visualization interface of the management end of the system. Detailed Embodiment
[0018] The present invention will be further described in detail below with reference to the accompanying drawings.
[0019] The present invention realizes the recognition and monitoring of super-sized vehicles based on multi-sensor data (camera, lidar, millimeter-wave radar) and a multi-modal large model. The system consists of a multi-modal sensor data acquisition module, a multi-modal large model processing module, a key vehicle monitoring and tracking module, a super-sized vehicle real-time report generation module, and a user management terminal. The overall architecture of the system is as Figure 1As shown below: First, multi-modal sensor data acquisition modules are deployed at key nodes such as highway toll stations, important bridge and tunnel entrances, etc. Sensors such as cameras, lidar, and millimeter-wave radars are used to collect multi-modal information such as vehicle dimensions and road conditions. The collected data is transmitted in real-time to the multi-modal large model processing module through both wireless and wired parallel methods. The processing process of the multi-modal large model is divided into two stages: training and inference. In the training stage, the parameters of the pre-trained model are fine-tuned to adapt it to the task of over-sized vehicle recognition and monitoring. In the inference stage, based on the real-time collected data, feature extraction and fusion are performed, the vehicle type is determined, and sensor scheduling information is generated and transmitted to the key vehicle monitoring and tracking module. The key vehicle monitoring and tracking module controls the sensors to perform trajectory tracking according to the scheduling information, generates a real-time report, and finally submits it to the user management terminal to achieve the accurate recognition and dynamic monitoring of over-sized vehicles.
[0020] Figure 2 The deployment and data transmission of multi-modal sensors are shown. The key to sensor deployment lies in the comprehensive coverage of the target monitoring area. Specifically, high-definition cameras are installed at key positions to obtain the appearance data of vehicles, especially the external contours and identification information of vehicles; lidar is used to capture the three-dimensional spatial data of vehicles, providing accurate depth information, which helps to build a high-precision vehicle contour model; while millimeter-wave radars can monitor the dynamic information of vehicles in real-time. Especially in adverse weather conditions, millimeter-wave radars can provide relatively stable distance and speed data, enhancing the robustness of the system. The data of all sensors are transmitted in parallel by wireless and wired methods and integrated through a data synchronization mechanism to ensure the time consistency and spatial alignment of data from different sensors.
[0021] Figure 3 The fine-tuning and inference flow chart of the multi-modal large model is shown. The multi-modal large model in the present invention is based on the general pre-trained large model (Qwen-VL). Since the general large model lacks a specific learning process for specific scenarios, in this stage, it is actually necessary to fine-tune the pre-trained large model. After the pre-trained multi-modal large model is loaded, the vehicle multi-modal information (camera images, radar point cloud maps) is used as the training input, and over-sized vehicle data is labeled for multi-modal feature extraction and fusion. During the fine-tuning process, the task performance is mainly optimized by fine-tuning the backend adaptation layer of the model (i.e., the fully connected layer of the cross-modal fusion module). After the fine-tuning is completed, the model parameters are saved and lightweight processing is performed according to the terminal deployment requirements, so that the system can adapt to small terminals and be deployed on edge computing devices (NVIDIA Jetson AGX Xavier). In the inference stage, after receiving the real-time collected data set, the multi-modal large model performs inference, realizes the recognition of over-sized vehicles and outputs decision-making information, and focuses on monitoring over-sized vehicles.
[0022] Figure 4 It shows the monitoring and early warning process of the system. After the key vehicle monitoring and tracking module receives the input of oversize vehicle abnormal information composed of the oversize vehicle judgment generated by the multimodal large model and the sensor scheduling information, on the one hand, it visualizes the positioning information of the oversize vehicle, and on the other hand, it realizes the scheduling of multimodal sensors according to the sensor scheduling instructions, tracks the oversize vehicle and obtains the trajectory information. These information are finally used to generate a real-time report of the oversize vehicle and upload it to the management terminal. Figure 5 It shows the schematic diagram of the visualization interface of the management terminal. The terminal interface consists of the visualization positioning information of the oversize vehicle, the warning situation of the oversize vehicle and the sensor status information, etc., and tracks the trajectory of the oversize vehicle.
[0023] The following will explain each step in more detail with reference to the accompanying drawings.
[0024] Step 1, Deployment of multimodal sensor devices. In order to improve the detection accuracy and monitoring coverage of oversize vehicles, the system combines the characteristics of the highway environment and optimizes the deployment strategy of sensors. First, at fixed monitoring points, the system preferentially selects key nodes with dense traffic flow and prone to oversize vehicle violations, such as highway toll stations, bridges, tunnel entrances, etc. Multimodal sensing devices such as high-definition cameras, lidar, and millimeter-wave radars are deployed here to statically and accurately monitor the appearance, size, and contour of passing vehicles. Second, to make up for the limitations of fixed monitoring points, the system deploys mobile multimodal sensors on patrol vehicles to realize the dynamic tracking of oversize vehicles. These sensors can flexibly patrol on the main line of the highway and remote sections, focus on monitoring abnormal vehicles, effectively improve the monitoring coverage, and reduce the regulatory blind spots.
[0025] Step 2, Multimodal sensor data acquisition. Figure 2 The deployment of multimodal sensors in [description] shows how different sensors work together to obtain comprehensive multimodal data. At fixed monitoring points, high-definition cameras capture the appearance images of vehicles and combine OCR technology to achieve license plate recognition to ensure the accurate matching of vehicle identity information. Lidar generates high-precision point cloud data through three-dimensional scanning, providing millimeter-level accuracy for vehicle body size measurement to ensure the accurate determination of oversize vehicles. At the same time, millimeter-wave radar, relying on its long-distance measurement ability, can still maintain high-precision perception of vehicle size in bad weather (such as rain and fog environments), making up for the deficiencies of optical sensors in low visibility environments. To ensure the efficient transmission of data, the system adopts a parallel mode of wireless and wired to ensure the real-time nature of data while taking into account the stability of transmission.
[0026] Step 3: Fusion and processing of data based on the multi-modal large model. The processing flow of the multi-modal large model includes two stages: training and inference, to ensure that the system can efficiently adapt to the task of over-sized vehicle recognition and monitoring in the intelligent transportation scenario. In the training stage, the system fine-tunes the parameters of the pre-trained multi-modal large model based on the pre-collected data set to make it adapt to the feature recognition of over-sized vehicles and improve the detection accuracy and generalization ability of the model. The optimized model can extract key information from multi-modal data more accurately to adapt to complex road environments. In the inference stage, the system inputs the real-time collected multi-modal data into the optimized model, and the model performs feature extraction and fusion, analyzes multi-source data such as images, point clouds, and radar signals, and generates high-dimensional comprehensive feature vectors. Subsequently, the model uses these feature vectors to accurately calculate the vehicle size and determines whether the vehicle is over-sized based on a preset threshold. Finally, the system generates sensor scheduling information and transmits the determination result to the key vehicle monitoring and tracking module to ensure the accuracy and real-time nature of the recognition result. Figure 3 The fine-tuning and inference process of the multi-modal large model shown is based on the pre-trained general multi-modal large model Qwen-VL. Since the pre-trained model lacks targeted learning, it needs to be fine-tuned at this stage. After loading the pre-trained model, the system takes the multi-modal data of the vehicle (camera images, radar point clouds) and the vehicle size information as inputs, and introduces the labeled data of over-sized vehicles to extract and fuse multi-modal features.
[0027] Specifically, after loading the pre-trained model, the system takes the multi-modal data of the vehicle as inputs, including: Camera images: Represented as , where , , are the height, width, and number of channels of the image respectively; Radar point clouds: Represented as , where represents the rd point cloud data, The feature contains three-dimensional coordinates and intensity information, is the number of points in the point cloud; The combined input information is represented as: ; where and are the image and radar point cloud encoders respectively, is the size embedding matrix, is the finally fused combined input information 's feature dimension; The multi-modal large model processing module introduces a cross-modal attention mechanism for information fusion: ; Among them, is the weight matrix for cross-modal information fusion, , are the query, key, and value matrices respectively, which are the feature representations from different modalities, , and are the weights of the query, key, and value matrices respectively, is the transpose symbol, is the activation function for normalization, is the scaling factor of the feature dimension, and the weight matrix ; During the fine-tuning process, the goal is to minimize the loss function: ; Among them, and are the loss function weights, indicating the importance of the two types of loss functions for the fine-tuning goal. The classification loss adopts the cross-entropy loss to optimize the recognition accuracy of over-sized vehicles: ; Among them, is the predicted vehicle size result, is the accurate vehicle size result; The localization loss adopts the L1 smooth loss to optimize the detection box position: ; Among them, is the detection box position output by the multi-modal large model processing module, is the accurate target position manually marked, is the smooth L1 loss function; After the fine-tuning is completed, the model parameters of the multi-modal large model processing module will be saved and lightweight processed according to the device requirements for the deployment of the multi-modal large model; in the inference stage, the system performs detection based on the real-time collected data. For each frame of input, the multi-modal large model processing module performs forward propagation: ; The output of the multi-modal large model includes: (1) The recognition result of over-sized vehicles: , where 1 represents the detection of an over-sized vehicle and 0 represents a normal vehicle; (2) The bounding box information of the vehicle: , indicating the position information of the vehicle.
[0028] Finally, the system generates sensor dispatch information based on the recognition and detection results. Specifically: If the vehicle size is close to the threshold, the system can dispatch more high-definition cameras for further appearance recognition to confirm whether it is an oversized vehicle. If the recognition confidence is low, the system can request additional lidars for additional scanning to improve detection accuracy. If the oversized vehicle is in the driving stage, the system analyzes its trajectory based on historical data, and dispatches monitoring equipment along the route in advance through trajectory prediction to improve monitoring coverage and accuracy.
[0029] Step 4: Key vehicle detection and tracking and real-time report generation. Figure 4 As shown in the figure, when the system detects an oversized or illegally modified vehicle, it will automatically trigger an early warning mechanism to quickly respond to potential traffic safety risks. First, the system will upload the basic information of the illegal vehicle (including license plate number, vehicle model, size data and driving trajectory) to the traffic management platform in real time. The management platform will take corresponding countermeasures according to the degree of violation, such as remote warning, guiding the vehicle to enter the designated detection area, or notifying the traffic management department to intercept on the spot. At the same time, based on the sensor dispatch information, the system will accurately control the multimodal sensors of the corresponding road section to focus on tracking and monitoring the target vehicle. Equipment such as high-definition cameras, lidars and millimeter-wave radars will work together to continuously record the driving trajectory of the target vehicle to ensure full tracking of oversized vehicles. For vehicles with serious over-limit risks, the system can analyze their violation frequency and driving patterns in combination with historical data, further optimize monitoring strategies, and improve supervision efficiency. In this process, all detection data will be recorded and stored for subsequent statistical analysis to provide data support for traffic planning.
[0030] Step 5: Data visualization and intelligent management. Figure 5 As shown in the figure, this system integrates an interactive visual management terminal to assist managers in real-time monitoring and decision support in an intuitive and efficient way. The terminal interface adopts a graphical design, combined with multi-layer data display, so that managers can clearly grasp the vehicle detection situation, the real-time traffic records of oversized vehicles and various monitoring indicators. The system supports historical data query function, and users can analyze the traffic patterns of oversized vehicles in a specific area to provide a decision-making basis for the optimization of traffic management strategies. In addition, the system provides an open terminal management interface to support user-defined instructions and tasks. Managers can flexibly configure personalized parameters such as detection rules, alarm strategies, and monitoring ranges, so that the system can adapt to different road conditions and regulatory needs.
[0031] The present invention proposes a system for identifying and monitoring over-sized vehicles based on a multi-modal large model. This system integrates various sensors such as high-definition cameras, lidar, and millimeter-wave radars, and combines multi-modal large model technology to achieve precise identification and tracking monitoring of over-sized vehicles. The system generates real-time reports of over-sized vehicles based on the results of identification and monitoring, and adopts a parallel scheme of wireless and wired transmission to transmit them to the user management terminal interface.
[0032] In terms of the overall architecture, this system consists of the following modules: (1) Multi-modal sensor data acquisition module; (2) Multi-modal large model processing module; (3) Key vehicle monitoring and tracking module; (4) Real-time report generation module for over-sized vehicles; (5) User management terminal. Among them, the multi-modal data acquisition module is deployed at key road points such as highway toll stations, important bridges, and tunnel entrances, and collects multi-dimensional multi-modal vehicle data (camera images, lidar point cloud data, millimeter-wave radar point cloud data) through various sensors (cameras, lidar, millimeter-wave radars). The collected data is transmitted to the multi-modal large model processing module.
[0033] The processing of the multi-modal large model processing module is divided into two main stages. In the training stage, the general large model is adapted to the current over-sized vehicle identification and monitoring scenario in intelligent transportation by means of parameter fine-tuning; in the actual deployment stage, the data collected by multi-modal sensors is fused and feature-extracted. The multi-modal large model processing module determines over-sized vehicles based on multi-modal data and gives sensor scheduling information, which is transmitted to the key vehicle monitoring and tracking module. Based on the scheduling information, the key vehicle monitoring and tracking module controls the multi-modal sensor data acquisition module to track the trajectory of over-sized vehicles, and at the same time generates real-time reports of over-sized vehicles and submits them to the user management terminal interface, providing a visual interface to support traffic control and decision-making.
[0034] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the guidance of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A super-sized vehicle recognition and monitoring system based on a multi-modal large model, characterized in that It includes a multi-modal sensor data acquisition module, a multi-modal large model processing module, a key vehicle monitoring and tracking module, an oversized vehicle real-time report generation module, and a user management terminal; The multi-modal data acquisition module collects multi-dimensional and multi-modal vehicle data through multiple sensors; the collected vehicle data is transmitted to the multi-modal large model processing module; The processing process of the multi-modal large model processing module is divided into a training stage and an inference stage. In the training stage, the parameters of the multi-modal large model processing module are fine-tuned to make the multi-modal large model processing module adapt to the oversized vehicle identification and monitoring tasks; In the inference stage, the data collected by the multi-modal data acquisition module is fused and feature-extracted; the multi-modal large model processing module determines the oversized vehicle based on the fused and feature-extracted data and generates sensor scheduling information, which is transmitted to the key vehicle monitoring and tracking module; Based on the scheduling information, the key vehicle monitoring and tracking module controls the multi-modal sensor data acquisition module to track the trajectory of the oversized vehicle, and at the same time generates a real-time report of the oversized vehicle and submits it to the user management terminal interface.
2. The ultra-size vehicle recognition and monitoring system based on a multi-modal large model according to claim 1, wherein The multiple sensors used by the multi-modal data acquisition module include cameras, lidars, and millimeter-wave radars. The multi-dimensional and multi-modal vehicle data collected includes camera images, lidar point cloud data, and millimeter-wave radar point cloud data.
3. The super-sized vehicle identification and monitoring system based on a multimodal large model according to claim 1, characterized in that The collected vehicle data is transmitted using a parallel wireless and wired transmission method to ensure real-time data transmission to the multi-modal large model processing module.
4. A method for implementing an oversize vehicle recognition and monitoring system based on a multimodal large model, using an oversize vehicle recognition and monitoring system based on a multimodal large model as described in any one of claims 1-3, characterized in that, It includes the following steps: Step 1, deploy the multi-modal sensor data acquisition module; Step 2, perform multi-modal sensor data acquisition; Various sensors work together to obtain multi-modal data; Step 3, fusion and processing based on the data of the multi-modal data acquisition module; the processing process includes two stages: training and inference; In the training stage, the system fine-tunes the parameters of the pre-trained multi-modal large model processing module based on the previously collected data set to make it adapt to the feature recognition of oversized vehicles; in the inference stage, the system inputs the real-time collected multi-modal data into the optimized multi-modal large model processing module. The multi-modal large model processing module performs feature extraction and fusion, analyzes multi-source data, and generates high-dimensional comprehensive feature vectors; subsequently, the multi-modal large model processing module calculates the vehicle size using these feature vectors and determines whether the vehicle is oversized based on a preset threshold; finally, the system generates sensor scheduling information and transmits the determination result to the key vehicle monitoring and tracking module; Step 4, key vehicle detection, tracking, and real-time report generation; Step 5, visualization and intelligent management of data.
5. The implementation method of an oversize vehicle recognition and monitoring system based on a multimodal large model according to claim 4, characterized in that, In Step 1, install the multi-modal data acquisition module at highway toll stations, bridges, and tunnel entrances, and deploy mobile multi-modal sensors on inspection vehicles to achieve dynamic monitoring of oversized vehicles.
6. The implementation method of an oversized vehicle recognition and monitoring system based on a multimodal large model according to claim 4, characterized in that The multiple sensors used by the multi-modal data acquisition module include cameras, lidars, and millimeter-wave radars. The multi-dimensional and multi-modal vehicle data collected includes camera images, lidar point cloud data, and millimeter-wave radar point cloud data. Step 3 specifically includes the following steps: Camera image: Represented as , where , , are the height, width, and number of channels of the image, respectively; Radar point cloud: expressed as , where represents the th point cloud data, The features include three-dimensional coordinates and intensity information, is the number of points in the point cloud; The combined input information is represented as: ; Among them and are the image and radar point cloud encoders respectively, is the size embedding matrix, is the jointly input information after final fusion of the feature dimension; The multi-modal large model processing module introduces a cross-modal attention mechanism for information fusion: ; Among them, is the weight matrix for cross-modal information fusion, , are the query, key, and value matrices, respectively, which are derived from feature representations of different modalities, , and are the weights of the query, key, and value matrices, respectively, is the transpose symbol, is the activation function for normalization, is the scaling factor of the feature dimension, and the weight matrix ; During fine-tuning, the goal is to minimize the loss function: ; Among them, and are the loss function weights, indicating the importance of the two types of loss functions for the fine-tuning target. The classification loss adopts cross-entropy loss and is used to optimize the recognition accuracy of over-sized vehicles: ; Among them, is the predicted vehicle size result, is the accurate result of the vehicle size; Localization loss The L1 smooth loss is adopted to optimize the position of the detection box: ; Among them, is the position of the detection box output by the multi-modal large model processing module, is the accurate position of the target manually marked, is the smooth L1 loss function; After the fine-tuning is completed, the model parameters of the multi-modal large model processing module will be saved and lightweight processed according to the requirements of the multi-modal large model deployment device; in the inference stage, the system performs detection based on the real-time collected data. For each frame of input, the multi-modal large model processing module performs forward propagation: ; Among them, is the new multimodal input data, which is the output of the multimodal large model processing module including: the recognition result of oversize vehicles , where 1 represents that an oversize vehicle is detected, and 0 represents a normal vehicle; the bounding box information of the vehicle ; If the vehicle size is close to the threshold, the system will dispatch more high-definition cameras for further appearance recognition to confirm whether it is an oversized vehicle; if the recognition confidence is low, the system will request additional lidar for supplementary scanning to improve detection accuracy; if the oversized vehicle is in the driving stage, the system will analyze its trajectory based on historical data and dispatch monitoring equipment along the route in advance through trajectory prediction.
7. The implementation method of an oversized vehicle recognition and monitoring system based on a multimodal large model according to claim 4, characterized in that, Step 4 specifically includes the following steps: When the system detects an abnormal oversized vehicle, it will automatically trigger the early warning mechanism; for oversized or illegally modified vehicles, the system will send relevant information to the traffic management platform to intercept or guide the illegal vehicles to the designated detection area; based on the sensor scheduling information, the system controls the multi-sensor data acquisition equipment of the corresponding road section to focus on tracking and monitoring such targets.
8. The implementation method of an oversized vehicle recognition and monitoring system based on a multimodal large model according to claim 4, characterized in that, Step five specifically includes the following steps: designing an interactive visual user management terminal; viewing vehicle inspection status and oversized vehicle passage records in real time through a PC or mobile terminal; the system supports historical data query.
Citation Information
Patent Citations
Whole-course overspeed monitoring system for vehicles running on expressway
CN118711377A
Radar and video information fusion coding method for vehicle over-limit early warning
CN119478858A
Toll vehicle type detection method and system based on fusion of laser radar and bayonet camera
CN120014728A
Intelligent driving multi-sensor fusion data processing system
CN120105350A
Three-Dimensional Object Detection
US20220214457A1