Traffic video multi-dimensional semantic understanding method and system based on multi-modal large model
By using a multimodal large model to perform multi-dimensional semantic understanding of traffic video streams, text data, and environmental sensor data, the problem of single semantic understanding dimension and insufficient recognition accuracy in complex scenes in existing technologies is solved, thus achieving more efficient traffic management and safety early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AI SUPER EYE TECH CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-01
AI Technical Summary
Existing traffic video semantic understanding technologies suffer from problems such as limited semantic understanding dimensions and insufficient recognition accuracy in complex scenarios.
The system employs real-time acquisition of traffic video streams, traffic management text data, and environmental sensor data. After preprocessing, it performs multimodal feature extraction and cross-modal feature alignment, inputting the data into a pre-trained multimodal large model for multidimensional semantic understanding. The system is then optimized and updated in conjunction with a traffic rule knowledge base.
It enriches the semantic coverage dimensions, improves the recognition accuracy in complex scenarios, and enhances the level of intelligence in traffic management.
Smart Images

Figure CN121963033A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation, and in particular to a method and system for multi-dimensional semantic understanding of traffic videos based on a multimodal large model. Background Technology
[0002] Traffic video semantic understanding is a core component of intelligent transportation systems. Its capabilities directly determine the level of intelligence in higher-level applications such as traffic management, safety warnings, and traffic scheduling, and are crucial for improving traffic efficiency and ensuring road safety. Current mainstream technologies for addressing this problem are mainly divided into two categories: single-modal processing and traditional multimodal fusion. Single-modal processing focuses solely on visual data to identify basic targets such as vehicles and pedestrians, while traditional multimodal fusion often uses simple feature stitching to integrate multi-source information. Existing methods suffer from limitations because single-modal processing cannot associate non-visual information such as traffic rules, and traditional multimodal fusion lacks deep integration of temporal dynamic features from videos. This results in a single dimension of semantic understanding, making it prone to problems such as target misidentification and semantic judgment bias in complex scenarios like adverse weather or target occlusion.
[0003] At present, traffic video semantic understanding suffers from technical problems such as a single semantic understanding dimension and insufficient recognition accuracy in complex scenarios. Summary of the Invention
[0004] This application provides a method and system for multi-dimensional semantic understanding of traffic videos based on a multimodal large model. It employs real-time acquisition of traffic video streams, traffic management text data, and environmental sensor data, followed by preprocessing to obtain a traffic data sequence. This sequence undergoes multimodal feature extraction and cross-modal feature alignment to establish a traffic feature alignment vector. This vector is then input into a pre-trained multimodal large model to obtain preliminary multi-dimensional semantic understanding results. After multi-dimensional semantic parsing and optimization, a structured semantic report is generated. Based on the feedback dataset from this report, the multimodal large model and traffic rule knowledge base are periodically and incrementally updated and optimized. These techniques address the technical problems of existing traffic video semantic understanding methods, such as limited semantic understanding dimensions and insufficient recognition accuracy in complex scenarios. The method achieves the technical effect of enriching semantic coverage dimensions and improving recognition accuracy in complex scenarios.
[0005] This application provides a method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model, comprising: real-time acquisition of multi-source collected data, including traffic video streams, traffic management text data, and environmental sensor data; preprocessing the multi-source collected data to obtain traffic data sequences; performing multimodal feature extraction and cross-modal feature alignment on the traffic data sequences to establish traffic feature alignment vectors; inputting the traffic feature alignment vectors into a pre-trained multimodal large model to obtain preliminary multi-dimensional semantic understanding results; performing multi-dimensional semantic parsing and optimization on the preliminary semantic understanding results to obtain a structured semantic report; and periodically incrementally updating and optimizing the multimodal large model and the traffic rule knowledge base based on the feedback dataset of the structured semantic report.
[0006] In a possible implementation, the multi-source acquired data is preprocessed to obtain a traffic data sequence, and the following processing is performed: frame extraction is performed on the traffic video stream to obtain a traffic video frame set; the traffic video frame set is enhanced using an adaptive histogram equalization algorithm and a median filtering algorithm to obtain an enhanced video frame set; moving target detection is performed on the enhanced video frame set based on the inter-frame difference method to obtain a video keyframe sequence; data cleaning is performed on the traffic management text data and the environmental sensing data to obtain a traffic management text sequence and an environmental sensing sequence; the traffic data sequence is generated based on the video keyframe sequence, the traffic management text sequence, and the environmental sensing sequence.
[0007] In a possible implementation, multimodal feature extraction and cross-modal feature alignment are performed on the traffic data sequence to establish a traffic feature alignment vector. The following processing is then performed: visual feature extraction is performed on the video keyframe sequence to obtain a visual feature vector; text feature extraction is performed on the traffic management text sequence to obtain a text semantic vector; environmental feature extraction is performed on the environmental sensing sequence to obtain an environmental feature vector; the visual feature vector, the text semantic vector, and the environmental feature vector are dimension-unified, and the traffic feature alignment vector is generated by combining the timestamp and location label.
[0008] In a possible implementation, the following processing is performed: the traffic feature alignment vector is subjected to feature deep fusion and generative inference through the cross-modal attention mechanism of the multimodal large model, and the preliminary semantic understanding results including target dimension, behavior dimension, scene dimension and event dimension are output.
[0009] In a possible implementation, the preliminary semantic understanding result is subjected to multi-dimensional semantic parsing and optimization to obtain a structured semantic report, and the following processing is performed: hierarchical semantic parsing is performed on the preliminary semantic understanding result to obtain multi-dimensional traffic semantic parsing result; the traffic rule knowledge base is called to verify and optimize the multi-dimensional traffic semantic parsing result to obtain initial traffic semantic optimization result; the initial traffic semantic optimization result is corrected in context using the target trajectory feature sequence and the semantic association of continuous frames to generate the structured semantic report.
[0010] In a possible implementation, the following processing is performed: the hierarchical semantic parsing includes target attribute parsing, behavior action parsing, scene type parsing, and event level parsing.
[0011] In a possible implementation, the multimodal large model and traffic rule knowledge base are periodically and incrementally updated and optimized based on the feedback dataset of the structured semantic report, and the following processes are performed: collecting the manual review results of the structured semantic report through the traffic management platform; annotating the structured semantic report according to the manual review results to obtain semantic review tags; collecting upper-layer application feedback data of the structured semantic report; and organizing the semantic review tags and the upper-layer application feedback data to obtain the feedback dataset.
[0012] This application also provides a traffic video multi-dimensional semantic understanding system based on a multimodal large model, comprising: a multi-source data acquisition module for real-time acquisition of multi-source data, including traffic video streams, traffic management text data, and environmental sensor data; a preprocessing module for preprocessing the multi-source data to obtain traffic data sequences; a cross-modal feature alignment module for extracting multimodal features and aligning cross-modal features of the traffic data sequences to establish traffic feature alignment vectors; a preliminary semantic understanding module for inputting the traffic feature alignment vectors into a pre-trained multimodal large model to obtain multi-dimensional preliminary semantic understanding results; a multi-dimensional semantic parsing module for performing multi-dimensional semantic parsing and optimization on the preliminary semantic understanding results to obtain a structured semantic report; and an incremental update module for periodically incrementally updating and optimizing the multimodal large model and the traffic rule knowledge base based on the feedback dataset of the structured semantic report.
[0013] The proposed method and system for multi-dimensional semantic understanding of traffic videos based on a multimodal large model, as described in this application, first acquires multi-source data in real time, including traffic video streams, traffic management text data, and environmental sensor data. Next, the multi-source data is preprocessed to obtain traffic data sequences. Then, multimodal feature extraction and cross-modal feature alignment are performed on the traffic data sequences to establish traffic feature alignment vectors. These alignment vectors are then input into a pre-trained multimodal large model to obtain preliminary multi-dimensional semantic understanding results. These preliminary semantic understanding results are then subjected to multi-dimensional semantic parsing and optimization to obtain a structured semantic report. Finally, the multimodal large model and the traffic rule knowledge base are periodically and incrementally updated and optimized based on the feedback dataset of the structured semantic report. Through this process, the proposed method and system achieve the technical effects of enriching semantic coverage dimensions and improving the recognition accuracy of complex scenes. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0015] Figure 1 This is a flowchart illustrating the multi-dimensional semantic understanding method for traffic videos based on a multimodal large model, as provided in an embodiment of this application.
[0016] Figure 2 This is a schematic diagram of the structure of a traffic video multi-dimensional semantic understanding system based on a multimodal large model, provided in an embodiment of this application.
[0017] Figure labeling: 10 for multi-source data acquisition module, 20 for preprocessing module, 30 for cross-modal feature alignment module, 40 for preliminary semantic understanding module, 50 for multi-dimensional semantic parsing module, and 60 for incremental update module. Detailed Implementation
[0018] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0019] This application provides a method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model, such as... Figure 1 As shown, the method includes:
[0020] Step S100: Acquire multi-source data in real time, including traffic video streams, traffic management text data, and environmental sensor data.
[0021] Specifically, corresponding data acquisition devices are deployed in key traffic areas to acquire three types of data through interfaces or direct acquisition methods. Among them, traffic video streams are acquired in real time by deploying high-definition network cameras with a frame rate of 30 frames per second and a resolution of no less than 4K at key locations at traffic intersections and road sections, covering visual targets such as vehicles, pedestrians, non-motorized vehicles, traffic lights, and traffic signs. Traffic management text data is acquired through interfaces with traffic management databases, in the form of structured text, including traffic regulations (such as speed limits and traffic priorities), intersection attribute information (such as the number of lanes and prohibited times), and historical event records (such as accident types and handling results). Environmental sensor data is acquired through environmental sensors deployed around the road sections, including real-time weather (sunny, rainy, snowy, foggy), light intensity (with a standard of no less than 100 lux as normal light intensity), and road surface conditions (dry, wet, icy), etc. The acquisition frequency is synchronized with the video stream to ensure data spatiotemporal consistency.
[0022] Step S200: Preprocess the multi-source collected data to obtain a traffic data sequence.
[0023] Specifically, preprocessing operations are performed according to data type. For video data, the focus is on frame extraction, enhancement, and keyframe selection; for text and environmental data, the focus is on cleaning and denoising. Finally, the data is integrated based on spatiotemporal correlation. Specifically, for traffic video streams, keyframe sequences are obtained through frame extraction, enhancement, and moving object detection; for traffic management text data, data cleaning is performed by removing redundant characters and correcting formatting errors; for environmental sensor data, cleaning is performed by removing outliers and filling in missing data. The three types of processed data are then linked and integrated according to timestamps and location information to generate a traffic data sequence.
[0024] In one possible implementation, the multi-source acquired data is preprocessed to obtain a traffic data sequence. Step S200 further includes step S210, which involves extracting frames from the traffic video stream to obtain a traffic video frame set. Specifically, the OpenCV video processing library is used, and its video reading interface is called to read the traffic video stream. The frame extraction interval is set to extract one keyframe per frame, meaning that every consecutive frame in the video stream is extracted and stored as a JPEG image frame. All extracted image frames are combined to form a traffic video frame set.
[0025] Step S220 involves enhancing the traffic video frame set using an adaptive histogram equalization algorithm and a median filtering algorithm to obtain an enhanced video frame set. Specifically, for each image in the traffic video frame set, an adaptive histogram equalization algorithm is first used to divide the image into multiple sub-blocks. Histogram equalization is then performed on each sub-block separately to avoid detail loss caused by overall equalization and to optimize image clarity in areas of uneven lighting. Next, a median filtering algorithm is used. The filter window size is set, and for each pixel in the image, the median value of its neighboring pixels is used to replace the original value of that pixel, removing image noise from rainy or snowy weather conditions and other random noise. All processed image frames constitute the enhanced video frame set.
[0026] Step S230: Motion target detection is performed on the enhanced video frame set based on the inter-frame difference method to obtain a video keyframe sequence. Specifically, for consecutive frames in the enhanced video frame set, the current frame and the previous frame are first converted to grayscale images. The pixel difference between the two grayscale images is calculated, and a pixel difference threshold is set. When the pixel difference between the two frames is greater than the threshold, it is determined that a moving target exists, and the current frame is retained as a keyframe. If the pixel difference is less than or equal to the threshold, it is determined that there is no valid moving target, and the frame is discarded to reduce the amount of invalid data processing.
[0027] For example, in the enhanced video frame set, the pixel difference between the first and second frames is 15, which is less than the threshold of 30, so the second frame is discarded; the pixel difference between the first and third frames is 40, which is greater than the threshold of 30, so the third frame is retained as a key frame, and so on, until all frames containing moving targets are finally selected and arranged in chronological order to form a video key frame sequence.
[0028] Step S240 involves data cleaning of the traffic management text data and the environmental sensing data to obtain a traffic management text sequence and an environmental sensing sequence. Specifically, for the traffic management text data, Python string processing functions are used to remove redundant spaces, line breaks, and special characters, correct typos, and categorize the data by data type to form a structured traffic management text sequence. For the environmental sensing data, the interquartile range (IQR) method is used to remove outliers. This involves calculating the first and third quartiles of the data, determining outlier boundaries (data values less than the first quartile minus 1.5 IQR or greater than the third quartile plus 1.5 IQR) and deleting them. For missing data, the average of adjacent data is used to fill in the missing data. After processing, the data is arranged in chronological order to form the environmental sensing sequence.
[0029] Step S250: Generate the traffic data sequence based on the video keyframe sequence, the traffic management text sequence, and the environmental sensor sequence. Specifically, add a unified format timestamp and location tag to the data in each type of sequence. Through database association query, associate the video keyframes, traffic management text data, and environmental sensor data corresponding to the same timestamp and location tag. Store the associated data in JSON format. All associated data are arranged in timestamp order to form the traffic data sequence.
[0030] Step S300: Perform multimodal feature extraction and cross-modal feature alignment on the traffic data sequence to establish a traffic feature alignment vector.
[0031] Specifically, a Transformer-based visual encoder is used to extract visual feature vectors from video keyframes, a text encoder is used to extract semantic vectors from traffic management text, and a feature encoding method is used to extract environmental feature vectors from environmental sensor data. All three types of vectors are mapped to 1024 dimensions to ensure dimensionality uniformity. Finally, based on timestamps and location labels, the three types of feature vectors from the same spatiotemporal scene are associated and aligned to form a traffic feature alignment vector.
[0032] In one possible implementation, multimodal feature extraction and cross-modal feature alignment are performed on the traffic data sequence to establish a traffic feature alignment vector. Step S300 further includes step S310, which extracts visual features from the video keyframe sequence to obtain a visual feature vector. Specifically, a ViT-L / 14 visual encoder is used to scale each keyframe in the video keyframe sequence to 224×224 pixels and input it into the encoder. The encoder's multi-layer Transformer structure extracts features from the image and outputs a 1024-dimensional visual feature vector. This vector contains information such as target location (boundary box coordinates, with the top left corner of the image as the origin and the coordinate value being the number of pixels), morphological features (encoded information such as vehicle type and pedestrian clothing), and dynamic features (quantized data of vehicle speed and pedestrian movement direction). Simultaneously, the DeepSORT target tracking algorithm is used to associate the same target in consecutive frames based on the appearance features and motion information of the target in the keyframes, generating a target trajectory feature sequence of 30 frames in length, which is combined with the single-frame visual feature vector to form a complete visual feature representation.
[0033] Step S320 involves extracting text features from the traffic management text sequence to obtain a text semantic vector. Specifically, a BERT-based pre-trained text encoder is used to first segment each text data in the traffic management text sequence, breaking the text into word units. For example, "speed limit 50 km / h" is segmented into "speed limit," "50," "km," and "hour." Then, the segmented word units are converted into word embedding vectors that the model can recognize and input into the BERT model. The model's multi-layer Transformer encoder performs semantic encoding, ultimately outputting a 1024-dimensional text semantic vector. This vector contains the semantic information of the text, such as the constraints of regulations and key parameters of intersection attributes.
[0034] Step S330: Extract environmental features from the environmental sensing sequence to obtain an environmental feature vector. Specifically, for discrete data in the environmental sensing sequence, such as weather type and road surface condition, one-hot encoding is used. For example, in weather type, sunny is encoded as (1, 0, 0, 0), rain as (0, 1, 0, 0), snow as (0, 0, 1, 0), and fog as (0, 0, 0, 1). In road surface condition, dry is encoded as (1, 0, 0), wet as (0, 1, 0), and icy as (0, 0, 1). For continuous data, such as light intensity, the Min-Max normalization method is used to map the data to the interval between 0 and 1. The normalization formula is: normalized value = (original value minus minimum value) / (maximum value minus minimum value), where the minimum value of light intensity is set to 0 lux and the maximum value is set to 1000 lux. The encoded discrete data vector is concatenated with the normalized continuous data vector to form a 1024-dimensional environmental feature vector.
[0035] Step S340: Unify the dimensions of the visual feature vector, the text semantic vector, and the environment feature vector, and generate the traffic feature alignment vector by combining the timestamp and location label. Specifically, for the visual feature vector, text semantic vector, and environment feature vector, if the original dimensions are not 1024, a fully connected neural network is used for dimension mapping. This neural network includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is the original vector dimension, the hidden layer contains two layers, each with 2048 neurons, and the ReLU function is used as the activation function. The number of neurons in the output layer is 1024. The trained neural network maps all types of vectors to 1024 dimensions. Then, using the timestamp and location label as a joint index, find the visual feature vector, text semantic vector, and environment feature vector under the same index, and concatenate the three in order to form a 3072-dimensional traffic feature alignment vector.
[0036] Step S400: Input the traffic feature alignment vector into a pre-trained multimodal large model to obtain a multi-dimensional preliminary semantic understanding result. The traffic feature alignment vector is subjected to feature deep fusion and generative inference through the cross-modal attention mechanism of the multimodal large model, and the preliminary semantic understanding result including target dimension, behavior dimension, scene dimension and event dimension is output.
[0037] Specifically, the Flamingo-8B pre-trained multimodal large model was selected. This model includes modules such as a visual encoder, a text encoder, a cross-modal attention layer, and a generator. The visual encoder is used to further extract visually relevant features, the text encoder is used to deepen text semantic understanding, the cross-modal attention layer achieves deep fusion of visual features, text features, and environmental features by calculating attention weights between features of different modalities, and the generator adopts an autoregressive language model structure. The traffic feature alignment vector was input into the model, and the maximum output length of generative inference was set to 512 characters, the temperature parameter was 0.7, the top_k parameter was 50, and the top_p parameter was 0.95. The model captures the associations between three types of features through a cross-modal attention mechanism, such as the association between vehicle speed in visual features and speed limit standards in text features, and the association between rainy weather in environmental features and vehicle braking status in visual features. It also performs autoregressive inference through a generator to output preliminary semantic understanding results, including target dimensions (target type, quantity, attributes), behavior dimensions (vehicle driving status, whether pedestrian behavior is compliant), scene dimensions (intersection type, current time scene label), and event dimensions (whether congestion, accidents, violations, etc. have occurred).
[0038] Step S500: Perform multi-dimensional semantic analysis and optimization on the preliminary semantic understanding results to obtain a structured semantic report.
[0039] Specifically, the preliminary semantic understanding results are first analyzed in layers to clarify the specific information of each dimension. Then, the traffic regulations text knowledge base is called to perform rule verification on the analysis results and eliminate misjudgments due to rule mismatch. Next, the semantic association between the target trajectory feature sequence and continuous frames is used to correct the single-frame analysis deviation. Finally, the optimized results are organized in a structured format, including analysis information of each dimension, confidence level, key frame screenshot references, target trajectory visualization data references, rule basis references, etc., to generate a structured semantic report.
[0040] In one possible implementation, the preliminary semantic understanding result is subjected to multi-dimensional semantic analysis and optimization to obtain a structured semantic report. Step S500 further includes step S510, which involves performing hierarchical semantic analysis on the preliminary semantic understanding result to obtain multi-dimensional traffic semantic analysis results. The hierarchical semantic analysis includes target attribute analysis, behavior and action analysis, scene type analysis, and event level analysis. Specifically, target attribute analysis is as follows: Based on the fused features and the preliminary semantic understanding result, the specific information of the target is clarified. For vehicles, the license plate number, vehicle type, color, and whether it is operating are analyzed; for pedestrians, whether they are adults and whether they are carrying items are analyzed; for traffic facilities, the status of traffic lights and the content of traffic signs are analyzed. The analysis accuracy is set to be no less than 95%. Behavior and action analysis is as follows: Based on traffic rules, it is determined whether vehicles are speeding, running red lights, illegally changing lanes, or driving against traffic; and whether pedestrians are running red lights, jaywalking, or walking in the motor vehicle lane are determined. The behavior recognition accuracy is set to be no less than 92%. Scene type analysis is as follows: Combining intersection attributes, time period, and environmental data, scene labels are added, such as school commuting scenarios, peak holiday scenarios in commercial areas, commuting scenarios on rainy days, and road construction scenarios. The scene matching accuracy is set to be no less than 90%. Event level analysis is as follows: Based on the target behavior and scene characteristics, the type and level of traffic events are determined. Violation events are divided into general violations and serious violations; congestion events are divided into light congestion, moderate congestion, and heavy congestion; and accident events are divided into minor accidents, general accidents, and major accidents. The event recognition delay is set to be no more than 2 seconds.
[0041] Step S520: The traffic rule knowledge base is invoked to verify and optimize the multi-dimensional traffic semantic analysis results, obtaining initial optimized traffic semantic results. Specifically, the traffic rule knowledge base is stored in a relational database and includes tables of legal clauses, intersection management rules, and historical cases. Keyword indexes are established for each table, such as road segment number, rule type, and violation name. For each judgment result in the multi-dimensional traffic semantic analysis, keywords are extracted, such as road segment number, violation, and speed value as the basis for judgment. The corresponding rule is searched in the knowledge base using the keyword index, such as the speed limit standard for the road segment and the threshold for speeding. The parameters in the analysis results are compared with the rule parameters in the knowledge base. If they match, the judgment result is retained; if they do not match, it is judged as a misjudgment and is removed or corrected to obtain the initial optimized traffic semantic results.
[0042] Step S530: Contextual correction is performed on the initial traffic semantic optimization results using the target trajectory feature sequence and the semantic association of consecutive frames, generating the structured semantic report. Specifically, for each judgment result in the initial traffic semantic optimization results, the corresponding target trajectory feature sequence is retrieved, and the target's motion state in consecutive frames is analyzed, such as changes in position, speed, and direction. Simultaneously, the semantic parsing results of consecutive frames are compared to determine whether the judgment result of the current frame is consistent with the target's motion trend and the semantics of the consecutive frames. If inconsistent, it is determined to be a single-frame parsing deviation, and correction is performed based on the target's motion trend and the semantics of the consecutive frames. After correction, all results are organized in a structured format. The report includes multi-dimensional parsing information, the confidence level of each result, keyframe screenshot paths, target trajectory visualization data paths, rule references, etc., generating a structured semantic report.
[0043] Step S600: Periodically incrementally update and optimize the multimodal large model and traffic rule knowledge base based on the feedback dataset of the structured semantic report.
[0044] Specifically, an update cycle of two weeks is used. Manual review results are collected through the traffic management platform, and structured semantic reports are labeled, including correct, incorrect, and missed judgments, with the reasons for incorrect judgments recorded. Simultaneously, data on the usage of structured semantic reports by upper-layer applications, such as violation detection and congestion warning applications, is collected, including whether the applications adopted the report results and the processing effects after adoption. This data is organized into a labeled feedback dataset, which is used for incremental fine-tuning of the multimodal large model, optimizing the model's cross-modal fusion weights and inference logic, with a target of reducing the misjudgment rate by at least 5% per month. New traffic rules, new types of violations, and special scenario information are also added to the traffic rule knowledge base, and historical semantic parsing cases are structured and stored to form a case knowledge base.
[0045] In one possible implementation, the multimodal large model and traffic rule knowledge base are periodically and incrementally updated and optimized based on the feedback dataset of the structured semantic report. Step S600 further includes step S610, collecting the manual review results of the structured semantic report through the traffic management platform. Specifically, a manual review module for structured semantic reports is developed in the traffic management platform. This module includes a report display interface, annotation options, a custom opinion input box, and a drop-down menu for reasons for misjudgment, such as target recognition errors due to severe weather, lack of coverage of new types of violations, rule matching errors, and misjudgments due to occlusion. After logging into the platform, traffic management personnel review the structured semantic reports one by one according to their assigned workload, viewing the multi-dimensional analysis information, keyframe screenshots, target trajectory data, etc. in the report, and selecting annotation options according to the actual situation. If it is a misjudgment or omission, the corresponding reason can be selected from the drop-down menu, or a detailed explanation can be entered in the custom opinion input box. After the staff submits the review results, the platform automatically associates and stores the review results with the corresponding structured semantic report to form a manual review result dataset.
[0046] Step S620: Annotate the structured semantic report according to the manual review results to obtain semantic review tags. Specifically, formulate semantic review tag rules. Tags are divided into basic tags and detailed tags. Basic tags include three categories: correct, misjudged, and missed. Detailed tags are subcategories of misjudgments and missed judgments, such as misjudgment - target recognition error in severe weather, misjudgment - new type violation not covered, misjudgment - rule matching error, misjudgment - occlusion, missed judgment - congestion event, missed judgment - minor accident, etc. Based on the annotation options and reasons for misjudgments in the manual review results, assign corresponding basic tags and detailed tags to each structured semantic report. If the review result is correct, only the basic tag "correct" is assigned; if it is a misjudgment or missed judgment, both the basic tag and the corresponding detailed tag are assigned. Associate the semantic review tags with the unique identifier of the structured semantic report and store them in the feedback dataset database.
[0047] Step S630: Collect feedback data from the upper-layer application of the structured semantic report. Specifically, a data collection module is embedded in the upper-layer application, such as a violation detection application or a congestion warning application. This module interfaces with the output interface of the structured semantic report and sets the types of feedback data to be collected, including: whether the application adopts the semantic report (options: adopt, partially adopt, do not adopt), the specific content of adoption (e.g., adopting the violation judgment result, adopting the event level judgment result, etc.), the reason for non-adoption (options: inaccurate result, incomplete information, incompatible format, etc.), the processing efficiency after using the report (e.g., the time it takes for the violation detection application to complete the violation processing based on the report), and the evaluation of the processing effect (e.g., whether the violation was successfully detected, whether the congestion warning was timely, etc.). The collection frequency is set to real-time collection, that is, every time the upper-layer application uses a structured semantic report and generates feedback, the collection module immediately collects the relevant feedback data, associates it with the unique identifier of the report, and transmits it to the feedback dataset database for storage.
[0048] Step S640: Organize the semantic review tags and the upper-layer application feedback data to obtain the feedback dataset. Specifically, clean the semantic review tag data, removing invalid tags without corresponding report numbers and tags with duplicate annotations. If the same report is annotated multiple times, retain the latest annotation result. Clean the upper-layer application feedback data, removing data with format errors (such as missing key fields for whether the report was adopted) and logical contradictions (such as annotating adoption but providing inaccurate reasons for non-adoption). Standardize the format of both types of data, converting text-based feedback information into encoded forms, such as encoding inaccurate results as 01, incomplete information as 02, successful violation detection as 10, and untimely warning as 11, etc. Finally, integrate the two types of data according to the structure of report number, basic semantic review tags, detailed semantic review tags, upper-layer application adoption status, adopted content encoding, non-adoption reason encoding, processing efficiency, and processing effect encoding to form a feedback dataset, which is stored as a CSV file for incremental fine-tuning of the multimodal large model and updating of the traffic rule knowledge base.
[0049] This application's embodiments employ real-time acquisition of traffic video streams, traffic management text data, and environmental sensor data, followed by preprocessing to obtain a traffic data sequence. Multimodal feature extraction and cross-modal feature alignment are performed on this sequence to establish a traffic feature alignment vector. This vector is then input into a pre-trained multimodal large model to obtain preliminary multidimensional semantic understanding results. Multidimensional semantic parsing and optimization generate a structured semantic report. Based on the feedback dataset from this report, the multimodal large model and traffic rule knowledge base are periodically and incrementally updated and optimized. These techniques address the technical problems of existing traffic video semantic understanding, such as a single semantic understanding dimension and insufficient recognition accuracy in complex scenarios. This achieves the technical effect of enriching semantic coverage dimensions and improving recognition accuracy in complex scenarios.
[0050] In the above text, refer to Figure 1 This paper describes in detail a method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model according to embodiments of the present invention. Next, reference will be made to... Figure 2 This invention describes a traffic video multi-dimensional semantic understanding system based on a multimodal large model according to an embodiment of the present invention.
[0051] The traffic video multi-dimensional semantic understanding system based on a multimodal large model according to embodiments of the present invention addresses the technical problems of existing traffic video semantic understanding systems, such as limited semantic understanding dimensions and insufficient recognition accuracy in complex scenes. It achieves the technical effects of enriching semantic coverage dimensions and improving recognition accuracy in complex scenes. The traffic video multi-dimensional semantic understanding system based on a multimodal large model includes: a multi-source data acquisition module 10, a preprocessing module 20, a cross-modal feature alignment module 30, a preliminary semantic understanding module 40, a multi-dimensional semantic parsing module 50, and an incremental update module 60.
[0052] The system includes a multi-source data acquisition module 10, which acquires multi-source data in real time, including traffic video streams, traffic management text data, and environmental sensor data; a preprocessing module 20, which preprocesses the multi-source data to obtain traffic data sequences; a cross-modal feature alignment module 30, which extracts multi-modal features and aligns cross-modal features in the traffic data sequences to establish traffic feature alignment vectors; a preliminary semantic understanding module 40, which inputs the traffic feature alignment vectors into a pre-trained multi-modal large model to obtain multi-dimensional preliminary semantic understanding results; a multi-dimensional semantic parsing module 50, which performs multi-dimensional semantic parsing and optimization on the preliminary semantic understanding results to obtain a structured semantic report; and an incremental update module 60, which performs periodic incremental updates and optimizations on the multi-modal large model and the traffic rule knowledge base based on the feedback dataset of the structured semantic report.
[0053] The preprocessing module 20 is described in detail below: As mentioned above, it preprocesses the multi-source collected data to obtain a traffic data sequence. The preprocessing module 20 may further include: a frame extraction unit for extracting frames from the traffic video stream to obtain a traffic video frame set; an enhancement processing unit for enhancing the traffic video frame set using an adaptive histogram equalization algorithm and a median filtering algorithm to obtain an enhanced video frame set; a moving target detection unit for detecting moving targets in the enhanced video frame set based on the inter-frame difference method to obtain a video keyframe sequence; a data cleaning unit for cleaning the traffic management text data and the environmental sensing data to obtain a traffic management text sequence and an environmental sensing sequence; and a traffic data sequence generation unit for generating the traffic data sequence based on the video keyframe sequence, the traffic management text sequence, and the environmental sensing sequence.
[0054] The cross-modal feature alignment module 30 is described in detail below: As mentioned above, it performs multimodal feature extraction and cross-modal feature alignment on the traffic data sequence to establish a traffic feature alignment vector. The cross-modal feature alignment module 30 may further include: a visual feature extraction unit for extracting visual features from the video keyframe sequence to obtain a visual feature vector; a text feature extraction unit for extracting text features from the traffic management text sequence to obtain a text semantic vector; an environmental feature extraction unit for extracting environmental features from the environmental sensing sequence to obtain an environmental feature vector; and a dimension unification unit for unifying the dimensions of the visual feature vector, the text semantic vector, and the environmental feature vector, and generating the traffic feature alignment vector by combining the timestamp and location label.
[0055] The detailed description of the specific configuration of the preliminary semantic understanding module 40 is explained as follows: As mentioned above, the preliminary semantic understanding module 40 may further include: performing feature deep fusion and generative reasoning on the traffic feature alignment vector through the cross-modal attention mechanism of the multimodal large model, and outputting the preliminary semantic understanding result including target dimension, behavior dimension, scene dimension and event dimension.
[0056] The detailed description of the specific configuration of the multi-dimensional semantic parsing module 50 is as follows: As mentioned above, the multi-dimensional semantic parsing module 50 performs multi-dimensional semantic parsing and optimization on the preliminary semantic understanding results to obtain a structured semantic report. The multi-dimensional semantic parsing module 50 may further include: a hierarchical semantic parsing unit for performing hierarchical semantic parsing on the preliminary semantic understanding results to obtain multi-dimensional traffic semantic parsing results; a verification and optimization unit for calling the traffic rule knowledge base to verify and optimize the multi-dimensional traffic semantic parsing results to obtain initial traffic semantic optimization results; and a context correction unit for using the target trajectory feature sequence and the semantic association of continuous frames to perform context correction on the initial traffic semantic optimization results to generate the structured semantic report.
[0057] The hierarchical semantic parsing unit may further include: target attribute parsing, behavior action parsing, scene type parsing, and event level parsing.
[0058] The detailed description of the specific configuration of the incremental update module 60 is explained as follows: As mentioned above, the multimodal large model and traffic rule knowledge base are periodically incrementally updated and optimized based on the feedback dataset of the structured semantic report. The incremental update module 60 may further include: a manual review result collection unit for collecting the manual review results of the structured semantic report through the traffic management platform; a labeling unit for labeling the structured semantic report based on the manual review results to obtain semantic review tags; an upper-layer application feedback data collection unit for collecting upper-layer application feedback data of the structured semantic report; and a feedback dataset acquisition unit for organizing the semantic review tags and the upper-layer application feedback data to obtain the feedback dataset.
[0059] The traffic video multi-dimensional semantic understanding system based on a multimodal large model provided in this embodiment of the invention can execute the traffic video multi-dimensional semantic understanding method based on a multimodal large model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0060] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0061] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model, characterized in that, The method includes: Real-time acquisition of multi-source data, including traffic video streams, traffic management text data, and environmental sensor data; The multi-source collected data is preprocessed to obtain a traffic data sequence; Multimodal feature extraction and cross-modal feature alignment are performed on the traffic data sequence to establish a traffic feature alignment vector; The traffic feature alignment vector is input into a pre-trained multimodal large model to obtain preliminary semantic understanding results in multiple dimensions. The preliminary semantic understanding results are subjected to multi-dimensional semantic analysis and optimization to obtain a structured semantic report; The multimodal large model and traffic rule knowledge base are periodically and incrementally updated and optimized based on the feedback dataset of the structured semantic report.
2. The method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model as described in claim 1, characterized in that, The multi-source collected data is preprocessed to obtain traffic data sequences, including: Frame extraction is performed on the traffic video stream to obtain a traffic video frame set; An adaptive histogram equalization algorithm and a median filtering algorithm are used to enhance the traffic video frame set to obtain an enhanced video frame set. The enhanced video frame set is subjected to moving target detection based on the inter-frame difference method to obtain the video keyframe sequence; Data cleaning is performed on the traffic management text data and the environmental sensor data to obtain traffic management text sequences and environmental sensor sequences; The traffic data sequence is generated based on the video keyframe sequence, the traffic management text sequence, and the environmental sensing sequence.
3. The method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model as described in claim 1, characterized in that, Multimodal feature extraction and cross-modal feature alignment are performed on the traffic data sequence to establish a traffic feature alignment vector, including: Visual feature vectors are obtained by extracting visual features from keyframe sequences in the video. Text feature extraction is performed on traffic management text sequences to obtain text semantic vectors; Environmental features are extracted from the environmental sensing sequences to obtain environmental feature vectors; The visual feature vector, the text semantic vector, and the environmental feature vector are unified in dimension, and the traffic feature alignment vector is generated by combining the timestamp and location label.
4. The method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model as described in claim 1, characterized in that, The traffic feature alignment vector is subjected to deep feature fusion and generative inference through the cross-modal attention mechanism of the multimodal large model, and the preliminary semantic understanding results including target dimension, behavior dimension, scene dimension and event dimension are output.
5. The method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model as described in claim 1, characterized in that, The preliminary semantic understanding results are subjected to multi-dimensional semantic analysis and optimization to obtain a structured semantic report, including: The preliminary semantic understanding results are subjected to hierarchical semantic parsing to obtain multi-dimensional traffic semantic parsing results; The traffic rule knowledge base is invoked to verify and optimize the multidimensional traffic semantic parsing results, and the initial optimization results of traffic semantics are obtained. The initial traffic semantic optimization results are corrected by using the target trajectory feature sequence and the semantic association of consecutive frames to generate the structured semantic report.
6. The method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model as described in claim 5, characterized in that, The hierarchical semantic parsing includes target attribute parsing, behavior and action parsing, scene type parsing, and event level parsing.
7. The method for multi-dimensional semantic understanding of traffic videos based on a multimodal large model as described in claim 1, characterized in that, Based on the feedback dataset from the structured semantic report, the multimodal large model and traffic rule knowledge base are periodically and incrementally updated and optimized, including: The results of manual review of the structured semantic reports are collected through the traffic management platform; The structured semantic report is annotated based on the results of the manual review to obtain semantic review tags; Collect feedback data from the upper-layer application of the structured semantic report; Organize the semantic review tags and the feedback data from the upper-layer application to obtain the feedback dataset.
8. A traffic video multi-dimensional semantic understanding system based on a multimodal large model, characterized in that, The system is used to implement the traffic video multi-dimensional semantic understanding method based on a multimodal large model as described in any one of claims 1-7, and the system comprises: The multi-source data acquisition module is used to acquire multi-source data in real time, including traffic video streams, traffic management text data, and environmental sensor data. The preprocessing module is used to preprocess the multi-source collected data to obtain traffic data sequences; The cross-modal feature alignment module is used to extract multimodal features and align cross-modal features of the traffic data sequence, and to establish a traffic feature alignment vector; The preliminary semantic understanding module is used to input the traffic feature alignment vector into a pre-trained multimodal large model to obtain multi-dimensional preliminary semantic understanding results; The multi-dimensional semantic parsing module is used to perform multi-dimensional semantic parsing and optimization on the preliminary semantic understanding results to obtain a structured semantic report; The incremental update module is used to perform periodic incremental updates and optimizations on the multimodal large model and the traffic rule knowledge base based on the feedback dataset of the structured semantic report.
Citation Information
Cited By
A large model-based traffic event review and adaptive optimization system
CN122336654A