Multi-modal large model collaborative sensing method and system for smart traffic
By employing a collaborative perception method combining multimodal sensors and large models, the problem of inaccurate data correlation in intelligent transportation has been solved, achieving comprehensive traffic perception and efficient decision feedback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTELLIGENT INTER CONNECTION TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-01
AI Technical Summary
The current intelligent transportation perception system suffers from inaccurate data association and limitations of single-modal perception, resulting in one-sided traffic perception, low data fusion accuracy, and weak decision feedback support.
Traffic data is collected by multimodal sensors and projected onto the traffic topology according to the spatiotemporal positioning of the data source. Multimodal large model is used to perform individual modal perception and modal fusion perception, output single modal and fused modal perception information, and identify the correlation effects in the traffic topology to output traffic perception feedback results.
It improves the comprehensiveness and accuracy of traffic perception, enhances the efficiency of decision feedback, and enables effective collaboration of multimodal data and accurate information correction.
Smart Images

Figure CN121963466A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation perception technology, specifically to a multimodal large-scale collaborative perception method and system for intelligent transportation. Background Technology
[0002] In the field of intelligent transportation, traffic perception is the core foundation for achieving refined management and efficient scheduling. However, existing perception technologies still face many bottlenecks: the data collected by multimodal sensors lacks a spatiotemporal positioning and projection mechanism with traffic scenarios, resulting in fragmented data that is difficult to effectively coordinate; single-modal perception networks have insufficient adaptability to feature extraction for specific types of data, easily leading to one-sided feature capture; modal fusion often remains at a single level, lacking a progressive fusion design from features to semantics to decision-making, resulting in limited fusion accuracy and depth; at the same time, there is a lack of effective information coordination, correction, and enhancement mechanisms between nodes, making it difficult to solve the problems of conflicting and redundant perception data. Ultimately, this leads to insufficient comprehensiveness and accuracy of traffic perception, failing to provide accurate and efficient support for traffic management decisions and hindering the intelligent upgrading of intelligent transportation systems.
[0003] Existing technologies for intelligent transportation perception suffer from inaccurate data association and limitations of single-modal perception, resulting in one-sided traffic perception, low data fusion accuracy, and weak decision feedback support. Summary of the Invention
[0004] This application provides a multimodal large-scale collaborative perception method and system for intelligent transportation, which addresses the technical problems in existing intelligent transportation perception, such as inaccurate data association, limitations of single-modal perception, resulting in one-sided traffic perception, low data fusion accuracy, and weak decision feedback support.
[0005] In view of the above problems, this application provides a multimodal large-scale collaborative perception method and system for intelligent transportation.
[0006] The first aspect of this application provides a multimodal large-model collaborative perception method for intelligent transportation, the method comprising:
[0007] Traffic data is collected by multimodal sensors and projected onto the traffic topology according to the spatiotemporal positioning of the data source. The collected traffic data is subjected to modal perception and modal fusion perception through a built-in multimodal large model, and single-modal perception information and fused modal perception information are output, wherein the modal fusion perception includes multiple fusion layers. Based on the single-modal perception information and fused modal perception information, the correlation influence is identified in the traffic topology, and the traffic perception feedback result is output.
[0008] A second aspect of this application provides a multimodal large-model collaborative perception system for intelligent transportation, the system comprising:
[0009] The traffic data acquisition module is used to collect traffic data through multimodal sensors and project it onto the traffic topology according to the spatiotemporal positioning of the acquisition source. The modal perception module is used to perform modal perception and modal fusion perception on the collected traffic data through a built-in multimodal large model, and output single-modal perception information and fused modal perception information, wherein the modal fusion perception includes multiple fusion levels. The feedback result output module is used to identify the correlation effects in the traffic topology based on the single-modal perception information and fused modal perception information, and output the traffic perception feedback result.
[0010] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0011] Traffic data is collected by multimodal sensors and projected onto the traffic topology according to the spatiotemporal positioning of the data source. A built-in multimodal large model performs modal perception and modal fusion perception on the collected traffic data, outputting single-modal perception information and fused modal perception information, where the modal fusion perception includes multiple fusion levels. Based on the single-modal perception information and fused modal perception information, correlation and influence identification is performed on the traffic topology, and traffic perception feedback results are output. This achieves the technical effect of improving the comprehensiveness, accuracy, and efficiency of traffic perception and decision feedback. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of the multimodal large-model collaborative perception method for intelligent transportation provided in an embodiment of this application;
[0014] Figure 2 This is a schematic diagram of the structure of a multimodal large-scale collaborative perception system for intelligent transportation provided in an embodiment of this application.
[0015] Figure labeling: Traffic data acquisition module 10, modal perception module 20, feedback result output module 30. Detailed Implementation
[0016] This application provides a multimodal large-scale collaborative perception method and system for intelligent transportation, which addresses the technical problems in existing intelligent transportation perception, such as inaccurate data association, limitations of single-modal perception, resulting in one-sided traffic perception, low data fusion accuracy, and weak decision feedback support.
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0018] Example 1, as Figure 1 As shown, this application provides a multimodal large-model cooperative perception method for intelligent transportation, the method comprising:
[0019] Step S100: Collect traffic data through multimodal sensors and project it into the traffic topology according to the spatiotemporal positioning of the data source.
[0020] Specifically, relying on multimodal sensors such as cameras, millimeter-wave radar, lidar, and infrared sensors, various types of traffic data, including images, radar signals, 3D point clouds, and infrared images, are comprehensively collected in traffic scenarios. First, a corresponding traffic topology is constructed with the perception-associated execution target as the central node. Then, based on the actual deployment location and perception coverage of each multimodal sensor, a positioning mapping relationship between the sensor and the traffic topology is established. Finally, according to the acquisition source identifier and acquisition timestamp corresponding to the multimodal acquisition data, various types of traffic data are accurately associated and projected into the traffic topology based on the established positioning mapping relationship, thereby achieving comprehensive feedback on the real-time monitoring status and historical monitoring status of traffic nodes.
[0021] Step S200: The collected traffic data is subjected to various modal perceptions and modal fusion perceptions through the built-in multimodal large model, and single-modal perception information and fused modal perception information are output, wherein the modal fusion perception includes multiple fusion layers.
[0022] Specifically, relying on the built-in multimodal large model, which includes a modality-specific network and a shared feature network, the collected traffic data is processed. First, the raw data of each modality is preprocessed for denoising. Then, feature extraction is performed through the corresponding modality-specific network to output single-modal perception information. The visual modality-specific network uses a convolutional neural network to extract key image features, the radar modality-specific network uses a regression network or a multilayer perceptron to extract the distance and velocity features of radar signals, the lidar modality-specific network uses a point cloud processing network to process 3D point cloud data and extract spatial features, and the infrared modality-specific network uses a convolutional neural network or a recurrent neural network to extract infrared image features to achieve target detection and abnormal behavior feature recognition. Subsequently, the recognition features of each modality are mapped to the shared feature network for multimodal shared fusion processing. This shared feature network includes a multi-level fusion path, including feature fusion path, semantic fusion path, and decision fusion path. It sequentially completes the unified fusion of primary features of different modalities, the extraction of high-level semantic and spatiotemporal correlation information based on the fused primary features, and the decision fusion based on the fused semantic features. Finally, single-modal perception information and fused modal perception information containing fused perception information at each level are output.
[0023] Step S300: Based on the single-modal perception information and the fused modal perception information, identify the correlation effects in the traffic topology and output the traffic perception feedback results.
[0024] Specifically, using traffic topology as the core framework, this study first analyzes the complementarity and conflict between single-modal and fused modal perception information at each node. Simultaneously, it assesses the consistency of multimodal descriptions of the same traffic element or state among adjacent or related nodes at the spatial, temporal, and logical levels, accurately identifying perception conflict areas and information redundancy areas in the topological association. Then, guided by the spatial connectivity and traffic flow constraints of the traffic topology, a cross-node perception coordination mechanism is constructed. For perception conflict areas, reliable modal or node information is determined based on the node's positional relationship within the topology and traffic rules, and conflict information is corrected online. For information redundancy areas, complementary fusion is performed based on the confidence level of multimodal information. Simultaneously, the perception results from high-confidence nodes or paths in the topology are used to guide and enhance the perception process of adjacent or downstream low-confidence nodes through a graph neural network message passing mechanism. Finally, based on the enhanced fused modal perception information obtained after processing by the perception coordination mechanism, in-depth correlation and impact analysis is conducted within the traffic topology, generating traffic perception feedback results that include event correlations, impact assessment scope, and decision-making suggestions.
[0025] In one possible implementation, step S100 further includes:
[0026] Step S110: Construct a relevant traffic topology structure based on the perception-related execution target as the central node.
[0027] Step S120: Establish a positioning mapping relationship with the traffic topology based on the deployment location and sensing range of the multimodal sensors.
[0028] Step S130: Based on the collection source and timestamp of the multimodal data, the data is correlated and projected into the traffic topology according to the positioning mapping relationship, so as to provide feedback on the real-time and historical monitoring status of traffic nodes.
[0029] Specifically, with the perception-related execution target as the core central node, this target can focus on traffic lights. It is necessary to focus on their control logic and scope of influence, while also incorporating related traffic elements such as road network structure and intersections directly related to traffic lights. The system systematically sorts out the spatial connection relationships and logical interaction rules between the central node and surrounding adjacent intersections, related lanes, travel paths, and surrounding traffic facilities, clarifying the subordination, association, and influence paths between each element, and constructing a relevant traffic topology structure that can accurately map the spatial layout, element association, and control logic of traffic scenarios.
[0030] First, determine the actual deployment coordinates of multimodal sensors such as cameras, millimeter-wave radar, lidar, and infrared sensors, and accurately define the effective sensing range of each type of sensor, including parameters such as spatial coverage boundaries, signal detection distance, and sensing angle. Then, combine the spatial distribution and range division of elements such as nodes, road segments, and intersections in the constructed traffic topology, and establish the positioning mapping relationship between each sensor and the corresponding node, associated road segment, or specific area in the traffic topology through spatial coordinate matching and sensing coverage area association, so as to ensure that the various types of traffic data collected by the sensors can be accurately mapped to the specific spatial location of the topology.
[0031] For various traffic data collected by multimodal sensors such as cameras, millimeter-wave radar, lidar, and infrared sensors, including images, radar signals, 3D point clouds, and infrared images, the system first extracts the source identifier for each data point to identify the specific sensor from which the data originates and the collection timestamp. Then, based on the established positioning mapping relationship between multimodal sensors and traffic topology, the system accurately associates and projects various traffic data to the corresponding nodes, road segments, or specific areas of the traffic topology according to the topological location corresponding to the collection source and the time sequence corresponding to the timestamp. By updating the current collected data in real time to provide feedback on the real-time operating status of traffic nodes, and by retaining historical collected data to form a complete historical monitoring record, the system achieves comprehensive and accurate feedback on the real-time and historical monitoring status of traffic nodes.
[0032] In one possible implementation, step S120 further includes:
[0033] The multimodal sensor includes at least: a camera, millimeter-wave radar, lidar, and infrared sensor.
[0034] Specifically, the multimodal sensors encompass various device types adapted to the data acquisition needs of intelligent transportation scenarios, including at least cameras, millimeter-wave radar, lidar, and infrared sensors. Each sensor performs its own function while complementing each other: cameras can accurately collect image data in traffic scenarios, providing a foundation for visual modal perception; millimeter-wave radar can capture key motion features such as distance and speed of targets, adapting to target detection in complex environments; lidar can generate high-precision three-dimensional point cloud data, helping to extract detailed features in spatial dimensions; and infrared sensors can collect infrared images in special environments such as low light and nighttime, enabling target detection and abnormal behavior feature recognition, collectively providing rich and comprehensive raw traffic data for the multimodal large model.
[0035] In one possible implementation, step S200 further includes:
[0036] The multimodal large model includes a modality-specific network and a shared feature network. After denoising preprocessing, the raw data of each modality is used to extract features through the corresponding modality-specific network to output single-modality perception information. The recognition features of each modality are mapped to the shared feature network for multimodal sharing and fusion processing to output fused modality perception information.
[0037] The shared feature network includes a multi-level fusion path, which consists of a feature fusion path, a semantic fusion path, and a decision fusion path, to obtain fused perception information at each level.
[0038] Specifically, the multimodal large model adopts a two-layer architecture of modality-specific extraction and shared fusion, taking into account both feature mining of each modality's data and deep integration of cross-modal information: First, targeted denoising preprocessing is performed on the raw traffic data of each modality, such as images, radar signals, 3D point clouds, and infrared images collected by cameras, millimeter-wave radar, lidar, and infrared sensors, to filter out invalid information such as environmental interference and equipment noise, ensuring data purity; then, the preprocessed modality data of each type are input into the corresponding modality-specific network, where the visual modality extracts key image features through a convolutional neural network, and the radar modality... The system uses regression networks or multilayer perceptrons to capture distance and velocity features, LiDAR modalities to extract 3D spatial features using point cloud processing networks, and infrared modalities to achieve target detection and abnormal behavior feature recognition through convolutional neural networks or recurrent neural networks. Each dedicated network independently completes feature extraction and outputs accurate single-modal perception information. Finally, the recognition features of each modality are uniformly mapped to a shared feature network. Through the multimodal sharing and fusion mechanism of this network, the complementary information between different modalities is fully explored and the associated features are integrated to achieve effective fusion of cross-modal data. The final output is fused modal perception information that is both comprehensive and accurate.
[0039] The shared feature network incorporates a multi-level fusion system consisting of a feature fusion path, a semantic fusion path, and a decision fusion path. Through progressive processing logic, it achieves deep integration of multimodal information, thereby obtaining fused perception information at each level: The feature fusion path, as the basic fusion stage, focuses on the initial integration of original features from different modalities. It unifies and fuses the primary features of visual, radar, point cloud, and infrared modalities from cameras, millimeter-wave radar, lidar, and infrared sensors, generating standardized and integrated feature representations to lay the foundation for subsequent higher-order fusion. The semantic fusion path, based on feature fusion, delves into the intrinsic relationships between modalities, conducting high-level semantic representation fusion based on unified primary features. It focuses on extracting the spatiotemporal correlation information contained in different modal data, achieving information dimensionality upgrade from the feature level to the semantic level. The decision fusion path, as the final fusion stage, uses the fused semantic features as the core, combining the actual needs of the traffic scenario with decision logic to perform multimodal decision fusion, integrating the advantages of each modality to form accurate and reliable decision results. The three paths are sequentially connected and work synergistically to output fused perception information at the feature level, semantic level, and decision level, respectively.
[0040] In one possible implementation, step S200 further includes:
[0041] The feature fusion pathway is used to fuse primary features from different modalities to generate a unified feature representation.
[0042] The semantic fusion pathway is used to perform high-level semantic representation fusion based on the fused primary features, and to extract spatiotemporal correlation information between modalities.
[0043] The decision fusion pathway is used to perform decision fusion based on the fused semantic features and output the final traffic perception feedback result.
[0044] Specifically, the feature fusion pathway specifically receives primary features extracted from dedicated networks of different modalities such as vision, radar, lidar, and infrared. These include key image features output by convolutional neural networks, distance and velocity features of radar signals extracted by regression networks or multilayer perceptrons, three-dimensional spatial features obtained by point cloud processing networks, and target and abnormal behavior features of infrared images extracted by convolutional neural networks or recurrent neural networks. Through algorithms such as unified feature dimension mapping, cross-modal redundancy filtering, and complementary information integration, the independent and scattered primary features of each modality are systematically fused to eliminate feature differences and information fragmentation between modalities, ultimately generating a standardized and integrated unified feature representation.
[0045] The semantic fusion pathway is based on the unified feature representation generated by the feature fusion pathway. It focuses on the in-depth mining and cross-modal integration of high-level semantic information. Through semantic parsing, association modeling and other technologies, it performs dimensionality-upgrading on the fused primary features. It focuses on capturing the spatiotemporal correlation information of traffic scenes contained in different modal data such as vision, radar, lidar and infrared. For example, the matching relationship between vehicle driving trajectory and traffic light timing, the spatial linkage characteristics of pedestrian flow at intersections and vehicle flow in adjacent road segments, and the temporal correlation of traffic flow changes at different times. At the same time, it explores the complementary and correlation logic of semantic levels between modalities, eliminates the semantic barriers between modalities, and realizes the effective transformation from basic features to semantic information.
[0046] As the final link in the multi-level fusion of the shared feature network, the decision fusion path takes the fused semantic features rich in spatiotemporal correlation information output by the semantic fusion path as the core input. Combining the traffic rule constraints, practical application needs and decision logic of the intelligent transportation scenario, it adopts a multimodal decision fusion strategy, which integrates the advantages of semantic features from various modalities such as vision, radar, lidar, and infrared. It effectively avoids the decision bias and information limitations that may occur in complex traffic environments with a single modality. Through comprehensive analysis of key dimensions such as the nature, scope of impact, and development trend of traffic events, it finally outputs accurate, reliable, and practically valuable traffic perception feedback results.
[0047] In one possible implementation, step S220 further includes:
[0048] A dedicated network for visual modalities uses convolutional neural networks to extract key features from images.
[0049] The dedicated network for radar modes uses regression networks or multilayer perceptrons to extract the range and velocity characteristics of radar signals.
[0050] The dedicated network for lidar modes uses a point cloud processing network to process 3D point cloud data and extract spatial features.
[0051] The dedicated network for infrared modalities uses convolutional neural networks or recurrent neural networks to extract infrared image features for target detection and abnormal behavior feature recognition.
[0052] Specifically, the dedicated network for visual modalities employs a Convolutional Neural Network (CNN) as its core architecture, fully leveraging its advantages in image feature extraction. It performs precise feature mining on intelligent traffic scene images captured by cameras, encompassing various traffic elements such as vehicles, pedestrians, traffic lights, traffic signs, and road markings. Through the synergistic effect of multiple convolutional and pooling layers, the network first downsamples and enhances the original image, gradually filtering redundant pixel information. Then, through local receptive fields and weight sharing mechanisms, it efficiently captures low-level visual features such as edges, textures, and shapes in the image. These features are further fused to generate mid-to-high-level key features such as vehicle outlines, traffic light colors, and sign text shapes, ultimately outputting a feature vector that accurately represents the visual information of the traffic scene.
[0053] The dedicated network for radar modes employs a regression network or multilayer perceptron as its core architecture, adapting to the characteristics of radar signal data in traffic scenarios acquired by millimeter-wave radar, and focusing on the accurate extraction of key motion parameters of target objects. The network first performs denoising preprocessing on the raw radar signal to filter out environmental interference and equipment noise. Then, through the fitting and prediction capabilities of the regression network or the nonlinear mapping function of the multilayer perceptron, feature mining and parameter estimation are performed on the preprocessed radar signal. By leveraging weight updates between network layers and the gradual transformation of signal features, the network accurately captures the distance characteristics between the target object and the radar sensor, as well as the target object's moving speed characteristics. Simultaneously, it effectively avoids interference caused by multi-target occlusion and signal superposition in complex traffic environments, ultimately outputting distance and velocity feature vectors that accurately characterize the target's motion state.
[0054] The dedicated network for LiDAR modal analysis uses a point cloud processing network as its core architecture, precisely adapting to the characteristics of 3D point cloud data acquired by LiDAR and focusing on the efficient extraction of spatial features of targets in traffic scenarios. The network first preprocesses the raw 3D point cloud data, filtering out environmental noise and redundant points through point cloud denoising, downsampling, and registration, thus optimizing data quality. Then, using core algorithms such as point cloud segmentation and feature encoding, it performs deep processing on the preprocessed point cloud data. This enables accurate identification of the basic spatial features of targets such as vehicles, pedestrians, and road facilities, including their spatial location, contour shape, and size, while also capturing the relative spatial relationships between targets, effectively mining the spatial structural information contained in the 3D point cloud. Finally, it outputs feature vectors that accurately characterize the spatial layout of traffic scenarios and the spatial attributes of targets, providing reliable LiDAR modal spatial feature support for multimodal fusion perception.
[0055] A dedicated network for infrared modal analysis is adapted to the acquisition characteristics of infrared sensors in special traffic scenarios such as low light and nighttime. Employing convolutional neural networks (CNNs) or recurrent neural networks (RNNs) as its core architecture, the network focuses on feature extraction and key information recognition from infrared images. The network first preprocesses the acquired infrared images, performing denoising and enhancement to optimize image quality and highlight target contours and temperature differences. Then, leveraging the local feature capture capabilities of CNNs or the temporal dependency modeling advantages of RNNs, it deeply mines the key features contained in the infrared images. This enables accurate detection of the position and shape of traffic targets such as vehicles, pedestrians, and obstacles, and also keenly identifies the feature information corresponding to abnormal traffic behaviors such as driving against traffic, lane obstruction, and stagnation, effectively overcoming the limitations of lighting conditions. Finally, it outputs feature vectors that combine target localization and abnormal behavior representation, providing all-time, interference-resistant infrared modal data support for multimodal fusion perception, ensuring the continuity and accuracy of traffic perception in complex environments.
[0056] In one possible implementation, step S300 further includes:
[0057] Step S310: Using the traffic topology as a framework, analyze the complementarity and conflict between single-modal perception information and fused modal perception information at each node. At the same time, evaluate the consistency of multimodal descriptions of the same traffic element or state between adjacent or related nodes in space, time and logic, and identify perception conflict areas and information redundancy areas in the topological association.
[0058] Step S320: Guided by the spatial connectivity of the traffic topology and traffic flow constraints, a cross-node perception coordination mechanism is constructed. For the identified perception conflict areas, reliable modal or node information is determined based on the positional relationship of nodes in the topology and traffic rules, and the conflict information is corrected online.
[0059] Step S330: For information redundancy areas, perform complementary fusion based on the confidence level of multimodal information.
[0060] Step S340: Using the perception results of high-confidence nodes or paths in the topology, guide and enhance the perception process of adjacent or downstream low-confidence nodes through the message passing mechanism of graph neural networks.
[0061] Step S350: Based on the enhanced fusion modal perception information obtained after processing by the perception collaboration mechanism, perform correlation impact analysis on the traffic topology to generate traffic perception feedback results that include event correlation, impact assessment scope, and decision suggestions.
[0062] Specifically, using a pre-constructed traffic topology as the core analytical framework, and based on the spatial distribution and logical connections of each node within this structure, the analysis examines the relationship between single-modal perception information (such as visual, radar, lidar, and infrared) and multi-level fused modal perception information at each node. This involves both exploring the supplementary and explanatory value of different modal information for the same traffic scene element (complementarity) and identifying potential contradictions (conflicts) between the two types of information when representing the same traffic state. Furthermore, the analysis comprehensively evaluates the degree of fit between adjacent or logically related nodes for the same traffic element, such as a specific vehicle, traffic light, or traffic state (e.g., road congestion level and traffic efficiency), considering the three dimensions of spatial location matching, temporal synchronization, and logical connection rationality. Through comprehensive comparison and verification of information throughout the entire topology, the analysis accurately identifies perceptual conflict areas with contradictory information and redundant areas with duplicate or invalid information within the topological connections.
[0063] Based on the spatial connectivity of nodes and road segments in the traffic topology, such as the connectivity logic between intersections and associated lanes, the road network connection between nodes, and traffic flow constraints, such as traffic direction, priority, and traffic light control timing, a perception collaboration mechanism capable of cross-node information interaction and collaborative optimization is constructed. For identified perception conflict areas, the relative position of the node in the conflict area within the topology, the strength of its association with surrounding nodes, and traffic rules, such as yield rules, traffic light phase logic, and traffic control requirements, are comprehensively evaluated to assess the reliability of different modal data in the conflict information, such as visual, radar, lidar, and infrared data, as well as the perception results of different nodes. Highly reliable information that conforms to spatial logic and traffic rules is selected as a benchmark, and contradictory conflict data is corrected online in real time to eliminate information bias and ensure the consistency and accuracy of perception information of each node in the traffic topology.
[0064] For the identified redundant information areas, the focus is on the repetitive representations or overlaps in multimodal sensing information such as vision, radar, lidar, and infrared within these areas. First, considering factors such as the acquisition accuracy, environmental adaptability, and historical data reliability of each modality sensor, a quantitative evaluation model is used to score the confidence level of each multimodal information, clarifying the credibility weight of different modal information. Then, based on these confidence weights, complementary fusion processing is carried out, prioritizing the retention of core effective content from high-confidence modal information while integrating supplementary information not covered by low-confidence modal information. Redundant, ineffective, and overlapping data are eliminated, avoiding information waste and enhancing the comprehensiveness and accuracy of information through multimodal complementarity. Ultimately, this results in concise, efficient, reliable, and complete fused information.
[0065] First, within the traffic topology, based on the accuracy, consistency, and reliability assessment results of the perceived information, nodes or paths with high perception confidence are selected. Their high-quality perception data, including key features of single-modality perception and effective information from fused modal perception, are used as benchmark references. Then, leveraging the message passing mechanism of graph neural networks, the feature information and decision logic corresponding to these high-confidence perception results are transmitted and disseminated to adjacent or downstream nodes with lower perception confidence through inter-node links. During this transmission process, low-confidence nodes, guided by the received high-confidence information, dynamically adjust the parameters of their own perception models, optimize feature extraction and information fusion strategies, and compensate for deficiencies in their own data collection or processing, thereby enhancing perception accuracy and reliability. This improves the overall performance and data quality of the entire traffic topology perception network, providing consistent and high-quality perception support for subsequent correlation impact analysis.
[0066] Based on the enhanced fusion modal perception information output by the perception collaboration mechanism, and after conflict correction, redundancy fusion, and cross-node perception enhancement processing, a graph neural network combined with a traffic scene knowledge graph is used to construct an association analysis model. Using the node connection relationships and road network weights of the traffic topology as a foundation, a node embedding algorithm is used to quantify the association strength between traffic events and surrounding nodes, and a time-series analysis model is used to trace the event propagation path. Simultaneously, a traffic flow dynamics model and spatiotemporal interpolation algorithm are integrated, and based on core indicators such as traffic density and speed in the enhanced perception data, the impact diffusion process of events on surrounding road segments is simulated, accurately defining the spatial boundaries and time span of the impact range. Finally, combined with a pre-set traffic rule base, road network capacity thresholds, and multi-objective optimization algorithms, for event types such as congestion and accidents, decision-making content such as traffic light timing adjustment schemes, detour route planning, and emergency resource dispatch suggestions is generated. Through structured data encapsulation, a traffic perception feedback result containing an event association graph, an impact range heatmap, and an executable decision list is generated, realizing the entire process from data analysis to decision output.
[0067] Example 2, based on the same inventive concept as the multimodal large-model collaborative perception method for intelligent transportation in the previous examples, such as... Figure 2 As shown, this application provides a multimodal large-scale collaborative perception system for intelligent transportation. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0068] Traffic data acquisition module 10 is used to acquire traffic data through multimodal sensors and project it into the traffic topology according to the spatiotemporal positioning of the acquisition source.
[0069] The modal perception module 20 is used to perform modal perception and modal fusion perception on the collected traffic data through the built-in multimodal large model, and output single modal perception information and fused modal perception information, wherein the modal fusion perception includes multiple fusion layers.
[0070] The feedback result output module 30 is used to identify the correlation effects in the traffic topology based on the single-modal perception information and the fused modal perception information, and output the traffic perception feedback result.
[0071] Furthermore, the system is also used to implement the following functions:
[0072] Based on the perception-related execution target as the central node, a relevant traffic topology is constructed; according to the deployment location and perception range of multimodal sensors, a positioning mapping relationship with the traffic topology is established; according to the acquisition source and timestamp of the multimodal data, data is associated and projected into the traffic topology based on the positioning mapping relationship to provide feedback on the real-time and historical monitoring status of traffic nodes.
[0073] Furthermore, the system is also used to implement the following functions:
[0074] The multimodal sensor includes at least: a camera, millimeter-wave radar, lidar, and infrared sensor.
[0075] Furthermore, the system is also used to implement the following functions:
[0076] The multimodal large model includes a modality-specific network and a shared feature network. After denoising preprocessing, the raw data of each modality is used to extract features through the corresponding modality-specific network to output single-modality perception information. The recognition features of each modality are mapped to the shared feature network for multimodal shared fusion processing to output fused modality perception information. The shared feature network includes multi-level fusion paths, namely feature fusion path, semantic fusion path, and decision fusion path, to obtain fused perception information at each level.
[0077] Furthermore, the system is also used to implement the following functions:
[0078] The feature fusion pathway is used to fuse primary features from different modalities to generate a unified feature representation; the semantic fusion pathway is used to perform high-level semantic representation fusion based on the fused primary features to extract spatiotemporal correlation information between modalities; the decision fusion pathway is used to perform decision fusion based on the fused semantic features to output the final traffic perception feedback result.
[0079] Furthermore, the system is also used to implement the following functions:
[0080] The dedicated network for the visual modality uses convolutional neural networks to extract key features from images; the dedicated network for the radar modality uses regressive networks or multilayer perceptrons to extract the distance and velocity features of radar signals; the dedicated network for the lidar modality uses point cloud processing networks to process 3D point cloud data and extract spatial features; and the dedicated network for the infrared modality uses convolutional neural networks or recurrent neural networks to extract infrared image features for target detection and abnormal behavior feature recognition.
[0081] Furthermore, the system is also used to implement the following functions:
[0082] Using the aforementioned traffic topology as a framework, this study analyzes the complementarity and conflict between single-modal perception information and fused modal perception information at each node. Simultaneously, it assesses the spatial, temporal, and logical consistency of multimodal descriptions of the same traffic element or state among adjacent or related nodes, identifying perception conflict areas and information redundancy areas in the topological association. Guided by the spatial connectivity and traffic flow constraints of the traffic topology, a cross-node perception coordination mechanism is constructed. For identified perception conflict areas, reliable modal or node information is determined based on the node's positional relationship within the topology and traffic rules, and conflict information is corrected online. For information redundancy areas, complementary fusion is performed based on the confidence level of multimodal information. Utilizing perception results from high-confidence nodes or paths in the topology, the perception process of adjacent or downstream low-confidence nodes is guided and enhanced through a graph neural network message passing mechanism. Based on the enhanced fused modal perception information obtained after processing by the perception coordination mechanism, an association impact analysis is performed within the traffic topology, generating traffic perception feedback results that include event associations, impact assessment scope, and decision-making suggestions.
[0083] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0084] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0085] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A multimodal large-scale collaborative perception method for intelligent transportation, characterized in that, include: Traffic data is collected by multimodal sensors and projected into the traffic topology according to the spatiotemporal location of the data source. The built-in multimodal large model performs multimodal perception and modal fusion perception on the collected traffic data, and outputs single-modal perception information and fused modal perception information, wherein the modal fusion perception includes multiple fusion layers; Based on the single-modal perception information and the fused modal perception information, the correlation influence is identified in the traffic topology, and the traffic perception feedback result is output.
2. The multimodal large-scale collaborative perception method for intelligent transportation according to claim 1, characterized in that, Traffic data is collected through multimodal sensors and projected onto the traffic topology according to the spatiotemporal location of the data source, including: Based on the perception-related execution target as the central node, a relevant traffic topology structure is constructed; Based on the deployment location and sensing range of the multimodal sensors, a positioning mapping relationship with the traffic topology is established; Based on the data acquisition source and timestamp of the multimodal data, the data is correlated and projected into the traffic topology according to the location mapping relationship, which is used to provide feedback on the real-time and historical monitoring status of traffic nodes.
3. The multimodal large-scale collaborative perception method for intelligent transportation according to claim 2, characterized in that, The multimodal sensor includes at least: a camera, millimeter-wave radar, lidar, and infrared sensor.
4. The multimodal large-scale collaborative perception method for intelligent transportation according to claim 1, characterized in that, The system utilizes a built-in multimodal large model to perform modal perception and modal fusion perception on the collected traffic data, outputting single-modal perception information and fused modal perception information, including: The multimodal large model includes a modality-specific network and a shared feature network. After the raw data of each modality is preprocessed by denoising, the corresponding modality-specific network is used to extract features and output single-modality perception information. The recognition features of each modality are mapped to the shared feature network for multimodal sharing and fusion processing, and fused modality perception information is output. The shared feature network includes a multi-level fusion path, which consists of a feature fusion path, a semantic fusion path, and a decision fusion path, to obtain fused perception information at each level.
5. The multimodal large-model collaborative perception method for intelligent transportation according to claim 4, characterized in that, The feature fusion pathway is used to fuse primary features from different modalities to generate a unified feature representation; The semantic fusion pathway is used to perform high-level semantic representation fusion based on the fused primary features, and to extract spatiotemporal correlation information between modalities; The decision fusion pathway is used to perform decision fusion based on the fused semantic features and output the final traffic perception feedback result.
6. The multimodal large-model collaborative perception method for intelligent transportation according to claim 4, characterized in that, A dedicated network for visual modalities uses convolutional neural networks to extract key features from images; The dedicated network for radar modes uses regression networks or multilayer perceptrons to extract the range and velocity characteristics of radar signals; The dedicated network for lidar modes uses a point cloud processing network to process 3D point cloud data and extract spatial features; The dedicated network for infrared modalities uses convolutional neural networks or recurrent neural networks to extract infrared image features for target detection and abnormal behavior feature recognition.
7. The multimodal large-scale collaborative perception method for intelligent transportation according to claim 4, characterized in that, Based on the single-modal perception information and the fused modal perception information, the correlation impact is identified in the traffic topology, and traffic perception feedback results are output, including: Using the traffic topology as a framework, the complementarity and conflict between single-modal perception information and fused modal perception information at each node are analyzed. At the same time, the consistency of multimodal descriptions of the same traffic element or state between adjacent or related nodes in space, time and logic is evaluated, and perception conflict areas and information redundancy areas in topological associations are identified. Guided by the spatial connectivity of traffic topology and traffic flow constraints, a cross-node perception collaboration mechanism is constructed. For the identified perception conflict areas, reliable modal or node information is determined based on the positional relationship of nodes in the topology and traffic rules, and the conflict information is corrected online. For areas with redundant information, complementary fusion is performed based on the confidence level of multimodal information; By utilizing the perception results of high-confidence nodes or paths in the topology, the perception process of adjacent or downstream low-confidence nodes is guided and enhanced through the message passing mechanism of graph neural networks. Based on the enhanced fusion modal perception information obtained after processing by the aforementioned perception collaboration mechanism, correlation impact analysis is performed on the traffic topology to generate traffic perception feedback results that include event correlation, impact assessment scope, and decision recommendations.
8. A multimodal large-scale collaborative perception system for intelligent transportation, characterized in that: The system is used to implement the multimodal large-model collaborative perception method for intelligent transportation as described in any one of claims 1-7, and the system comprises: The traffic data acquisition module is used to collect traffic data through multimodal sensors and project it into the traffic topology according to the spatiotemporal positioning of the data source. The modal perception module is used to perform modal perception and modal fusion perception on the collected traffic data through the built-in multimodal large model, and output single modal perception information and fused modal perception information, wherein the modal fusion perception includes multiple fusion layers; The feedback result output module is used to identify the correlation effects in the traffic topology based on the single-modal perception information and the fused modal perception information, and output the traffic perception feedback result.