Vehicle and road cloud collaborative sensing method and system based on multi-source data fusion
By using a vehicle-road-cloud collaborative perception method, and by integrating multi-source data from vehicle-side, roadside, and cloud control platforms, the problem of inaccurate decision-making caused by data heterogeneity is solved, thereby improving traffic management efficiency and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTELLIGENT INTER CONNECTION TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing vehicle-road cooperative perception technologies suffer from data heterogeneity, making it difficult to efficiently integrate multi-source data. This leads to inaccurate decision-making in complex traffic environments, affecting traffic management efficiency and safety.
By working together with vehicle-side perception units, roadside perception units, and cloud control platforms, a multi-source perception dataset is constructed, features are extracted and fused, vehicle-road situational awareness results are generated, risk assessment and decision reasoning are performed, and vehicle cooperative control commands are generated.
It enables optimal decision-making for vehicle-road cooperative systems in complex traffic environments, improving traffic management efficiency and road safety.
Smart Images

Figure CN121963465A_ABST
Abstract
Description
Vehicle-Road-Cloud Cooperative Perception Method and System Based on Multi-Source Data Fusion Technical Field
[0001] This application relates to the field of intelligent transportation technology, specifically to a vehicle-road-cloud collaborative perception method and system based on multi-source data fusion. Background Technology
[0002] Existing vehicle-road cooperative perception technologies face significant challenges in multi-source data fusion. With the increasing number of data sources from vehicles, roadside equipment, and cloud platforms, these heterogeneous data differ in format, type, and collection frequency, making efficient data fusion difficult. The accuracy and completeness of real-time data collected by vehicle sensors, roadside equipment, and cloud platforms in different environments also vary, further exacerbating the difficulty of fusion. Currently, existing solutions often rely on simple rules or models for data fusion, failing to fully explore the potential value within the data. This results in inaccurate perception results, impacting decision-making capabilities in complex traffic scenarios and ultimately affecting the overall efficiency and safety of traffic management.
[0003] In summary, existing technologies suffer from technical problems due to data heterogeneity, which makes it difficult to integrate multi-source data in vehicle-road cooperative perception, leading to inaccurate decision-making in complex traffic environments and further affecting traffic management efficiency. Summary of the Invention
[0004] The purpose of this application is to provide a vehicle-road-cloud cooperative perception method and system based on multi-source data fusion, in order to solve the technical problem in the prior art that due to data heterogeneity, vehicle-road cooperative perception is difficult to fuse multi-source data, resulting in inaccurate decision-making in complex traffic environments, which further affects the efficiency of traffic management.
[0005] To achieve the above objectives, this application provides a vehicle-road-cloud collaborative perception method and system based on multi-source data fusion.
[0006] Firstly, this application provides a vehicle-road-cloud cooperative perception method based on multi-source data fusion. This method is implemented through a vehicle-road-cloud cooperative perception system based on multi-source data fusion. The method includes: heterogeneous perception through vehicle-side sensing units, roadside sensing units, and a cloud control platform to form a multi-source sensing dataset; feature extraction based on the multi-source sensing dataset to generate a data feature set; vehicle-road-cloud cooperative perception based on the data feature set to generate a vehicle-road situational awareness result, where the result includes traffic target state parameters; risk assessment based on the traffic target state parameters; decision reasoning based on the risk assessment value; and construction of vehicle cooperative control and cooperative perception commands.
[0007] Optionally, vehicle-side perception data is collected through the vehicle-side perception unit, which includes a dataset of vehicle motion status and raw sensor data. Roadside perception data is collected through the roadside perception unit, which includes traffic participant information and road environment information collected by roadside cameras, millimeter-wave radar, and lidar. Cloud control platform data is collected through the cloud control platform, which includes historical traffic flow data and regional traffic event information.
[0008] Optionally, a network time protocol is constructed, and the vehicle-side sensing data, roadside sensing data, and cloud control platform data are traversed according to the network time protocol to generate unified timestamp data; a regional map network of the target area is constructed, and a coordinate system transformation is performed based on the regional map network to construct a global coordinate system; according to the unified timestamp data, the vehicle-side sensing data is mapped to the global coordinate system, and the target local coordinates are determined based on the motion state dataset according to the first mapping result; the roadside sensing data is mapped to the global coordinate system according to the unified timestamp data, and the target detection coordinates are determined based on the road environment information according to the second mapping result; the target local coordinates, the target detection coordinates, and the historical traffic flow data and the regional traffic event information are spatiotemporally aligned to construct the multi-source sensing dataset.
[0009] Optionally, a vehicle-side feature set is constructed by analyzing the target's local coordinates in conjunction with the vehicle-side perception data; a roadside feature set is constructed by analyzing the target's detection coordinates in conjunction with the roadside perception data; traffic flow distribution features are constructed by performing traffic flow analysis based on historical traffic flow data; semantic features are constructed by performing semantic analysis based on regional traffic event information; a spatiotemporal perception map is constructed; heterogeneous data fusion inference is performed based on the spatiotemporal perception map in conjunction with the vehicle-side feature set, the roadside feature set, the traffic flow distribution features, and the semantic features to generate a global traffic fusion state parameter; collaborative perception is performed based on the global traffic fusion state parameter to obtain short-term behavior prediction data; multi-layer fusion is performed based on the global traffic fusion state parameter in conjunction with the short-term behavior prediction data to generate traffic target state parameters, and the traffic target state parameters are added to the vehicle-road situational awareness result.
[0010] Optionally, the vehicle-side feature set and the roadside feature set are uploaded in real time via a low-latency communication link, and the traffic flow distribution features and the semantic features are periodically distributed or pushed on demand by the cloud control platform.
[0011] Optionally, vehicle-side image data is obtained by parsing the raw sensor data; convolution calculations are performed on the vehicle-side image data to extract pixel-level semantic features and target-level visual features; three-dimensional scene analysis is performed based on the target's local coordinates to extract local geometric structure features and global scene features; the motion state dataset is analyzed through a long short-term memory network to extract vehicle dynamic features; and the pixel-level semantic features, target-level visual features, local geometric structure features, global scene features, and vehicle dynamic features are concatenated to construct the vehicle-side feature set.
[0012] Optionally, based on the road environment information, roadside wide-angle image data is obtained through analysis; the roadside wide-angle image data is then segmented into a panoramic view to extract pixel-level semantic segmentation features and bounding box features; the target detection coordinates are tracked based on the traffic participant information to obtain target motion features and target interaction features; the pixel-level semantic segmentation features, the bounding box features, the target motion features, and the target interaction features are then fused to generate the roadside feature set.
[0013] Secondly, this application also provides a vehicle-road-cloud cooperative perception system based on multi-source data fusion, used to execute the vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in the first aspect. The vehicle-road-cloud cooperative perception system based on multi-source data fusion includes: a heterogeneous perception module, used to perform heterogeneous perception through vehicle-side perception units, roadside perception units, and a cloud control platform to form a multi-source perception dataset; a vehicle-road-cloud cooperative perception module, used to extract features based on the multi-source perception dataset to generate a data feature set, perform vehicle-road-cloud cooperative perception based on the data feature set, and generate a vehicle-road situational awareness result, the vehicle-road situational awareness result including traffic target state parameters; and a risk assessment module, used to perform risk assessment according to the traffic target state parameters, perform decision reasoning based on the risk assessment value, and construct vehicle cooperative control cooperative perception instructions.
[0014] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0015] Heterogeneous sensing is achieved through vehicle-mounted sensing units, roadside sensing units, and a cloud control platform, forming a multi-source sensing dataset. Feature extraction is performed on this dataset to generate a data feature set. Vehicle-road-cloud collaborative sensing is then performed based on this feature set to generate vehicle-road situational awareness results, which include traffic target state parameters. Risk assessment is conducted according to these traffic target state parameters, and decision reasoning is performed based on the risk assessment values to construct vehicle cooperative control and collaborative sensing commands. In other words, by combining data from vehicle-mounted sensing units, roadside sensing units, and the cloud control platform to form a multi-source sensing dataset, performing feature extraction on this dataset to generate a data feature set for vehicle-road-cloud collaborative sensing, conducting risk assessment based on the traffic target state parameters in the vehicle-road situational awareness results, and performing decision reasoning based on the risk assessment values to construct vehicle cooperative control commands, the vehicle-road cooperative system ensures that it can make optimal decisions in complex and ever-changing traffic environments, improving traffic management efficiency and road safety.
[0016] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 is a flowchart illustrating the vehicle-road-cloud collaborative perception method based on multi-source data fusion proposed in this application.
[0019] Figure 2 is a schematic diagram of the structure of the vehicle-road-cloud cooperative perception system based on multi-source data fusion in this application.
[0020] Figure labeling: Heterogeneous sensing module 11, vehicle-road-cloud collaborative sensing module 12, risk assessment module 13. Detailed Implementation
[0021] This application provides a vehicle-road-cloud cooperative perception method and system based on multi-source data fusion. It addresses the technical problem in existing technologies where data heterogeneity makes it difficult to integrate multi-source data for vehicle-road cooperative perception, leading to inaccurate decision-making in complex traffic environments and further impacting traffic management efficiency. By combining data from vehicle-side perception units, roadside perception units, and the cloud control platform, a multi-source perception dataset is formed. Features are extracted from this dataset to generate a data feature set for vehicle-road-cloud cooperative perception. Risk assessment is performed based on traffic target state parameters from the vehicle-road situational awareness results, and decision reasoning is based on the risk assessment values to construct vehicle cooperative control commands. This ensures that the vehicle-road cooperative system can make optimal decisions in complex and ever-changing traffic environments, improving traffic management efficiency and road safety.
[0022] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0023] Example 1, please refer to Figure 1. This application provides a vehicle-road-cloud cooperative perception method based on multi-source data fusion. The method is applied to a vehicle-road-cloud cooperative perception system based on multi-source data fusion. The method specifically includes the following steps:
[0024] S100: Heterogeneous sensing is achieved through vehicle-side sensing units, roadside sensing units, and cloud control platforms to form a multi-source sensing dataset.
[0025] Furthermore, S100 of this application includes: acquiring vehicle-side perception data through the vehicle-side perception unit, the vehicle-side perception data including a vehicle motion state dataset and raw sensor data; acquiring roadside perception data through the roadside perception unit, the roadside perception data including traffic participant information and road environment information collected by roadside cameras, millimeter-wave radar and lidar; and acquiring cloud control platform data through the cloud control platform, the cloud control platform data including historical traffic flow data and regional traffic event information.
[0026] Specifically, the vehicle-side perception unit consists of various sensors and devices installed on the vehicle, responsible for collecting real-time information about the vehicle's surrounding environment and its own dynamic state. These include cameras, radar, lidar, GPS, and IMU. The vehicle-side perception data includes a motion state dataset of the vehicle and raw sensor data. The motion state dataset includes the vehicle's speed, acceleration, direction, and position; the raw sensor data includes image data, radar wave reflection data, and environmental information.
[0027] The vehicle-mounted perception unit collects real-time motion data of the vehicle through its onboard sensors. For example, the vehicle's GPS records its position coordinates, and the IMU records its acceleration and steering angle. Onboard cameras, radar, and lidar collect images and depth information of the surrounding environment. For instance, on a highway, the onboard GPS records the vehicle's position as (40.7128°N, 74.0060°W) and its speed as 100 km / h. The onboard radar detects a stationary obstacle 500 meters ahead, possibly another vehicle, while the onboard camera provides real-time image data of that obstacle. The IMU provides the vehicle's acceleration and steering angle information.
[0028] Roadside sensing units refer to sensors and devices deployed on roads to collect information about the surrounding environment. Roadside devices include roadside cameras, millimeter-wave radar, and lidar, used to acquire information about traffic participants and the road environment. Roadside cameras, millimeter-wave radar, and lidar are used to capture road images, detect vehicle speed, identify traffic flow, and detect distances, respectively.
[0029] Roadside sensing units collect surrounding traffic information in real time through roadside equipment. Roadside cameras capture road conditions and traffic flow, millimeter-wave radar detects vehicle speed and lane occupancy, and lidar measures the precise distance to objects ahead. For example, on the same highway section, roadside millimeter-wave radar detects a vehicle speed of 95 km / h in the lane, roadside cameras record a traffic flow of 100 vehicles per minute, and lidar measures the distance to an obstacle at 450 meters.
[0030] The cloud control platform is a centralized data processing and control platform that receives and processes real-time perception data from vehicles and roadside systems, enabling more advanced analysis and decision-making. It provides cross-regional traffic data analysis, generating traffic flow predictions and event monitoring. The platform's data includes historical traffic flow data and regional traffic event information. Historical traffic flow data represents traffic volume, vehicle speed, and lane occupancy within a specific area over a past time period. Regional traffic event information includes events affecting traffic such as traffic accidents, road closures, and weather changes.
[0031] By collaborating with vehicle-mounted and roadside equipment, historical traffic flow data and regional traffic incident information are acquired in real time. Historical data is used to analyze traffic flow trends, while regional incident information helps the platform understand current traffic anomalies, such as accidents or road construction sites. For example, in the past 5 minutes, traffic flow in this area reached 120 vehicles per minute, an increase of 20% compared to the average flow. At the same time, the platform receives a notification from the traffic management center that a traffic accident has occurred on a certain road segment, causing a 50% decrease in traffic flow in that area.
[0032] Through the coordinated operation of vehicle-mounted sensing units, roadside sensing units, and cloud control platforms, comprehensive perception of vehicle movement status, surrounding traffic environment, and historical traffic flow can be achieved.
[0033] Furthermore, this application also includes the following steps: constructing a network time protocol, traversing the vehicle-side perception data, the roadside perception data, and the cloud control platform data according to the network time protocol to perform unified time identification and generate unified timestamp data; constructing a regional map network for the target area, performing coordinate system transformation based on the regional map network to construct a global coordinate system; mapping the vehicle-side perception data to the global coordinate system according to the unified timestamp data, performing target localization on the original sensor data according to the motion state dataset based on the first mapping result to determine the target local coordinates; mapping the roadside perception data to the global coordinate system according to the unified timestamp data, performing target localization on the traffic participant information according to the road environment information based on the second mapping result to determine the target detection coordinates; and spatiotemporally aligning the target local coordinates, the target detection coordinates, the historical traffic flow data, and the regional traffic event information to construct the multi-source perception dataset.
[0034] Specifically, to ensure that all data from vehicle-mounted sensors, roadside sensing units, and the cloud control platform has a unified time signature, the time of each device is synchronized using a network time protocol. The network time protocol ensures clock consistency across devices by synchronizing with a precise time source. Whenever new data is collected, a precise timestamp is added to that data. This timestamp ensures that even if data collection times differ between different devices, they can still be sorted and compared according to a unified time standard. The network time protocol is a protocol used to synchronize computer system time over the internet. By comparing and adjusting the system time with a precise time source, it ensures time synchronization between devices.
[0035] According to the network time protocol, data from vehicle-side sensing, roadside sensing, and cloud control platforms are traversed and uniformly time-stamped to generate unified timestamp data. Unified timestamp data refers to the synchronization of data from different sensors and devices via the network time protocol, enabling data from these heterogeneous data sources to be processed and analyzed within the same time frame. Each data point is marked with a precise timestamp, ensuring temporal consistency and facilitating subsequent data fusion and analysis.
[0036] A regional map network is constructed for the target area, representing its road layout, terrain information, etc. This network is also used to represent traffic facilities, road layout, traffic signs, etc. Based on the regional map network, coordinate system transformation is performed, converting the local coordinates from vehicle-mounted and roadside sensing data into coordinates in the global coordinate system. The global coordinate system is a unified spatial coordinate framework that enables data fusion across devices and regions. It covers the entire target area or a wider region. Under the global coordinate system, all sensor data can be uniformly located, facilitating cross-device and cross-regional data integration and comparison.
[0037] Based on unified timestamp data, vehicle-mounted sensing data is mapped to a global coordinate system. Target localization is performed on the raw sensor data from the vehicle-mounted sensing unit using the motion state dataset. Target localization determines the relative position of the vehicle and surrounding traffic participants, i.e., the target's local coordinates. The target's local coordinates are the relative position coordinates within the target area based on the data collected by the vehicle-mounted sensing unit.
[0038] Using unified timestamp data, roadside perception data is mapped to a global coordinate system. Roadside perception units acquire information about the road environment and traffic participants through sensors such as cameras and radar, and then determine the positions of traffic participants based on the target detection coordinates. Target detection coordinates are the precise location coordinates of traffic participants or obstacles within the target area, determined based on the roadside perception data.
[0039] By spatiotemporally aligning the local coordinates of targets from vehicle-mounted and roadside sensing data with historical traffic flow data and regional traffic event information, data from different times and sources are combined to form a complete multi-source sensing dataset. For example, on a certain road segment, the target coordinates collected by vehicle-mounted and roadside sensors have already been combined with historical traffic flow data through spatiotemporal alignment. Assuming that the historical traffic flow is 120 vehicles per minute during the same time period, vehicle-mounted sensing data records a vehicle slowing down ahead, while roadside sensing data indicates a traffic accident in the area. Through spatiotemporal alignment, this information is synthesized to generate an accurate multi-source sensing dataset.
[0040] By constructing a network time protocol and synchronizing time, data from different data sources can be accurately integrated into a unified time standard. Coordinate system transformation and target positioning can ensure the consistency of geographical location information in each data source, enabling seamless data exchange across devices and regions.
[0041] S200: Based on the multi-source perception dataset, feature extraction is performed to generate a data feature set. Based on the data feature set, vehicle-road-cloud collaborative perception is performed to generate vehicle-road situational perception results. The vehicle-road situational perception results include traffic target state parameters.
[0042] Furthermore, this application also includes the following steps: analyzing the target local coordinates in conjunction with the vehicle-side perception data to construct a vehicle-side feature set; analyzing the target detection coordinates in conjunction with the roadside perception data to construct a roadside feature set; performing traffic flow analysis based on the historical traffic flow data to construct traffic flow distribution characteristics; performing semantic analysis based on regional traffic event information to construct semantic features; constructing a spatiotemporal perception map; performing heterogeneous data fusion reasoning based on the spatiotemporal perception map in conjunction with the vehicle-side feature set, the roadside feature set, the traffic flow distribution characteristics, and the semantic features to generate a global traffic fusion state parameter; performing collaborative perception based on the global traffic fusion state parameter to obtain short-term behavior prediction data; performing multi-layer fusion based on the global traffic fusion state parameter in conjunction with the short-term behavior prediction data to generate traffic target state parameters; and adding the traffic target state parameters to the vehicle-road situational awareness result.
[0043] Specifically, the vehicle-mounted perception unit acquires the local coordinates of the target and combines them with vehicle-side perception data, such as vehicle speed and acceleration, to analyze and extract dynamic features of the vehicle, including vehicle speed, acceleration, and steering angle, forming a vehicle-side feature set. The vehicle-side feature set is a set of features extracted through analysis of vehicle-side perception data, including the vehicle's motion state, environmental state, and the detected target location.
[0044] The roadside sensing unit analyzes target detection coordinates and roadside sensing data to generate a roadside feature set, including traffic density, flow information, and traffic events. The roadside feature set is a set of features extracted from the roadside sensing data, including traffic flow information, the status of traffic participants, and environmental features, used to analyze the overall road traffic situation.
[0045] The cloud-based traffic control platform collects historical traffic flow data, analyzes the data's flow trends and congestion conditions, and generates traffic flow distribution characteristics, which helps determine the current road traffic conditions. For example, suppose historical traffic flow data for a city's roads shows that 1,000 vehicles pass through per hour during the morning rush hour and 1,500 vehicles per hour during the evening rush hour. Traffic flow distribution characteristics refer to the distribution status and flow information of vehicles on the road at a specific moment or within a certain period of time, including vehicle density and flow speed.
[0046] Regional traffic incident information includes data on traffic accidents, construction areas, and traffic control measures. Semantic analysis is performed on this information to generate semantic features. Semantic features are high-level information extracted from roadside perception data that describes the specific meaning of the traffic environment. For example, information such as the type of traffic signs, lane type, and traffic light status helps in understanding traffic rules and dynamic changes within the scenario.
[0047] Spatiotemporal perception maps are graphical structures that combine temporal and spatial information. They are typically used to represent the spatiotemporal evolution of traffic flow and traffic targets, enabling data visualization and analysis across both temporal and spatial dimensions. They reflect the state, behavior, and evolution of various targets within a traffic system. Spatiotemporal perception maps map traffic targets and their dynamic behaviors across time and space. Examples include changes in vehicle position and speed over time, and changes in traffic light states over time.
[0048] Based on spatiotemporal perception maps combined with vehicle-side feature sets, roadside feature sets, traffic flow distribution features, and semantic features, spatiotemporal alignment and fusion inference are performed to generate a comprehensive traffic fusion state parameter, which describes the current traffic situation. The comprehensive traffic fusion state parameter is a comprehensive state parameter obtained by combining multi-source data from vehicles, roadside sources, traffic flow, and semantic features. It describes the current state of the global traffic environment, including traffic flow, the status of traffic participants, and the impact of traffic events.
[0049] Based on the fusion of state parameters across the entire traffic domain, collaborative perception is performed to predict the behavior of traffic participants (such as vehicles and pedestrians) within a short period of time. Short-term behavior prediction data is based on the current traffic state, target behavior, and trends to predict the behavior of traffic participants within a short timeframe, such as vehicle trajectories and pedestrian walking paths. Traffic target state parameters refer to the target state information generated by analyzing the fusion of state parameters across the entire traffic domain and short-term behavior prediction data, including the location, speed, direction, and dynamic state of traffic flow of traffic targets.
[0050] By combining comprehensive traffic fusion state parameters and short-term behavior prediction data, multi-source data fusion is performed to generate traffic target state parameters, which include information such as the current motion state and behavior prediction of traffic targets. Traffic target state parameters are the obtained state parameters for specific traffic targets, including the target's location, speed, and predicted behavior. Through comprehensive analysis of multi-source data such as vehicle-side and roadside perception data, traffic flow, and event information, comprehensive traffic fusion state parameters are generated and combined with short-term behavior prediction for real-time traffic perception. This improves the intelligence level of the traffic management system, enabling accurate prediction and control of traffic targets, thereby improving traffic safety, optimizing traffic flow, and reducing the probability of accidents.
[0051] Furthermore, this application also includes the following steps: the vehicle-side feature set and the roadside feature set are uploaded in real time through a low-latency communication link, and the traffic flow distribution features and the semantic features are periodically issued or pushed on demand by the cloud control platform.
[0052] Specifically, during vehicle operation, the onboard sensing unit continuously collects vehicle-side feature data, including the vehicle's dynamic status and environmental characteristics. Simultaneously, the roadside sensing unit also collects feature data about the surrounding environment, including road conditions and the distribution of traffic participants. This feature data is uploaded in real-time to the cloud control platform or other management systems via a low-latency communication link. Based on the collected data, the cloud control platform generates traffic flow distribution features and semantic features, periodically sending them to the vehicle-side and roadside sensing units as needed, and also pushing them to relevant systems on demand during emergencies. For example, assuming that during peak hours, the cloud control platform sends traffic flow distribution feature data every 5 minutes, including the number of vehicles passing through a major road per hour and the current traffic density; simultaneously, when the cloud control platform detects a traffic accident on a road segment, the platform can push the latest semantic features (such as traffic signs and signal light status) to the vehicle-side devices as needed, helping the vehicle dynamically adjust around the accident scene.
[0053] The cloud control platform can receive real-time sensing data from vehicles and roadside devices, and send real-time or periodically updated traffic information to relevant equipment. Low-latency communication links ensure real-time data transmission, thereby ensuring that the system can respond promptly in rapidly changing traffic environments.
[0054] Furthermore, this application also includes the following steps: parsing the raw sensor data to obtain vehicle-side image data; performing convolution calculations on the vehicle-side image data to extract pixel-level semantic features and target-level visual features; performing three-dimensional scene analysis based on the target's local coordinates to extract local geometric structure features and global scene features; analyzing the motion state dataset through a long short-term memory network to extract vehicle dynamic features; and concatenating the pixel-level semantic features, target-level visual features, local geometric structure features, global scene features, and vehicle dynamic features to construct the vehicle-side feature set.
[0055] Specifically, after preprocessing the raw video streams or image data captured by the cameras of the vehicle-mounted perception unit, the vehicle-mounted image data is extracted, which includes surrounding traffic participants and various road signs and facilities. The vehicle-mounted image data is image or video data captured by the onboard cameras, containing visual information about the vehicle's surroundings.
[0056] By processing vehicle-side image data using convolutional neural networks, pixel-level semantic features are first extracted, such as road surface, pedestrians, and lane markings. Then, target-level visual features are extracted to identify and locate specific objects in the image, such as other vehicles and pedestrians ahead. Assuming that pixel-level semantic features are extracted from the image through convolutional layers, the image might show areas marking lane markings, pedestrian positions, and the shapes of vehicles ahead. The convolutional neural network successfully extracts these key visual features through the computation of multiple convolutional layers.
[0057] Convolution is a mathematical operation, and in computer vision, convolutional neural networks are commonly used to process image data. Convolutional layers scan image data through multiple filters to extract different features from the image. Pixel-level information is extracted progressively so that the system can recognize complex objects or scenes. Pixel-level semantic features refer to the category information of each pixel extracted from the image. For example, in a traffic scene, one pixel might represent lane markings, while another might represent a pedestrian; this information is crucial for understanding the semantics of the image. Object-level visual features refer to the overall features of the target extracted from the image; these features typically represent the object's shape, position, size, and motion state.
[0058] During vehicle movement, data from different perspectives collected by onboard sensors is used for 3D scene analysis to extract local geometric features and global features of the entire scene. 3D scene analysis refers to the process of transforming data from a two-dimensional plane into a three-dimensional space by combining data from different sensors. Global scene features refer to the overall structural information of the entire scene, such as traffic signs, road types, and lane markings, which helps in understanding the macroscopic layout of the traffic environment.
[0059] Long Short-Term Memory (LSTM) networks are used to perform time-series analysis on vehicle motion state data to extract dynamic features such as acceleration, steering angle, and speed. For example, assuming a vehicle's acceleration is 0.2 m / s² and its speed is 60 km / h, LSTM analysis can extract the vehicle's current dynamic features, such as its potential acceleration or deceleration within the next 5 seconds and its predicted trajectory within the current lane. LSTM networks are a special type of recurrent neural network capable of processing and predicting long-term dependent data based on time series, making them suitable for analyzing time-correlated dynamic data, such as vehicle motion states. Vehicle dynamic features refer to the characteristics of vehicle motion extracted through analysis of vehicle motion state datasets, including the vehicle's direction of motion and speed changes at specific moments.
[0060] The pixel-level semantic features and target-level visual features extracted from image data, the local geometric structure features and global scene features obtained through 3D scene analysis, and the vehicle dynamic features are spliced together to form a complete vehicle-side feature set, which contains comprehensive perception information of the vehicle-side perception system on the current traffic environment, surrounding targets, and vehicle status.
[0061] Vehicle-mounted perception systems can extract rich features from different data sources, including the vehicle's current dynamic state, the geometric information of the surrounding environment, and the state of other traffic participants. This improves the accuracy of vehicle-mounted perception systems, helps to accurately identify objects in traffic scenes, understand complex environmental information, and predict future vehicle behavior.
[0062] Furthermore, this application also includes the following steps: analyzing the road environment information to obtain roadside wide-angle image data; performing panoramic segmentation on the roadside wide-angle image data to extract pixel-level semantic segmentation features and bounding box features; tracking the target detection coordinates based on the traffic participant information to obtain target motion features and target interaction features; and fusing the pixel-level semantic segmentation features, the bounding box features, the target motion features, and the target interaction features to generate the roadside feature set.
[0063] Specifically, roadside sensing units capture various areas of the road by collecting wide-angle image data in real time, including various information about the road environment such as traffic signs, traffic lights, road conditions, other vehicles, and pedestrians. The roadside wide-angle image data is collected by cameras installed on the roadside; these cameras typically have a wide field of view and can capture a large area of the traffic environment on the road.
[0064] Panoramic segmentation algorithms are used to process acquired wide-angle roadside images, classifying them pixel-wise. For example, each pixel in the image is labeled as a road, pedestrian, or vehicle. Simultaneously, bounding box features are extracted for each target (e.g., vehicle, pedestrian) to indicate its location and size. Panoramic segmentation is a technique in computer vision used to classify and label different regions in an image. It typically analyzes images at the pixel level, assigning a category label to each pixel to help identify specific objects and their locations within the scene. Pixel-level semantic segmentation features refer to the category information represented by each pixel obtained through image segmentation. Bounding box features are rectangles drawn around an object in the image, determining the object's location and size.
[0065] Traffic participant information refers to information collected by sensing units about various traffic participants on the road, including their current location, speed, and behavior. Trajectory tracking is performed based on target detection coordinates extracted from images, that is, tracking the target's motion trajectory through continuous frame data. By analyzing the trajectory, the target's motion characteristics and interaction characteristics are obtained. Target motion characteristics refer to the characteristic information about the target's motion obtained through the analysis of the target trajectory, including the target's speed, acceleration, and direction of travel. Target interaction characteristics refer to the interactions or relationships between targets in a multi-target scenario. For example, a vehicle may interact with the vehicle in front, such as decelerating or changing lanes; these interaction characteristics are crucial for traffic prediction and decision-making.
[0066] By fusing pixel-level semantic segmentation features, bounding box features, target motion features, and target interaction features, a comprehensive roadside feature set is generated, fully reflecting the roadside environment, the state of traffic participants, and the interaction information between them. The roadside feature set refers to a complete dataset generated through the analysis of roadside perception data, combining multiple features such as pixel-level semantic segmentation, target detection, target motion, and interaction, encompassing a comprehensive understanding of the road environment and traffic participants.
[0067] Roadside perception units can accurately identify, track, and analyze traffic participants in the road environment, extracting rich visual features from wide-angle image data, such as pixel-level semantic information, target bounding boxes, target motion states, and interactive behaviors. They can detect potential traffic risks in a timely manner and predict future traffic flow and events through target interaction features, thereby improving road safety and traffic flow efficiency.
[0068] S300: Perform risk assessment according to the traffic target state parameters, make decision-making inferences based on the risk assessment values, and construct vehicle cooperative control and cooperative perception commands.
[0069] Specifically, after acquiring the state parameters of traffic targets, an assessment is performed based on a predefined risk model. A risk value is calculated for each traffic target's state. For example, if a vehicle is traveling at excessive speed and approaches other vehicles, a high collision risk is assessed. If traffic congestion occurs at an intersection, a high congestion risk is assessed. For instance, suppose the system acquires the following traffic target state parameters: vehicle A's speed is 80 km / h, and its distance from vehicle B ahead is 5 meters. Based on this data and a collision prediction algorithm, the system will assess a collision risk value of 0.85, indicating a high collision risk. Simultaneously, roadside traffic flow data is acquired, showing severe traffic congestion at the intersection, and a congestion risk value of 0.7 is assessed.
[0070] Once the risk assessment is complete, the next step is decision reasoning. Various risk assessment values will be integrated, and appropriate decision instructions will be generated based on factors such as real-time traffic flow, target behavior predictions, and road conditions. For example, when a high collision risk is assessed, instructions to slow down or stop will be generated; when traffic congestion is assessed, it will be suggested that the vehicle change its route or adjust its speed. For instance, if a traffic target ahead (vehicle A) faces a high collision risk (value 0.85), and there is heavy traffic congestion in the area, the following decision will be generated through reasoning: Vehicle A needs to slow down to 40 km / h, maintain a safe distance from the vehicle ahead, and, to reduce congestion, the driver is advised to choose an alternative route.
[0071] After decision-making and reasoning, the generated collaborative perception commands guide vehicles to work collaboratively with surrounding traffic targets. The core objective of vehicle cooperative control is to achieve coordination between vehicles, such as collaborative lane changing, collaborative acceleration, and collaborative braking. Commands are issued to vehicles to guide them on how to interact with the surrounding traffic environment, avoid accidents, and optimize traffic flow. Collaborative perception commands are instructions generated based on multi-source data fusion and collaborative analysis, used to guide the collaborative operation between vehicles and traffic facilities, helping vehicles work collaboratively with other traffic targets and traffic infrastructure to achieve intelligent traffic management.
[0072] Vehicle status and traffic environment changes are monitored in real time through onboard sensors, roadside perception data, and a cloud control platform. Based on this real-time feedback data, vehicle cooperative perception commands are continuously optimized to ensure that vehicle control behavior matches the traffic environment. Within the framework of vehicle-road cooperation, the driving behavior of multiple vehicles is controlled through mutual coordination. By coordinating with surrounding vehicles and traffic facilities (such as traffic lights and roadside sensors), traffic flow is optimized, traffic conflicts are avoided, and traffic safety is ensured.
[0073] Based on the assessment of traffic target state parameters and combined with collaborative perception commands based on real-time feedback, vehicles are effectively guided to drive safely and smoothly in complex traffic environments. Especially in high-risk scenarios, collaborative control commands can help vehicles avoid collisions, alleviate traffic congestion, and improve the efficiency of overall traffic flow, significantly improving the level of intelligence in traffic management, reducing the probability of traffic accidents, and optimizing traffic flow.
[0074] In summary, the vehicle-road-cloud cooperative perception method based on multi-source data fusion provided in this application has the following technical effects:
[0075] Heterogeneous sensing is achieved through vehicle-mounted sensing units, roadside sensing units, and a cloud control platform, forming a multi-source sensing dataset. Feature extraction is performed on this dataset to generate a data feature set. Vehicle-road-cloud collaborative sensing is then performed based on this feature set to generate vehicle-road situational awareness results, which include traffic target state parameters. Risk assessment is conducted according to these traffic target state parameters, and decision reasoning is performed based on the risk assessment values to construct vehicle cooperative control and collaborative sensing commands. In other words, by combining data from vehicle-mounted sensing units, roadside sensing units, and the cloud control platform to form a multi-source sensing dataset, performing feature extraction on this dataset to generate a data feature set for vehicle-road-cloud collaborative sensing, conducting risk assessment based on the traffic target state parameters in the vehicle-road situational awareness results, and performing decision reasoning based on the risk assessment values to construct vehicle cooperative control commands, the vehicle-road cooperative system ensures that it can make optimal decisions in complex and ever-changing traffic environments, improving traffic management efficiency and road safety.
[0076] Example 2: Based on the same inventive concept as the vehicle-road-cloud cooperative perception method based on multi-source data fusion in Example 1, this application also provides a vehicle-road-cloud cooperative perception system based on multi-source data fusion. Please refer to Figure 2. The vehicle-road-cloud cooperative perception system based on multi-source data fusion includes:
[0077] The heterogeneous perception module 11 is used to perform heterogeneous perception through vehicle-side perception units, roadside perception units, and cloud control platforms to form a multi-source perception dataset; the vehicle-road-cloud collaborative perception module 12 is used to extract features based on the multi-source perception dataset, generate a data feature set, perform vehicle-road-cloud collaborative perception based on the data feature set, and generate vehicle-road situational perception results, which include traffic target state parameters; the risk assessment module 13 is used to perform risk assessment according to the traffic target state parameters, perform decision reasoning based on the risk assessment value, and construct vehicle collaborative control collaborative perception instructions.
[0078] Furthermore, the heterogeneous perception module 11 in the vehicle-road-cloud cooperative perception system based on multi-source data fusion is also used to: collect vehicle-end perception data through the vehicle-end perception unit, the vehicle-end perception data including vehicle motion state dataset and raw sensor data; collect roadside perception data through the roadside perception unit, the roadside perception data including traffic participant information and road environment information collected by roadside cameras, millimeter-wave radar and lidar; and collect cloud control platform data through the cloud control platform, the cloud control platform data including historical traffic flow data and regional traffic event information.
[0079] Furthermore, the heterogeneous sensing module 11 in the vehicle-road-cloud cooperative perception system based on multi-source data fusion is also used for: constructing a network time protocol, traversing the vehicle-end sensing data, the roadside sensing data, and the cloud control platform data according to the network time protocol to perform unified time identification and generate unified timestamp data; constructing a regional map network of the target area, performing coordinate system transformation based on the regional map network to construct a global coordinate system; mapping the vehicle-end sensing data to the global coordinate system according to the unified timestamp data, performing target local positioning on the original sensor data according to the motion state dataset based on the first mapping result, and determining the target local coordinates; mapping the roadside sensing data to the global coordinate system according to the unified timestamp data, performing target local positioning on the traffic participant information according to the road environment information based on the second mapping result, and determining the target detection coordinates; and spatiotemporally aligning the target local coordinates, the target detection coordinates, the historical traffic flow data, and the regional traffic event information to construct the multi-source sensing dataset.
[0080] Furthermore, the vehicle-road-cloud collaborative perception module 12 in the vehicle-road-cloud collaborative perception system based on multi-source data fusion is also used for: analyzing the target local coordinates in conjunction with the vehicle-end perception data to construct a vehicle-end feature set; analyzing the target detection coordinates in conjunction with the roadside perception data to construct a roadside feature set; performing traffic flow analysis based on the historical traffic flow data to construct traffic flow distribution characteristics; performing semantic analysis based on regional traffic event information to construct semantic features; constructing a spatiotemporal perception map; performing heterogeneous data fusion reasoning based on the spatiotemporal perception map in conjunction with the vehicle-end feature set, the roadside feature set, the traffic flow distribution characteristics, and the semantic features to generate full-domain traffic fusion state parameters; performing collaborative perception based on the full-domain traffic fusion state parameters to obtain short-term behavior prediction data; performing multi-layer fusion based on the full-domain traffic fusion state parameters in conjunction with the short-term behavior prediction data to generate traffic target state parameters, and adding the traffic target state parameters to the vehicle-road situational perception result.
[0081] Furthermore, the vehicle-road-cloud collaborative perception module 12 in the vehicle-road-cloud collaborative perception system based on multi-source data fusion is also used for: the vehicle-side feature set and the roadside feature set are uploaded in real time through a low-latency communication link, and the traffic flow distribution features and the semantic features are periodically issued or pushed on demand by the cloud control platform.
[0082] Furthermore, the vehicle-road-cloud collaborative perception module 12 in the vehicle-road-cloud collaborative perception system based on multi-source data fusion is also used for: parsing the original sensor data to obtain vehicle-end image data; performing convolution calculations on the vehicle-end image data to extract pixel-level semantic features and target-level visual features; performing three-dimensional scene analysis based on the target's local coordinates to extract local geometric structure features and global scene features; analyzing the motion state dataset through a long short-term memory network to extract vehicle dynamic features; and concatenating the pixel-level semantic features, target-level visual features, local geometric structure features, global scene features, and vehicle dynamic features to construct the vehicle-end feature set.
[0083] Furthermore, the vehicle-road-cloud collaborative perception module 12 in the vehicle-road-cloud collaborative perception system based on multi-source data fusion is also used for: analyzing the road environment information to obtain roadside wide-angle image data; performing panoramic segmentation on the roadside wide-angle image data to extract pixel-level semantic segmentation features and bounding box features; tracking the target detection coordinates based on the traffic participant information to obtain target motion features and target interaction features; and fusing the pixel-level semantic segmentation features, the bounding box features, the target motion features, and the target interaction features to generate the roadside feature set.
[0084] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The vehicle-road-cloud cooperative perception method and specific examples based on multi-source data fusion in the foregoing embodiment one are also applicable to the vehicle-road-cloud cooperative perception system based on multi-source data fusion in this embodiment. Through the foregoing detailed description of the vehicle-road-cloud cooperative perception method based on multi-source data fusion, those skilled in the art can clearly understand the vehicle-road-cloud cooperative perception system based on multi-source data fusion in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.
[0085] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0086] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A vehicle-road-cloud cooperative perception method based on multi-source data fusion, characterized in that, include: Heterogeneous sensing is achieved through vehicle-side sensing units, roadside sensing units, and cloud control platforms to form a multi-source sensing dataset. Feature extraction is performed based on the multi-source perception dataset to generate a data feature set. Vehicle-road-cloud collaborative perception is then performed based on the data feature set to generate vehicle-road situational awareness results, which include traffic target state parameters. Risk assessment is then performed based on the traffic target state parameters, and decision reasoning is conducted based on the risk assessment values to construct vehicle collaborative control and collaborative perception commands.
2. The vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in claim 1, characterized in that, Vehicle-side perception data is acquired through the vehicle-side perception unit, which includes vehicle motion state datasets and raw sensor data. Roadside perception data is acquired through the roadside perception unit, which includes traffic participant information and road environment information collected by roadside cameras, millimeter-wave radar, and lidar. Cloud control platform data is acquired through the cloud control platform, which includes historical traffic flow data and regional traffic event information.
3. The vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in claim 2, characterized in that, Heterogeneous sensing is achieved through vehicle-mounted sensing units, roadside sensing units, and a cloud control platform to form a multi-source sensing dataset. The method includes: constructing a network time protocol; traversing the vehicle-mounted sensing data, roadside sensing data, and cloud control platform data according to the network time protocol to generate unified timestamp data; constructing a regional map network for the target area; performing coordinate system transformation based on the regional map network to construct a global coordinate system; mapping the vehicle-mounted sensing data to the global coordinate system according to the unified timestamp data; performing target localization on the original sensor data according to the motion state dataset based on the first mapping result to determine the target's local coordinates; mapping the roadside sensing data to the global coordinate system according to the unified timestamp data; performing target localization on the traffic participant information according to the road environment information based on the second mapping result to determine the target detection coordinates; and spatiotemporally aligning the target's local coordinates, the target detection coordinates, historical traffic flow data, and regional traffic event information to construct the multi-source sensing dataset.
4. The vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in claim 3, characterized in that, Based on the multi-source perception dataset, feature extraction is performed to generate a data feature set. Vehicle-road-cloud collaborative perception is then performed based on the data feature set to generate vehicle-road situational awareness results. The method includes: analyzing the target's local coordinates in conjunction with vehicle-side perception data to construct a vehicle-side feature set; analyzing the target's detection coordinates in conjunction with roadside perception data to construct a roadside feature set; performing traffic flow analysis based on historical traffic flow data to construct traffic flow distribution characteristics; performing semantic analysis based on regional traffic event information to construct semantic features; constructing a spatiotemporal perception map; performing heterogeneous data fusion inference based on the spatiotemporal perception map in conjunction with the vehicle-side feature set, the roadside feature set, the traffic flow distribution characteristics, and the semantic features to generate a global traffic fusion state parameter; performing collaborative perception based on the global traffic fusion state parameter to obtain short-term behavior prediction data; and performing multi-layer fusion based on the global traffic fusion state parameter and the short-term behavior prediction data to generate traffic target state parameters, which are then added to the vehicle-road situational awareness results.
5. The vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in claim 4, characterized in that, The vehicle-side feature set and the roadside feature set are uploaded in real time through a low-latency communication link, and the traffic flow distribution features and the semantic features are periodically distributed or pushed on demand by the cloud control platform.
6. The vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in claim 4, characterized in that, The method for constructing a vehicle-side feature set based on the target's local coordinates and the vehicle-side perception data includes: parsing the original sensor data to obtain vehicle-side image data; performing convolution calculations on the vehicle-side image data to extract pixel-level semantic features and target-level visual features; performing 3D scene analysis based on the target's local coordinates to extract local geometric structure features and global scene features; analyzing the motion state dataset through a long short-term memory network to extract vehicle dynamic features; and concatenating the pixel-level semantic features, target-level visual features, local geometric structure features, global scene features, and vehicle dynamic features to construct the vehicle-side feature set.
7. The vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in claim 4, characterized in that, The method for constructing a roadside feature set based on the target detection coordinates and the roadside perception data includes: parsing the road environment information to obtain roadside wide-angle image data; performing panoramic segmentation on the roadside wide-angle image data to extract pixel-level semantic segmentation features and bounding box features; tracking the target detection coordinates based on the traffic participant information to obtain target motion features and target interaction features; and fusing the pixel-level semantic segmentation features, the bounding box features, the target motion features, and the target interaction features to generate the roadside feature set.
8. A vehicle-road-cloud cooperative perception system based on multi-source data fusion, characterized in that, The steps for implementing the vehicle-road-cloud cooperative perception method based on multi-source data fusion as described in any one of claims 1 to 7, wherein the vehicle-road-cloud cooperative perception system based on multi-source data fusion comprises: a heterogeneous perception module, used for heterogeneous perception through vehicle-side perception units, roadside perception units, and cloud control platforms to form a multi-source perception dataset; a vehicle-road-cloud cooperative perception module, used for feature extraction based on the multi-source perception dataset to generate a data feature set, and for performing vehicle-road-cloud cooperative perception based on the data feature set to generate a vehicle-road situational perception result, wherein the vehicle-road situational perception result includes traffic target state parameters; and a risk assessment module, used for risk assessment according to the traffic target state parameters, and for decision reasoning based on the risk assessment value to construct vehicle cooperative control cooperative perception instructions.