Highway vehicle-road cloud multi-modal data streaming acquisition and edge-cloud fusion method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-13
- Publication Date
- 2026-08-11
AI Technical Summary
[0008]为了解决上述技术问题,本发明提供高速公路车路云多模态数据流式采集与边云融合方法,解决了现有高速公路车路云系统中多模态流式数据时空同步精度低、复杂场景融合鲁棒性差、边云数据分流无法兼顾带宽时延与感知精度、云端全局融合能力弱且无闭环优化的技术问题
时空同步精度显著提升,本发明采用自创动态时延预测与时间戳修正公式,结合北斗全局时空基准进行统一校准,将多模态数据平均时间同步误差由现有技术的15ms以上降至2.3ms以内,平均空间对齐误差由40cm以上降至12.5cm以内,从根源上避免高速移动场景下目标检测与轨迹跟踪错位,大幅提升数据基础可靠性。
Smart Images

Figure CN122551580A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent transportation and vehicle-road cooperative technology, and more specifically, it relates to a method for multimodal data streaming acquisition and edge-cloud fusion of highway vehicle-road-cloud systems. Background Technology
[0002] With the rapid advancement of vehicle-road cooperation and the construction of smart highways, multimodal perception and data fusion based on the integrated vehicle-road-cloud architecture have become core technologies for accurate traffic situation identification and autonomous driving safety assurance. Highway scenarios are characterized by high vehicle speeds, complex environments, high accident risks, and large data volumes. The streaming acquisition, real-time synchronization, and edge-cloud collaborative processing of multi-source heterogeneous data from sources such as vision, LiDAR, millimeter-wave radar, vehicle bus, and meteorological sensors directly determine the reliability of traffic perception and control.
[0003] Existing highway vehicle-road-cloud data processing technologies generally suffer from the following technical deficiencies: Multimodal streaming data suffers from insufficient spatiotemporal synchronization accuracy. Various sensor hardware trigger delays and network transmission delays exhibit random fluctuations. Traditional fixed timestamp alignment methods do not consider dynamic delay prediction and compensation, which can easily lead to target spatial misalignment in high-speed moving scenarios. Time synchronization errors typically exceed 15ms, and spatial alignment errors are greater than 40cm, making it difficult to meet the requirements of high-precision sensing.
[0004] Multimodal data fusion at the edge has poor robustness. Existing methods mostly use equal weight or fixed weight fusion, without dynamically allocating modal confidence based on meteorological and traffic flow scenarios. In harsh environments such as rain, fog, and snow, the performance of visual and other sensors drops sharply, and the fusion perception accuracy is significantly reduced. The target detection accuracy in foggy weather is generally below 70%, making it unsuitable for complex highway scenarios.
[0005] The edge-cloud data offloading system lacks a globally optimal decision-making mechanism. Existing solutions either upload all raw data to the cloud, resulting in high bandwidth consumption and end-to-end latency exceeding 200ms; or process all data at the edge, causing a loss of over 5% in global perception accuracy, making it difficult to achieve a dynamic balance between bandwidth, latency, and perception accuracy.
[0006] The cloud-based global fusion capability is weak and lacks a closed-loop optimization mechanism. The cloud only achieves simple data aggregation and does not carry out temporal and semantic dual-dimensional correlation fusion, resulting in low accuracy of the overall road network situational awareness. At the same time, it cannot adjust the fusion weights and traffic splitting strategies of edge nodes based on global results, and the system's scenario adaptability and long-term stability are insufficient.
[0007] Currently, the relevant technologies have not yet formed a complete solution that takes into account streaming acquisition, dynamic spatiotemporal synchronization, adaptive edge-cloud diversion, global dual-dimensional fusion and closed-loop optimization. In particular, there is a lack of original quantitative fusion and decision formulas for highway scenarios, which makes it difficult to break through the performance bottleneck of intelligent highway vehicle-road-cloud systems. Summary of the Invention
[0008] To address the aforementioned technical issues, this invention provides a method for multimodal data streaming acquisition and edge-cloud fusion of vehicle-road-cloud systems on highways. This method solves the technical problems of low spatiotemporal synchronization accuracy of multimodal streaming data in existing highway vehicle-road-cloud systems, poor robustness of fusion in complex scenarios, inability of edge-cloud data diversion to balance bandwidth latency and perception accuracy, and weak global fusion capability in the cloud without closed-loop optimization.
[0009] The method for multimodal data streaming acquisition and edge-cloud fusion of highway vehicle-road-cloud data includes the following steps: S1. Construction of Global Spatiotemporal Reference and Acquisition of Multimodal Streaming Data: Based on BeiDou / GPS timing, a unified global spatiotemporal reference at the vehicle-road-cloud level is constructed, and multimodal streaming data in highway scenarios is collected synchronously. The multimodal streaming data includes vehicle-mounted vision, millimeter-wave radar, lidar, and CAN bus data; roadside unit (RSU), vision sensor, millimeter-wave radar, lidar, and meteorological sensor data; and cloud-based road network basic data and traffic control data. S2. Dynamic spatiotemporal synchronization and alignment of multimodal streaming data: The collected multimodal streaming data is segmented and the timestamp is corrected and the spatial coordinates are aligned for each data segment to obtain synchronized multimodal data under the global spatiotemporal reference. Specifically, for the k-th streaming data slice of the i-th type of multimodal sensing data source, its standard timestamp is corrected through dynamic latency prediction, and the correction formula is as follows: In the formula, This is the original acquisition timestamp of the k-th segment. For the inherent hardware trigger latency of the i-th type of data source, The current fragment transmission latency is predicted based on the exponentially weighted moving average. The initial clock offset between the i-th type of data source and the global spatiotemporal reference; Simultaneously, the target detection coordinates of each modal data are uniformly transformed to the global road network coordinate system through the external parameter calibration matrix; S3, Lightweight Multimodal Feature Fusion on the Roadside: At the roadside edge computing node, feature extraction of the synchronized multimodal data is performed in a unified dimension, the fusion confidence of each modality feature is calculated, weighted feature fusion is completed based on the fusion confidence, and the edge-side fusion features and local traffic situation perception results are output. The formula for calculating the fusion confidence of the i-th modality feature is as follows: In the formula, M represents the total number of modes participating in the fusion. Let be the signal-to-noise ratio of the i-th modal feature. Assign scene adaptation weights to the i-th modality, and satisfy the following conditions: =1; The final side-to-side fusion feature is , Let i be the feature vector of the i-th mode; S4. Edge-cloud adaptive traffic splitting decision based on optimal cost function: Construct the total cost function of edge-cloud traffic splitting decision, solve the optimal traffic splitting decision sequence with the goal of minimizing the total cost, and determine the processing path of each data shard. The processing path is either local processing on the edge or uploading to the cloud for processing. The total cost function formula is as follows: In the formula, N is the total number of data fragments to be decided. ∈{0,1} are the flow splitting decision variables. =1 indicates that the data is uploaded to the cloud in chunks. =0 indicates local processing on the edge. As the bandwidth cost factor, As a delay cost factor, Let be the accuracy loss cost factor, and be the weight coefficients of the three types of cost factors, respectively, and satisfy . + + =1; S5. Cloud-side global temporal-semantic dual-dimensional fusion: Global association matching is performed on the fragmented data received from the cloud and the fusion features from the edge side. The temporal-semantic dual-dimensional global fusion is completed through iterative weight updates to construct a global map of the traffic situation of the entire road network. The formula for iteratively updating the weights in the global fusion is as follows: ; In the formula, The fusion weights are the features of the i-th edge node in the k-th iteration. Let be the semantic similarity between the i-th side feature and the global fused feature. Let represent the spatiotemporal reliability of the i-th edge node; After iterating until the weights converge, the final global fusion feature is output. S6. Edge-Cloud Fusion Closed-Loop Optimization: Based on the global fusion results and global situation map in the cloud, the fusion confidence scenario adaptation weight and the traffic diversion decision cost function weight coefficient of the edge nodes are updated in reverse to achieve closed-loop dynamic optimization of the entire vehicle-road-cloud fusion process.
[0010] Preferably, in step S2, the transmission delay prediction value The formula for calculating the exponentially weighted moving average is: In the formula, α is the smoothing coefficient, and its value ranges from 0 to 1. The actual transmission delay of the (k-1)th fragment, initial value = ; The smoothing coefficient α is adaptively adjusted according to the degree of link delay fluctuation; the larger the variance of delay fluctuation, the larger the value of α.
[0011] Preferably, in step S3, the scene adaptation weight The calculation is dynamically calculated based on the current highway scene classification results. The calculation formula is as follows: ; In the formula, The adaptation probability of the i-th modality output by the scene classifier in the current scene, wherein the current scene includes weather scenes (sunny / rainy / foggy / snowy) and traffic flow scenes (smooth / slow / congested / accident).
[0012] Preferably, in step S3, the signal-to-noise ratio of the i-th modal feature The calculation formula is: ; In the formula, The effective signal power of the i-th mode eigenvector is obtained by calculating the eigenvariance. The noise power of the i-th modal eigenvector is calculated using the sum of squared feature residuals.
[0013] Preferably, in step S4, the bandwidth cost factor Delay cost factor Accuracy loss cost factor The calculation formulas are as follows: ; ; ; In the formula, Let i be the size of the i-th data segment. This represents the maximum available bandwidth for the current link. To reduce upload latency in chunks, To handle latency in the cloud, To account for edge processing delay, To improve the accuracy of edge-side fusion sensing, To improve the accuracy of global fusion perception in the cloud.
[0014] Preferably, in step S4, the weighting coefficients , , Based on the current highway scenario, the system dynamically and adaptively adjusts the settings according to the following rules: In congestion / accident scenarios, the weight of λD is reduced and the weight of λL is increased to prioritize perception accuracy and real-time decision-making. In low-bandwidth / weak network scenarios, increase the weight of λB and decrease the weight of λD to prioritize reducing link bandwidth usage; In scenarios with high computing power consumption on the edge, the weight of λD is increased to prioritize processing in the cloud.
[0015] Preferably, in step S5, semantic similarity With spatiotemporal credibility The calculation formulas are as follows: ; ; In the formula, Let be the fusion feature of the i-th edge node at time t. For the (k-1)th iteration, the global fusion feature is... Let be the spatial distance between the i-th side node and the perceived target. The time synchronization deviation between the data of the i-th side node and the global reference is... This is the distance normalization coefficient. This is the time deviation normalization coefficient.
[0016] Preferably, in step S5, the termination condition for iteration convergence is: <ε, where ε is the convergence threshold, and its value ranges from 10⁻⁶ ≤ ε ≤ 10⁻³. The convergence threshold is adaptively adjusted according to the cloud computing load.
[0017] Preferably, before step S2, a multimodal streaming data preprocessing step is included to remove outliers, denoise and compress the raw collected streaming data. Among them, outlier removal adopts the 3σ criterion based on sliding window, which removes and interpolates outlier data that exceeds the mean ± 3 times the standard deviation within the window. At the same time, high-dimensional point cloud and image data are subjected to lightweight compression based on region of interest, and the compression ratio is dynamically adjusted according to the side computing power and link bandwidth.
[0018] Preferably, in step S4, a data priority factor is added to the diversion decision, and a mandatory upload rule is set for high-priority data such as accidents, abnormal events, and traffic violations, forcibly setting their diversion decision variable to 1. At the same time, the weight of the latency cost factor for high-priority data is doubled in the total cost function to ensure that the end-to-end transmission and processing latency of high-priority data is minimized.
[0019] Compared with the prior art, the present invention has the following beneficial effects: The spatiotemporal synchronization accuracy is significantly improved. This invention adopts a self-created dynamic delay prediction and timestamp correction formula, combined with the BeiDou global spatiotemporal reference for unified calibration. The average time synchronization error of multimodal data is reduced from more than 15ms in the existing technology to less than 2.3ms, and the average spatial alignment error is reduced from more than 40cm to less than 12.5cm. This fundamentally avoids misalignment of target detection and trajectory tracking in high-speed moving scenarios and greatly improves the reliability of the data foundation.
[0020] The robustness of complex scene fusion is greatly enhanced. By creating a self-developed fusion confidence formula and dynamically allocating the fusion weights of each modality in combination with the signal-to-noise ratio and scene adaptation weights, the perception strategy can be adaptively adjusted according to scenarios such as sunny, rainy, foggy, congested, and accidental events. In severe environments such as foggy days, the average accuracy of target detection has increased from about 60% to over 92%, and the accuracy decay rate has decreased from over 30% to less than 6%, significantly improving the system's ability to adapt to all scenes.
[0021] The edge-cloud collaboration efficiency achieves global optimization. Based on the self-created edge-cloud traffic offloading optimal cost function, it makes adaptive decisions by comprehensively balancing three major indicators: bandwidth, latency, and accuracy loss. Compared with the full upload solution, the average uplink bandwidth utilization rate is reduced from 100% to below 33%, the end-to-end processing latency is reduced from more than 200ms to less than 90ms, and the perception accuracy loss is controlled within 1.2%, achieving the optimal balance between resource efficiency, real-time performance, and perception quality.
[0022] With high global perception accuracy and closed-loop optimization capabilities, the cloud adopts a time-series-semantic dual-dimensional weight iteration formula to achieve global feature fusion across roadside nodes. The target detection accuracy of the entire road segment exceeds 99%, and the abnormal event recall rate exceeds 99.5%. At the same time, the cloud global results are used to back-optimize the edge-side fusion weight, traffic diversion cost coefficient, and perception model, forming an edge-cloud collaborative closed loop. The system can adapt to scene changes over a long period of time and maintain high-performance and stable operation. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the present invention. Detailed Implementation
[0024] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.
[0025] System overall architecture and hardware / software infrastructure: This invention adopts a three-level collaborative architecture of vehicle-road-cloud, and all hardware and software parameters are clearly defined and reproducible, with no ambiguity. Hardware deployment and configuration: Vehicle-mounted components: 100 test vehicles (50 commercial vehicles and 50 passenger vehicles), each equipped with: a Beidou-3 high-precision positioning / timing module (timing accuracy ±10ns, positioning accuracy ±2cm), an 8-megapixel forward vision sensor (30fps), a 192-channel 4D millimeter-wave radar (10Hz), a 16-line lidar (10Hz), a CAN bus acquisition module (50Hz), an onboard computing unit (Jetson Xavier NX, 21TOPSINT8), and a 5G-V2X communication module.
[0026] Roadside edge: A 10km test section of the G42 Shanghai-Chengdu Expressway in Jiangsu Province was selected, and 20 roadside edge nodes were deployed at 500m intervals. Each node was equipped with: a Beidou-3 timing module (±10ns), 4 8-megapixel bullet cameras (covering the entire lane, 30fps), 2 4D millimeter-wave radars (10Hz), a 32-line lidar (10Hz), a weather sensor (collecting rainfall, visibility, and wind speed, 1Hz), an edge computing server (2 Intel Xeon D-1700 processors, 64GB of memory, NVIDIA A2 accelerator card, 40 TOPS INT8 computing power), and a fiber optic / 5G backhaul link (uplink available bandwidth up to 100Mbps).
[0027] Cloud-based: Deployed on the provincial-level traffic cloud control platform, configured with: a distributed computing cluster (32 nodes, each node with 32 cores of Intel Xeon Platinum 8358 CPU), a GPU acceleration cluster (16 NVIDIA A100 80GB cards), a distributed storage cluster, a highway GIS road network platform, and interfaces for the traffic control system.
[0028] Basic software environment: Streaming data processing frameworks: ROS2Humble, Apache Kafka 2.8; Deep learning frameworks: PyTorch 1.13, TensorRT 8.5 (lightweight and accelerated model); Distributed computing framework: Apache Spark 3.3; Coordinate transformation and spatiotemporal reference: WGS84 global coordinate system, BeiDou UTC time reference; Model pre-training and fine-tuning: The feature extraction model is pre-trained on the COCO dataset and fine-tuned using 1 million frames of multimodal data of highway scenes; the scene classifier is trained on the ResNet-18 lightweight model, with a scene classification accuracy of ≥98.5%.
[0029] The complete process of this invention includes 6 core steps: Step S1: Construction of global spatiotemporal benchmark and acquisition of multimodal streaming data Global spatiotemporal reference construction: Using the cloud platform as the time reference source, UTC standard time is obtained through the BeiDou-3 PNT service. Time synchronization messages are sent to all roadside nodes and vehicle-mounted OBUs every 100ms to complete the initial clock deviation calibration of the three-level nodes. After calibration, the initial clock deviation of all nodes with the global reference is determined. ≤5ns; At the same time, a global road network coordinate system was constructed based on the WGS84 coordinate system to complete the initial calibration of the external parameters of all roadside and vehicle-mounted sensors.
[0030] Multimodal streaming data synchronous acquisition: Vehicle-mounted data collection includes visual images, millimeter-wave radar points, lidar point clouds, and vehicle CAN bus data (vehicle speed, steering, and braking status), which are then transmitted in real time to the nearest roadside edge node via 5G-V2X. Roadside data acquisition: full-lane visual images, millimeter-wave radar spot data, lidar point cloud data, meteorological and environmental data, while also receiving data uploaded from the vehicle-mounted terminal; Synchronous cloud-based data collection includes: basic GIS data of highway network, traffic control signal data, and vehicle passage data at toll stations / gantry.
[0031] Step S2: Dynamic spatiotemporal synchronization and alignment of multimodal streaming data Streaming data fragmentation: For streaming data of all modalities, fragmentation is performed according to a fixed time window of 100ms, and each fragment is bound to a unique original acquisition timestamp. , where i is the modality type number and k is the segment number.
[0032] Dynamic timestamp correction: For the k-th segment of the i-th modality, the standard timestamp is calculated using a correction formula. The specific definitions and calculation rules for each parameter are as follows: ; Inherent hardware trigger latency The values were obtained through laboratory calibration and are fixed as follows: 33ms for vision sensors, 100ms for millimeter-wave radar, 100ms for lidar, 20ms for CAN bus, and 100ms for weather sensors. Initial clock skew : The clock deviation between the current sensor node and the global reference is obtained in real time through BeiDou synchronization messages; Transmission delay prediction The exponentially weighted moving average model is used for prediction, and the formula is as follows: , where the initial value = (Actual transmission delay of the first fragment); The smoothing coefficient α is adaptively adjusted based on the variance of the transmission delay of the most recent 10 fragments: variance < 10ms 2 At time α = 0.2, 10 ms 2 ≤variance<50ms 2 When α=0.5, variance≥50ms 2 When α = 0.8.
[0033] Spatial coordinate unification and alignment: Through the sensor extrinsic calibration matrix, the local coordinates of target detection of all modes are uniformly transformed to the WGS84 global road network coordinate system, eliminating coordinate misalignment caused by vehicle movement and roadside installation deviation, and completing spatial alignment.
[0034] Pre-processing steps: After step S2 and before step S3, preprocessing is performed on the synchronized fragmented data: Outlier removal: The sliding window 3σ criterion is used to remove outlier data that exceed the mean ± 3 times the standard deviation within the window, and linear interpolation is used to complete missing data; Lightweight compression: Only lane area ROIs are retained in the visual image, with a compression ratio of 4:1; ground point culling and voxel downsampling are performed on the LiDAR point cloud, with a compression ratio of 2:1; the compression ratio is dynamically adjusted according to the side computing power load and the available bandwidth of the link.
[0035] Step S3: Lightweight Multimodal Feature Fusion Completed at roadside edge nodes, the core principle is weighted feature fusion based on dynamically fused confidence levels, avoiding the poor scene adaptability problem of fixed weights. Unified Dimension Feature Extraction: A lightweight model is used to extract feature vectors from each modality and uniformly map them to a 256-dimensional feature space. Visual features: extracted using the YOLOv8n lightweight model, outputting a 256-dimensional feature vector; LiDAR point cloud features: extracted using a lightweight PointPillars model, outputting a 256-dimensional feature vector; Millimeter-wave radar features: extracted using 1D-CNN, outputting a 256-dimensional feature vector; Vehicle status / weather data: mapped by a fully connected layer, outputting a 256-dimensional feature vector.
[0036] Fusion confidence calculation: The fusion confidence Ci of the i-th modal feature is calculated using the following formula: In the formula, M is the total number of modes participating in the fusion, and satisfies =1, calculation rules for each parameter: Feature signal-to-noise ratio The formula is: ,in The effective signal power of the eigenvector (calculated through the eigenvariance). The noise power of the eigenvector (calculated by the sum of squared feature residuals); Scene adaptation weight : Dynamically calculated based on the classification results of the current scene, the formula is as follows ,in The adaptation probability of the i-th modality output by the scene classifier in the current scene; Scene categories include: weather scenes (sunny / rainy / foggy / snowy) and traffic flow scenes (smooth / slow / congested / accidents). For example, in a foggy scene, the compatibility probability of millimeter-wave radar... =0.8, visual adaptation probability =0.2.
[0037] Weighted feature fusion: via formula Calculate the final edge-side fusion features and output the edge-side local traffic situation awareness results (target detection, trajectory tracking, and abnormal event identification).
[0038] Step S4: Edge-cloud adaptive traffic splitting decision based on optimal cost function The core is to construct a globally optimal cost function, achieving a balance between bandwidth, latency, and accuracy, and then solving for the optimal traffic splitting decision. Construction of the total cost function: Using the total cost function, with the objective of minimizing the total cost, we solve for the optimal solution of the 0-1 integer programming problem. In the formula, N is the total number of data fragments to be decided. ∈{0,1} are the diversion decision variables ( =1 indicates uploading to the cloud. =0 indicates local processing on the side). , , The weighting coefficients for bandwidth, latency, and accuracy loss are respectively, satisfying λB+λD+λL=1.
[0039] Calculation rules for each cost factor: Bandwidth cost factor: ,in Let i be the size of the i-th slice. The available bandwidth limit for the current link is detected in real time and updated every 100ms. Delay cost factor: ,in To reduce upload latency in chunks, To handle latency in the cloud, To mitigate edge processing latency, all data are updated in real time using historical statistical values. Accuracy loss cost factor: ,in To improve the accuracy of edge-side fusion sensing, To improve the accuracy of global fusion perception in the cloud, the test set is pre-calibrated.
[0040] Weight coefficient adaptive adjustment rule: For smooth operation scenarios: λB=0.5, λD=0.3, λL=0.2, prioritize reducing bandwidth usage; Congestion / accident scenarios: λB=0.2, λD=0.2, λL=0.6, prioritizing perception accuracy; In weak network scenarios (available bandwidth < 20Mbps): λB = 0.6, λD = 0.3, λL = 0.1, prioritize reducing bandwidth consumption; For high-load edge scenarios (computing power utilization > 80%): λB = 0.2, λD = 0.6, λL = 0.2, prioritize offloading to the cloud.
[0041] High-priority data forced diversion rule: For high-priority data such as accidents, abnormal parking, wrong-way driving, and traffic violations, xi=1 is forcibly set, and the weight of the latency cost factor λD is doubled to ensure the lowest end-to-end processing latency.
[0042] Optimal solution: The branch and bound method is used to solve the 0-1 integer programming problem, output the optimal split decision sequence, and determine the processing path of each segment.
[0043] Step S5: Cloud-side global temporal-semantic dual-dimensional fusion The cloud-based system performs global correlation and fusion of the received fragmented data and edge-side fusion features across the entire road segment to construct a global traffic situation map. Global Feature Association Matching: Based on a global spatiotemporal benchmark, features uploaded by 20 roadside nodes are spatiotemporally associated to complete cross-node target trajectory matching and deduplication.
[0044] Iterative update of temporal-semantic dual-dimensional fusion weights: The fusion weights are updated using an iterative formula. In the formula, The fusion weights for the features of the i-th edge node in the k-th iteration are calculated according to the following rules: semantic similarity : represents the cosine similarity between the edge features and the global fused features, and the formula is: ,in, Let be the fusion feature of the i-th edge node at time t. This represents the global fusion feature for the (k-1)th iteration; Spatiotemporal credibility The formula is based on spatial distance and time deviation: ,in, Let be the spatial distance between the i-th side node and the perceived target. σd = 1000m is the time synchronization deviation between the data of the i-th side node and the global reference, σT = 50ms is the distance normalization coefficient, and σd = 1000m is the time deviation normalization coefficient.
[0045] Iterative convergence and global feature output: The iteration termination condition is <ε, where the convergence threshold ε is set to 1e-4 by default, and is adjusted to 1e-3 when the cloud computing load is >80%, and to 1e-6 when the load is <30%; After iterative convergence, the final global fusion features are output to construct a global traffic situation map for the entire road segment.
[0046] Step S6: Edge-Cloud Fusion Closed-Loop Optimization The cloud sends optimization parameters to all roadside edge nodes every 5 seconds, completing the entire closed loop process. Based on the global scene recognition results in the cloud, the scene adaptation weights wi of the edge nodes are updated in reverse. For example, in foggy scenes, the visual weights are lowered and the millimeter-wave radar weights are raised. Based on the bandwidth, computing power, and traffic conditions of the entire road segment, the weight coefficients of the cost function for edge diversion decisions are updated in reverse. , , ; Based on the global perception results in the cloud, the edge feature extraction model and scene classifier are incrementally fine-tuned to continuously improve the accuracy of edge perception.
[0047] Examples, comparative examples, and validation data: Experimental setup: Test section: 10km test section of G42 Shanghai-Chengdu Expressway in Jiangsu Province, with 20 roadside edge nodes and 100 test vehicles; Test scenarios: Covering four typical scenarios: clear skies with smooth traffic, rainy days with slow traffic, foggy days (visibility 200m), and accident-related congestion. Each scenario was tested continuously for 1 hour, repeated 100 times, and the average value was taken. Comparison group settings:
[0048] Performance metrics comparison data: Table 1 Comparison of Spatiotemporal Synchronization Accuracy:
[0049] Table 2 Comparison of mAP@0.5 for edge target detection in different scenarios:
[0050] Table 3 Comparison of edge-cloud traffic offloading performance:
[0051] Table 4 Comparison of Cloud-based Global Situation Awareness Performance:
[0052] Example of a typical single-cycle processing flow: Taking a foggy weather incident scenario with a 100ms time window as an example, the complete execution flow is as follows: After global spatiotemporal reference calibration, the clock deviation of all nodes is ≤3ns; Multimodal data was collected from roadside nodes, segmented into 100ms segments, and spatiotemporal synchronization was completed after preprocessing. After timestamp correction, the average time synchronization error was 2.1ms and the spatial alignment error was 11.8cm. 256-dimensional features of each modality were extracted from the edge. In the current foggy scene, the calculated visual SNR was 22dB with an adaptation weight of 0.2, the millimeter-wave radar SNR was 30dB with an adaptation weight of 0.5, and the lidar SNR was 28dB with an adaptation weight of 0.3. The fusion confidence scores were 0.18, 0.52, and 0.30, respectively. After weighted fusion, the edge features were output, and the target detection mAP was 92.1%. Triage Decision: In the event scenario, λ_B=0.2, λ_D=0.2, and λ_L=0.6 are set. High-priority event-related data is forcibly uploaded to the cloud, while raw data is processed locally on the edge. Only feature and abnormal data are uploaded, with a bandwidth utilization of 28.7%, an end-to-end latency of 82ms, and an accuracy loss of 1.1%. The cloud receives features from 20 nodes, iterates 6 times until the weights converge (ε=1e-4), and outputs global fused features. The target detection accuracy is 99.2% across the entire road segment, and the accident event recall rate is 100%. The cloud distributes foggy scene adaptation weights to all nodes to complete closed-loop optimization.
[0053] The present invention achieves the following effects: Significantly improved spatiotemporal synchronization accuracy: The dynamic delay prediction and timestamp correction method reduces the average time synchronization error from over 18ms to 2.3ms and the spatial alignment error from 45cm to 12.5cm, solving the core pain point of mismatch in multimodal data target matching in high-speed mobile scenarios; The robustness of the system in harsh scenarios is significantly enhanced: the dynamic weighted fusion method based on fusion confidence improves the target detection mAP in foggy scenarios from 62.5% to 92.3%, and reduces the accuracy decay rate from 31.6% to 5.9%, which greatly improves the adaptability to all scenarios. Edge-cloud traffic offloading achieves the optimal balance among the three: based on the globally optimal cost function, adaptive traffic offloading reduces bandwidth utilization from 100% to 32.6% and end-to-end latency from 245ms to 89ms compared to the full upload solution, while the accuracy loss is only 1.2%, taking into account bandwidth efficiency, real-time performance and perception accuracy. The global situational awareness capability has been fully upgraded: the temporal and semantic dual-dimensional global fusion combined with closed-loop optimization has achieved a target detection accuracy of 99.1% across the entire road segment and an abnormal event detection recall rate of 99.5%, providing a highly reliable perception foundation for intelligent management and control of highways and autonomous vehicle-road cooperation.
[0054] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A method for expressway vehicle-road cloud multi-modal data streaming acquisition and edge cloud fusion, characterized in that, Includes the following steps: S1. Construction of Global Spatiotemporal Reference and Acquisition of Multimodal Streaming Data: Based on BeiDou / GPS timing, a unified global spatiotemporal reference at the vehicle-road-cloud level is constructed, and multimodal streaming data in highway scenarios is collected synchronously. The multimodal streaming data includes vehicle-mounted vision, millimeter-wave radar, lidar, and CAN bus data; roadside unit (RSU), vision sensor, millimeter-wave radar, lidar, and meteorological sensor data; and cloud-based road network basic data and traffic control data. S2. Dynamic spatiotemporal synchronization and alignment of multimodal streaming data: The collected multimodal streaming data is segmented and the timestamp is corrected and the spatial coordinates are aligned for each data segment to obtain synchronized multimodal data under the global spatiotemporal reference. Specifically, for the k-th streaming data slice of the i-th type of multimodal sensing data source, its standard timestamp is corrected through dynamic latency prediction, and the correction formula is as follows: In the formula, This is the original acquisition timestamp of the k-th segment. For the inherent hardware trigger latency of the i-th type of data source, The current fragment transmission latency is predicted based on the exponentially weighted moving average. The initial clock offset between the i-th type of data source and the global spatiotemporal reference; Simultaneously, the target detection coordinates of each modal data are uniformly transformed to the global road network coordinate system through the external parameter calibration matrix; S3, Lightweight Multimodal Feature Fusion on the Roadside: At the roadside edge computing node, feature extraction of the synchronized multimodal data is performed in a unified dimension, the fusion confidence of each modality feature is calculated, weighted feature fusion is completed based on the fusion confidence, and the edge-side fusion features and local traffic situation perception results are output. The formula for calculating the fusion confidence of the i-th modality feature is as follows: In the formula, M represents the total number of modes participating in the fusion. Let be the signal-to-noise ratio of the i-th modal feature. Assign scene adaptation weights to the i-th modality, and satisfy the following conditions: =1; Final side fusion features are , is the feature vector for the i-th modality; S4. Edge-cloud adaptive traffic splitting decision based on optimal cost function: Construct the total cost function of edge-cloud traffic splitting decision, solve the optimal traffic splitting decision sequence with the goal of minimizing the total cost, and determine the processing path of each data shard. The processing path is either local processing on the edge or uploading to the cloud for processing. The total cost function formula is as follows: In the formula, N is the total number of data fragments to be decided. ∈{0,1} are the flow splitting decision variables. =1 indicates that the data is uploaded to the cloud in chunks. =0 indicates local processing on the edge. As the bandwidth cost factor, As a delay cost factor, Let be the accuracy loss cost factor, and be the weight coefficients of the three types of cost factors, respectively, and satisfy . + + =1; S5. Cloud-side global temporal-semantic dual-dimensional fusion: Global association matching is performed on the fragmented data received from the cloud and the fusion features from the edge side. The temporal-semantic dual-dimensional global fusion is completed through iterative weight updates to construct a global map of the traffic situation of the entire road network. The formula for iteratively updating the weights in the global fusion is as follows: ; In the formula, The fusion weights are the features of the i-th edge node in the k-th iteration. Let be the semantic similarity between the i-th side feature and the global fused feature. Let represent the spatiotemporal reliability of the i-th edge node; After iterating until the weights converge, the final global fusion feature is output. S6. Edge-Cloud Fusion Closed-Loop Optimization: Based on the global fusion results and global situation map in the cloud, the fusion confidence scenario adaptation weight and the traffic diversion decision cost function weight coefficient of the edge nodes are updated in reverse to achieve closed-loop dynamic optimization of the entire vehicle-road-cloud fusion process. 2.The highway cloud of vehicle-road multi-modal data streaming acquisition and edge cloud fusion method according to claim 1, characterized in that, In step S2, the transmission delay prediction value The exponential weighted moving average calculation formula is: In the formula, α is the smoothing coefficient, and its value ranges from 0 to 1. The actual transmission delay of the (k-1)th fragment, initial value = ; The smoothing coefficient α is adaptively adjusted according to the degree of link delay fluctuation; the larger the variance of delay fluctuation, the larger the value of α. 3.The highway car-road cloud multi-modal data streaming acquisition and edge cloud fusion method of claim 1, characterized in that, In step S3, the scene adaptation weight Based on the current highway scene classification result dynamic calculation, the calculation formula is: ; wherein is the adaptation probability of the i-th modality output by the scene classifier under the current scene, which includes the weather scene (sunny / rainy / foggy / snowy) and the traffic flow scene (free-flow / slow / congestion / accident).
4. The method for multimodal data streaming acquisition and edge-cloud fusion of highway vehicle-road-cloud according to claim 1, characterized in that, In step S3, the signal-to-noise ratio of the i-th modal feature The calculation formula is: ; In the formula, is the effective signal power of the i-th modal feature vector, which is calculated by the feature variance, is the noise power of the i-th modal feature vector, which is calculated by the feature residual sum of squares.
5. The highway car-road cloud multi-modal data streaming acquisition and edge cloud fusion method according to claim 1, characterized in that, In step S4, the bandwidth cost factor Delay cost factor Accuracy loss cost factor The calculation formulas are as follows: ; ; ; In the formula, Let i be the size of the i-th data segment. This represents the maximum available bandwidth for the current link. To reduce upload latency in chunks, To handle latency in the cloud, To account for edge processing delay, To improve the accuracy of edge-side fusion sensing, To improve the accuracy of global fusion perception in the cloud.
6. The method for multimodal data streaming acquisition and edge-cloud fusion of highway vehicles, roads, and cloud as described in claim 1, is characterized in that... In step S4, the weight coefficient 、 、 According to the current highway scene dynamic self-adaptive adjustment, the adjustment rule is: In congestion / accident scenarios, the weight of λD is reduced and the weight of λL is increased to prioritize perception accuracy and real-time decision-making. In low-bandwidth / weak network scenarios, increase the weight of λB and decrease the weight of λD to prioritize reducing link bandwidth usage; In scenarios with high computing power consumption on the edge, the weight of λD is increased to prioritize processing in the cloud.
7. The highway car-road cloud multi-modal data streaming acquisition and edge cloud fusion method according to claim 1, characterized in that, In step S5, the semantic similarity and the spatiotemporal credibility are calculated according to the following formulas, respectively: ; ; In the formula, Let be the fusion feature of the i-th edge node at time t. For the (k-1)th iteration, the global fusion feature is... Let be the spatial distance between the i-th side node and the perceived target. The time synchronization deviation between the data of the i-th side node and the global reference is... This is the distance normalization coefficient. This is the time deviation normalization coefficient. 8.The highway cloud of vehicle-road multi-modal data streaming acquisition and edge cloud fusion method according to claim 1, characterized in that, In step S5, the termination condition for iterative convergence is: <ε, where ε is the convergence threshold, and its value ranges from 10⁻⁶ ≤ ε ≤ 10⁻³. The convergence threshold is adaptively adjusted according to the cloud computing load.
9. The highway car-road cloud multi-modal data streaming acquisition and edge cloud fusion method according to claim 1, characterized in that, Before step S2, there is also a multimodal streaming data preprocessing step, which removes outliers, denoises, and performs lightweight compression on the raw acquired streaming data; Among them, outlier removal adopts the 3σ criterion based on sliding window, which removes and interpolates outlier data that exceeds the mean ± 3 times the standard deviation within the window. At the same time, high-dimensional point cloud and image data are subjected to lightweight compression based on region of interest, and the compression ratio is dynamically adjusted according to the side computing power and link bandwidth.
10. The highway car-road cloud multi-modal data streaming acquisition and edge cloud fusion method according to claim 1, characterized in that, In step S4, a data priority factor is added to the diversion decision. A mandatory upload rule is set for high-priority data such as accidents, abnormal events, and traffic violations, and their diversion decision variables are forced to be set to 1. At the same time, the weight of the latency cost factor for high-priority data is doubled in the total cost function to ensure that the end-to-end transmission and processing latency of high-priority data is minimized.