Spatiotemporal alignment-based vehicle-road-cloud integrated end-to-end autonomous driving device and method

By using the spatiotemporal alignment technology of vehicle-road-cloud integration, deep fusion of vehicle and road information is achieved, which solves the shortcomings of end-to-end autonomous driving algorithms in long-distance perception and occlusion scenarios, improves the reliability and real-time response capability of the system, adapts to different roadside sensor deployment environments, and reduces the computing load on the vehicle.

CN121578722BActive Publication Date: 2026-04-03AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-27
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing end-to-end autonomous driving algorithms are insufficient in long-distance perception, small object detection, and occlusion scenarios. The information fusion in vehicle-road cooperative systems is insufficient, decision response is lagging, and generalization ability is insufficient, making it difficult to meet the requirements of high precision, high reliability, and high real-time performance.

Method used

An end-to-end autonomous driving device based on spatiotemporal alignment is adopted, which extracts roadside features through roadside equipment and integrates them with vehicle features for modeling. Historical time series information is introduced, and a deep learning model is used for feature time alignment and fusion to optimize computational efficiency and achieve deep fusion of vehicle and road information.

Benefits of technology

It improves perception and decision-making capabilities for small targets at long distances and in occluded scenarios, enhances the reliability and versatility of autonomous driving systems, solves the scenario adaptation defects of single-vehicle intelligent algorithms, reduces the computing load on the vehicle, and ensures real-time response performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121578722B_ABST
    Figure CN121578722B_ABST
Patent Text Reader

Abstract

This invention discloses an end-to-end autonomous driving device and method based on spatiotemporal alignment and vehicle-road-cloud integration, aiming to solve the limitations of single-vehicle intelligent end-to-end algorithms and the incompatibility issues in vehicle-road cooperative perception fusion. The device includes roadside equipment and a vehicle-side component. The roadside equipment extracts and transmits roadside features, while the vehicle-side component compensates for transmission delays through a spatiotemporal alignment network module, projects the roadside features onto a unified BEV space, and then inputs them into a prediction and planning network module after dynamic fusion. The method includes roadside processing, vehicle-side processing, spatiotemporal alignment, feature fusion, and end-to-end planning steps, and employs a multi-task loss function for optimization. This invention achieves deep fusion of vehicle and road information, adapts to heterogeneous sensor scenarios, reduces vehicle-side load, improves perception and planning accuracy in long-tail scenarios, and enhances the reliability and safety of autonomous driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of new energy vehicle technology, and in particular to an end-to-end autonomous driving device and method based on spatiotemporal alignment and vehicle-road-cloud integration. Background Technology

[0002] While autonomous driving technology has developed rapidly in recent years, it still faces several critical issues in practical applications, severely restricting its reliability and widespread adoption: First, the complexity of the environment presents adaptation challenges. In complex scenarios such as urban traffic congestion and non-standard intersection layouts, the accuracy of system perception and decision-making is easily affected, making stable adaptation difficult. Second, the algorithm's generalization ability is insufficient. When faced with rare driving scenarios or unexpected situations, traditional algorithms have limited scenario transfer capabilities and cannot guarantee consistent performance under different operating conditions. Third, the data processing and computing load is high. Autonomous vehicles need to integrate massive amounts of data from multiple sensors in real time, placing extremely high demands on the system's parallel computing capabilities and data fusion efficiency. Existing solutions struggle to balance real-time performance and processing accuracy. Fourth, the system integration complexity is high. The independent division of perception, prediction, planning, and control modules in traditional architectures leads to information loss during inter-module transmission, and the performance fluctuations of individual modules directly affect the overall system performance, making integration and optimization extremely difficult.

[0003] To address these challenges, the industry has proposed end-to-end autonomous driving software solutions. The core idea is to establish a direct mapping between raw sensor data and vehicle control commands, abandoning the explicit separation of functional modules in traditional architectures. Instead, deep learning models are used to integrate perception, prediction, planning, and control functions, forming a unified end-to-end processing flow. Theoretically, this type of algorithm offers significant advantages: simplified system architecture, reduced information loss and computational redundancy between modules; data-driven characteristics, theoretically adaptable to different driving scenarios, improving generalization capabilities; elimination of intermediate processing steps, optimizing computational efficiency and real-time response speed; enhanced adaptability to complex scenarios, improving system robustness in dynamic environments.

[0004] However, in the past two years, the practical application of end-to-end autonomous driving algorithms has remained limited by the technical framework of single-vehicle intelligence, exhibiting long-tail problems that are difficult to overcome: for example, insufficient perception accuracy for distant small objects and weak ability to handle sudden scenarios such as "ghost pedestrians." Several recent safety incidents related to intelligent connected vehicles are directly related to these limitations. Vehicle-road cooperative technology, by integrating long-range, wide-field-of-view information from roadside sensors, could have effectively addressed these long-tail problems, but current mainstream end-to-end algorithms still focus on single-vehicle intelligence scenarios. Neither numerous well-known algorithms in academia nor autonomous driving systems launched by mainstream companies in the industry have achieved effective fusion of vehicle and roadside information. Only one algorithm exists that integrates roadside and vehicle information end-to-end, but it suffers from multiple technical defects and cannot meet the application requirements of actual autonomous driving systems.

[0005] Roadside features are expressed using complex features such as OCC and BEV, which leads to a surge in computational load. Existing roadside equipment is unable to handle this computational demand. If the execution is moved to the vehicle, it will significantly increase the vehicle's computational load and affect the overall operating efficiency of the vehicle.

[0006] Feature data relies on a single camera for acquisition, resulting in a limited field of view coverage. This severely restricts the system's ability to perceive complex environments in all directions and easily leads to blind spots.

[0007] Decision-making and reasoning are based solely on the current frame data, without incorporating historical time-series information. This results in insufficient stability and smoothness of the output, making it unable to adapt to dynamically changing driving scenarios.

[0008] The model design did not take into account the relative positional differences between roadside and vehicle-mounted sensors, resulting in a lack of scene transfer capability, poor generalization performance, and difficulty in adapting to different roadside deployment environments and vehicle configurations.

[0009] In the technological and patent landscape related to vehicle-road cooperative autonomous driving, existing solutions still have significant limitations: most technologies adopt a "perception-fusing" model or rely solely on roadside perception results as the basis for decision-making, failing to achieve deep integration of vehicle and road information. For example, some related patents acquire roadside perception results through vehicle-road cooperative communication devices and then perform post-fusion processing with vehicle-side perception data; some patents obtain numerical results of road traffic conditions from the roadside for vehicle positioning fusion optimization; some patents directly use the long-distance obstacle identification results obtained by roadside sensors as the final perception conclusion to guide vehicle obstacle avoidance operations; and some patents obtain traffic participant status information through vehicle-road cooperative networks in intersection areas, only as data support for the decision-making process. None of the above solutions break through the limitations of traditional architectures, exhibiting problems such as insufficient information fusion, delayed decision response, and insufficient generalization ability, making it difficult to meet the actual needs of autonomous driving systems for high precision, high reliability, and high real-time performance. Summary of the Invention

[0010] The purpose of this invention is to provide an end-to-end autonomous driving device and method based on spatiotemporal alignment and vehicle-road-cloud integration. It mainly addresses the shortcomings of single-vehicle intelligent end-to-end algorithms in long-distance perception, small object detection, and occlusion scenarios, as well as the incompatibility issues of existing vehicle-road cooperative systems that use perception-based fusion.

[0011] To achieve the above objectives, the present invention provides a spatiotemporally aligned, vehicle-road-cloud integrated end-to-end autonomous driving device, comprising:

[0012] Roadside equipment, used to extract roadside features based on a roadside sensing network;

[0013] The vehicle-side component includes a vehicle-road-cloud integrated end-to-end network, which receives vehicle sensor data and roadside features.

[0014] The integrated vehicle-road-cloud end-to-end network includes a roadside feature acquisition unit, a vehicle-side feature acquisition unit, a vehicle-road feature fusion module, and a prediction and planning network module.

[0015] The roadside feature acquisition unit includes a spatiotemporal alignment network module corresponding to the roadside equipment.

[0016] The spatiotemporal alignment network module includes a historical roadside feature update submodule, a historical roadside feature pool, a roadside feature time alignment network, and a vehicle-side BEV encoder; wherein: the historical roadside feature update submodule is used to take the roadside features of the n times before the current roadside feature a as historical roadside feature h1, input them into the historical roadside feature pool, and store them in the historical roadside feature pool, while deleting the (n+1)th roadside feature from the historical roadside feature pool; the roadside feature time alignment network predicts the roadside features at the future time (t+Δt) based on the current roadside feature a, the transmission delay display value Δt, and the historical roadside feature h1, in order to compensate for the transmission delay within a preset range, and at the same time, projects the roadside features at the future time (t+Δt) onto the BEV space consistent with the vehicle-side features through the vehicle-side BEV encoder to obtain spatiotemporally aligned roadside features; wherein, the loss value Lpre of the roadside feature time alignment network is expressed as Equation (1):

[0017] Lpre = sum ( Wm * | PRE[ L(m,t),Δt) ]- L(m,t+Δt) | 2 (1)

[0018] Where, sum represents the weighted summation function, L(m,t) represents the feature prediction value from the m-th roadside device at time t, L(m,t+Δt) represents the feature value of the m-th roadside device at time (t+Δt), PRE is the roadside feature spatiotemporal alignment network, Wm represents the weight of the m-th roadside device; * represents multiplication, | |2 Represents the L2 norm of a vector;

[0019] The vehicle-side feature acquisition unit is used to extract vehicle-side features based on vehicle sensor data. The vehicle-side BEV feature encoder transforms the vehicle-side features into a vehicle-centric BEV space to obtain spatiotemporally aligned vehicle-side features.

[0020] The vehicle-road feature fusion module is used to dynamically fuse spatiotemporally aligned roadside features and spatiotemporally aligned vehicle-end features to output vehicle-road fused features;

[0021] The predictive planning network module predicts the motion of other vehicles based on vehicle-road fusion features and outputs the vehicle's planned trajectory. During training, the predictive planning network module adopts a multi-task loss function L: L = Ldet + Lmap + Lmotion + Lplan + Lpre; where Ldet, Lmotion, Lmap, and Lplan represent the dynamic element detection loss, motion prediction loss, map building loss, and planning loss calculated by the predictive planning network module, respectively.

[0022] Furthermore, Wm is obtained using equation (2):

[0023] Wm =(e -Δs_m ) / sum(e -Δs_m (2)

[0024] Where Δs_m is the distance between the vehicle and the m-th roadside device.

[0025] Furthermore, the roadside feature time alignment network uses a deep learning model F(a, Δt, h1) to predict the time-aligned (t+Δt) roadside features by taking the current roadside feature a, the transmission delay display value Δt, and the historical roadside feature h1 as inputs. Δt is represented by 1 to 5 to indicate a transmission delay of 100 to 500 ms.

[0026] Furthermore, the vehicle-road feature fusion module is also used to support handling situations where the number of roadside sensors is inconsistent through zero-value padding. When the number of roadside sensors in the actual scenario is less than the preset number, zero-value padding is used during fusion.

[0027] Furthermore, the predictive planning network module includes:

[0028] The historical vehicle-road fusion feature update submodule is used to manage the time series of historical vehicle-road fusion features and store them in the historical vehicle-road fusion feature pool.

[0029] The historical vehicle-road fusion feature attention submodule, based on the data stored in the historical vehicle-road fusion feature pool, analyzes the temporal dependencies through an attention network, assigns weights to vehicle-road fusion features at different times, and calculates weighted historical vehicle-road fusion features;

[0030] The historical vehicle-road fusion feature network submodule is used to fuse weighted historical vehicle-road fusion features with current vehicle-road fusion features to generate enhanced vehicle-road fusion features;

[0031] The map network comprises a map perception submodule, a map perception header, and a map attention network. The map perception submodule is used to extract map-related features from the enhanced vehicle-road fusion features and identify the spatial distribution of map elements. The map perception header is used to decode map features during the model training phase, output real-time map perception results, and calculate Lmap to supervise the model. The map attention network is used to apply an attention mechanism to map features, filter key map information, and output the results.

[0032] The Traffic Participant Network is used to perceive and encode the state of traffic participants. It includes a Traffic Participant Perception Submodule and a Traffic Participant Perception Header. The Traffic Participant Perception Submodule is used to extract traffic participant features from vehicle-road fusion features and identify the position, speed and category of the participants. The Traffic Participant Perception Header is used to decode traffic participant features during the training phase, output real-time perception results, and calculate Ldet to supervise the model.

[0033] The prediction network predicts the future trajectories of other vehicles or pedestrians based on map and traffic participant features. It includes a prediction submodule and a prediction header. The prediction submodule is used to fuse the weighted map features output by the map attention network and the traffic participant perception features to predict the future state of the traffic participants. The prediction header is used to decode the prediction features during the training phase, output the motion prediction results, and calculate Lmotion for supervision.

[0034] The planning submodule integrates features from the prediction network output, map features, and traffic participant features. It then uses a planning algorithm to generate the vehicle's planned trajectory and calculates Lplan to optimize driving smoothness and safety.

[0035] This invention also includes a spatiotemporally aligned, vehicle-road-cloud integrated end-to-end autonomous driving method, comprising:

[0036] Roadside processing steps: Acquire roadside sensor data and extract roadside features using a roadside sensing network;

[0037] Vehicle-side processing steps: Extract vehicle-side features based on vehicle sensor data, and transform the vehicle-side features into a vehicle-centric BEV space through a vehicle-side BEV feature encoder to obtain spatiotemporally aligned vehicle-side features;

[0038] Spatiotemporal alignment steps: The spatiotemporal alignment network module takes the roadside features n times before the current roadside feature a as the historical roadside feature h1, and predicts the roadside features at the future time (t+Δt) based on the current roadside feature a, the transmission delay display value Δt, and the historical roadside feature h1 to compensate for the transmission delay within the preset range. At the same time, the roadside features at the future time (t+Δt) are projected onto the BEV space consistent with the vehicle-side features through the vehicle-side BEV encoder to obtain the spatiotemporally aligned roadside features; where the loss value Lpre of the roadside feature temporal alignment network is expressed as Equation (1):

[0039] Lpre = sum ( Wm * | PRE[ L(m,t),Δt) ]- L(m,t+Δt) | 2 (1)

[0040] Where, sum represents the weighted summation function, L(m,t) represents the feature prediction value from the m-th roadside device at time t, L(m,t+Δt) represents the feature value of the m-th roadside device at time (t+Δt), PRE is the roadside feature spatiotemporal alignment network, Wm represents the weight of the m-th roadside device; * represents multiplication, | | 2 Represents the L2 norm of a vector;

[0041] Feature fusion steps: Dynamically fuse spatiotemporally aligned roadside features and spatiotemporally aligned vehicle-side features to output vehicle-road fusion features;

[0042] End-to-end planning steps: Based on vehicle-road fusion features, the prediction planning network module predicts the motion of other vehicles and outputs the vehicle's planned trajectory. During training, the prediction planning network module adopts a multi-task loss function L: L = Ldet + Lmap + Lmotion + Lplan + Lpre; where Ldet, Lmotion, Lmap, and Lplan represent the dynamic element detection loss, motion prediction loss, map building loss, and planning loss calculated by the prediction planning network module, respectively.

[0043] Furthermore, Wm is obtained using equation (2):

[0044] Wm =(e -Δs_m ) / sum(e -Δs_m (2)

[0045] Where Δs_m is the distance between the vehicle and the m-th roadside device.

[0046] Furthermore, in the spatiotemporal alignment step, the roadside feature temporal alignment network uses a deep learning model F(a, Δt,h1) to predict the time-aligned (t+Δt) roadside features, with the current roadside feature a, the transmission delay display value Δt, and the historical roadside feature h1 as inputs. Δt is represented by 1 to 5 to indicate a transmission delay of 100 to 500 ms.

[0047] Furthermore, in the feature fusion step, the vehicle-road feature fusion module is also used to support handling situations where the number of roadside sensors is inconsistent through zero-value padding. When the number of roadside sensors in the actual scenario is less than the preset number, zero-value padding is used during fusion.

[0048] Furthermore, in the end-to-end planning step, the predictive planning network module includes:

[0049] The historical vehicle-road fusion feature update submodule is used to manage the time series of historical vehicle-road fusion features and store them in the historical vehicle-road fusion feature pool.

[0050] The historical vehicle-road fusion feature attention submodule, based on the data stored in the historical vehicle-road fusion feature pool, analyzes the temporal dependencies through an attention network, assigns weights to vehicle-road fusion features at different times, and calculates weighted historical vehicle-road fusion features;

[0051] The historical vehicle-road fusion feature network submodule is used to fuse weighted historical vehicle-road fusion features with current vehicle-road fusion features to generate enhanced vehicle-road fusion features;

[0052] The map network comprises a map perception submodule, a map perception header, and a map attention network. The map perception submodule is used to extract map-related features from the enhanced vehicle-road fusion features and identify the spatial distribution of map elements. The map perception header is used to decode map features during the model training phase, output real-time map perception results, and calculate Lmap to supervise the model. The map attention network is used to apply an attention mechanism to map features, filter key map information, and output the results.

[0053] The Traffic Participant Network is used to perceive and encode the state of traffic participants. It includes a Traffic Participant Perception Submodule and a Traffic Participant Perception Header. The Traffic Participant Perception Submodule is used to extract traffic participant features from vehicle-road fusion features and identify the position, speed and category of the participants. The Traffic Participant Perception Header is used to decode traffic participant features during the training phase, output real-time perception results, and calculate Ldet to supervise the model.

[0054] The prediction network predicts the future trajectories of other vehicles or pedestrians based on map and traffic participant features. It includes a prediction submodule and a prediction header. The prediction submodule is used to fuse the weighted map features output by the map attention network and the traffic participant perception features to predict the future state of the traffic participants. The prediction header is used to decode the prediction features during the training phase, output the motion prediction results, and calculate Lmotion for supervision.

[0055] The planning submodule integrates features from the prediction network output, map features, and traffic participant features. It then uses a planning algorithm to generate the vehicle's planned trajectory and calculates Lplan to optimize driving smoothness and safety.

[0056] The present invention has the following advantages due to the adoption of the above technical solutions:

[0057] 1. Because this invention constructs an end-to-end fusion architecture for vehicle-road cooperation, it abandons the traditional "perception-after fusion" mode and directly performs integrated modeling of the original data from the vehicle and the roadside. It does not rely on explicit perception results and realizes an end-to-end closed loop of data fusion and control command generation. Therefore, it achieves deep end-to-end fusion of vehicle and road information, successfully adapts to the end-to-end autonomous driving paradigm, breaks through the technical limitations of single-vehicle intelligent end-to-end algorithms, and solves the core contradiction of the incompatibility between the existing vehicle-road cooperative "perception-after fusion" mode and the end-to-end autonomous driving paradigm. It breaks through the technical bottleneck of traditional post-fusion relying on explicit perception results and being unable to adapt to the end-to-end architecture, while avoiding the problem of fusion result distortion caused by perception confidence deviation.

[0058] 2. Because this invention introduces a historical time-series information modeling mechanism, it inputs both roadside historical data and vehicle-side real-time data into the model. Through a time-series prediction algorithm, it infers the current roadside feature status based on the received lagging roadside features, and specifically eliminates the impact of data transmission delay. Therefore, it effectively eliminates the time-series deviation caused by the 100-500ms roadside feature transmission delay, improves the accuracy of vehicle-road data fusion, and ensures the stability of decision-making basis in dynamic scenarios. This solves the roadside feature transmission delay problem, eliminates the time-series difference between the roadside features received by the vehicle and its own real-time data, and avoids decision-making deviations caused by direct fusion.

[0059] 3. Because this invention uses a heterogeneous sensor adaptive adaptation module and a dynamic adjustment mechanism for input dimensions, the model can be compatible with heterogeneous scenarios with different numbers of roadside sensors deployed, ensuring the stability of data fusion under different deployment environments. Therefore, it achieves adaptive compatibility with different numbers of roadside sensors deployed, greatly improves the model's scenario generalization ability, reduces the deployment threshold and adaptation cost of vehicle-road cooperative systems, and solves the problem of heterogeneous deployment adaptation with inconsistent numbers of roadside sensors. It also overcomes the problem of inconsistent data input dimensions and difficulty in adapting existing models caused by differences in the number of sensors deployed at different intersections.

[0060] 4. Because this invention adopts a roadside-vehicle distributed computing strategy, it fully utilizes the computing resources of roadside computing devices to preprocess and extract roadside features, and only transmits low-dimensional effective features to the vehicle, thus optimizing the efficiency of computing power allocation. Therefore, it significantly reduces the computing load on the vehicle, ensures the real-time response performance of the autonomous driving system, and meets the stringent requirements of decision-making timeliness in actual road driving. This solves the problem of the surge in vehicle computing load introduced by multiple sensors. Through distributed computing power allocation, it avoids the problem of real-time inference being impossible due to insufficient vehicle computing power.

[0061] This invention significantly improves perception and decision-making capabilities in long-tail scenarios such as distant small targets, severe occlusion, and "ghost peeks" (sudden appearances by pedestrians), thereby significantly enhancing the reliability, safety, and versatility of autonomous driving systems. It addresses the issues of insufficient perception and decision-making capabilities, low reliability, and limited versatility in long-tail scenarios for autonomous driving systems, compensates for the scenario adaptation deficiencies of single-vehicle intelligent end-to-end algorithms, and improves driving safety and stability in complex environments. This invention is applied to the new energy vehicle industry, integrating vehicle-side and roadside information to solve problems such as poor long-distance perception accuracy in autonomous driving. Attached Figure Description

[0062] Figure 1 This is an architecture diagram of an end-to-end autonomous driving device based on spatiotemporal alignment for vehicle-road-cloud integration according to an embodiment of the present invention.

[0063] Figure 2 This is a diagram of the roadside feature spatiotemporal alignment network module architecture according to an embodiment of the present invention.

[0064] Figure 3 This is a diagram of the predictive planning network module architecture according to an embodiment of the present invention. Detailed Implementation

[0065] In the accompanying drawings, the same or similar reference numerals are used to denote the same or similar elements or elements having the same or similar functions. The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0066] In the description of this invention, the terms "center," "longitudinal," "lateral," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the scope of protection of this invention.

[0067] In scenarios such as congested intersections and obstructed views, the performance of single-vehicle intelligent autonomous driving is highly unstable due to significant interference with the information perceived by vehicle-mounted sensors. In these scenarios, roadside sensors, deployed on high light poles, and typically with 1 to 4 sets of roadside equipment at a single intersection, offer a higher field of view and a wider field of view compared to single-vehicle intelligent systems. This provides inherent advantages in resolving intersection traffic congestion, large vehicle obstruction at close range, unexpected pedestrian appearances, and high-speed remote object detection. This embodiment, combining the concept of end-to-end autonomous driving, proposes a vehicle-road-cloud integrated end-to-end autonomous driving device and method based on spatiotemporal alignment.

[0068] The vehicle-road-cloud integrated end-to-end autonomous driving device based on spatiotemporal alignment provided in this embodiment includes roadside equipment and vehicle-side equipment.

[0069] Roadside equipment is used to extract roadside features based on roadside sensing networks.

[0070] In one embodiment, such as Figure 1 As shown, m roadside devices are pre-installed near the autonomous vehicle (hereinafter referred to as the vehicle). If we assume that there is one set of roadside devices at each corner of the intersection, then typically there are a total of 4 roadside devices near the autonomous vehicle, where m = 1, 2, 3, 4. Each roadside device includes a roadside sensor and a GPU (Graphics Processing Unit) deep learning network. The GPU deep learning network corresponds to the roadside perception network mentioned in this paper.

[0071] Roadside sensors can be conventional sensors such as cameras, millimeter-wave radar, lidar, and weather sensors, used to acquire environmental data around the roadside equipment. This embodiment uses four roadside sensor inputs, and each sensor can be arranged clockwise according to its relationship with the vehicle's pose, but is not limited to this.

[0072] GPU deep learning networks are used for roadside perception inference, including the roadside perception network backbone. The roadside perception network backbone acquires data collected by roadside sensors and outputs roadside features. The roadside equipment also includes an FPGA (Field-Programmable Gate Array) for data processing, including a compression module and an output module. The compression module compresses the roadside features in real time to obtain compressed roadside features, which are then broadcast through the output module.

[0073] The GPU deep learning network also includes a roadside perception network (hender), which receives roadside features from the backbone of the roadside perception network and outputs roadside perception results. In this embodiment, these roadside perception results are not used.

[0074] This embodiment compresses the intermediate feature results of the GPU deep learning network and outputs them via FPGA. Based on the programmability of FPGA, this embodiment requires no hardware modification to the roadside equipment, only a simple modification to the FPGA programming, thus significantly reducing the cost of roadside equipment upgrades. It fully utilizes the computing power of the roadside equipment, reduces the computational load on the vehicle side, and allows the computational results of the roadside equipment to be reused by multiple vehicles, resulting in excellent economic efficiency.

[0075] The vehicle-side includes onboard sensors and a vehicle-road-cloud integrated end-to-end network. The vehicle-road-cloud integrated end-to-end network is used to receive vehicle sensor data and compressed roadside features output by the FPGA.

[0076] In one embodiment, combined Figure 1 The vehicle-road-cloud integrated end-to-end network includes a roadside feature acquisition unit, a vehicle-side feature acquisition unit, a vehicle-road feature fusion module, and a prediction and planning network module.

[0077] The roadside feature acquisition unit is used to receive and decompress roadside features, align the decompressed roadside features with vehicle-end features in time and space, and input them into the vehicle-road feature fusion module.

[0078] In a preferred embodiment of the roadside feature acquisition unit, the roadside feature acquisition unit includes a decompression module and a spatiotemporal alignment network module corresponding to the roadside equipment. The decompression module is used to decompress the corresponding compressed roadside features and send them to the corresponding spatiotemporal alignment network module.

[0079] In one embodiment, such as Figure 2 As shown, Figure 2The roadside features above the dashed box refer to the decompressed roadside features output by the decompression module. The spatiotemporal alignment network module includes a historical roadside feature update submodule, a historical roadside feature pool, a roadside feature temporal alignment network, and a vehicle-side BEV (Bird's Eye View) encoder.

[0080] The historical roadside feature update submodule is used to receive the decompressed roadside features, take the roadside features n times before the current roadside feature a as historical roadside features h1, and input them into the historical roadside feature pool for storage.

[0081] The historical roadside feature pool is used to store historical roadside features h1. When a new roadside feature is input into the historical roadside feature update submodule, the (n+1)th roadside feature is deleted from the historical roadside feature pool to maintain a fixed-length historical queue and ensure the continuity and real-time nature of the time-series information.

[0082] In one embodiment, the roadside feature time alignment network uses a deep learning model F(a, Δt, h1) as input, taking the current roadside feature a, the transmission delay display value Δt, and the historical roadside feature h1 as input, and outputs the time-aligned predicted roadside feature value for the future time (t+Δt). The predicted roadside feature value is the roadside feature for the future time (t+Δt) after compensating for the delay Δt. The acquisition method may include, for example, concatenating the historical roadside feature h1 stored in the historical roadside feature pool with the roadside feature at the current time (concat) to predict the time-aligned roadside feature for the future time (t+Δt) to compensate for the transmission delay within a preset range. The transmission delay within the preset range is set for the transmission delay of roadside features. In order to ensure that the roadside features are aligned with the vehicle-side features in the time dimension, in this implementation, Δt is used to represent a transmission delay of 100 to 500 ms using 1 to 5. That is, when Δt is displayed as 1, it means a transmission delay of 100 ms; when Δt is displayed as 2, it means a transmission delay of 200 ms; and so on for other values ​​of Δt.

[0083] The vehicle-side BEV encoder is used to project roadside features at a future time (t+Δt) onto a BEV space that is consistent with the vehicle-side features, thereby obtaining spatiotemporally aligned roadside features.

[0084] In one embodiment, the vehicle-side feature acquisition unit is used to receive vehicle sensor data, extract vehicle-side features based on the vehicle sensor data, and convert the vehicle-side features into a vehicle-centric BEV space through a vehicle-side BEV feature encoder to obtain spatiotemporally aligned vehicle-side features.

[0085] In one embodiment, the vehicle-side feature acquisition unit includes a vehicle-side backbone and a vehicle-side BEV feature encoder. The vehicle-side backbone is used to extract vehicle-side features based on vehicle sensor data. The vehicle-side features are encoded into BEV features by the vehicle-side BEV feature encoder and converted into the vehicle's BEV space to obtain spatiotemporally aligned vehicle-side features.

[0086] Combination Figure 1 The spatiotemporally aligned roadside features and spatiotemporally aligned vehicle-end features are simultaneously input into the vehicle-road feature fusion module. The vehicle-road fusion module dynamically fuses the spatiotemporally aligned roadside features and spatiotemporally aligned vehicle-end features through processes such as splicing, to obtain vehicle-road fused features, which are then output.

[0087] In one embodiment, the vehicle-road feature fusion module is also used to support handling situations where the number of roadside sensors is inconsistent through zero-value padding. That is, when the number of roadside sensors in the actual scenario is less than the preset number (4 units), zero-value padding is used during fusion. For example, to improve the stability of the model in handling inconsistent roadside sensor numbers, some roadside sensors may be randomly disabled during model training. In this case, zero-value padding is needed to enhance the robustness of the model.

[0088] The predictive planning network module receives vehicle-road fusion features, predicts the movement of other vehicles based on the current vehicle-road fusion features, and outputs the vehicle's planned trajectory.

[0089] In one embodiment, such as Figure 3 As shown, the prediction planning network module includes a historical vehicle-road fusion feature update submodule, a historical vehicle-road fusion feature attention submodule, a historical vehicle-road fusion feature network submodule, a map network, a traffic participant network, a prediction network, and a planning submodule.

[0090] The historical vehicle-road fusion feature update submodule manages the time series of historical vehicle-road fusion features. Specifically, it receives the vehicle-road fusion features at the current moment and updates the historical vehicle-road fusion feature pool. For example, the historical vehicle-road fusion feature pool retains features from the past 5 frames, with 2 frames per second, for a total of 2.5 seconds of data. This submodule deletes the oldest features each time a new feature is input to maintain a fixed-length historical queue, ensuring the continuity and real-time nature of the time series information.

[0091] The historical vehicle-road fusion feature attention submodule is used to apply an attention mechanism to historical vehicle-road fusion features and calculate weighted historical vehicle-road fusion features. This submodule, based on multi-frame data in the historical vehicle-road fusion feature pool, analyzes temporal dependencies through an attention network (such as Transformer or a similar structure) and assigns weights to vehicle-road fusion features at different times, thereby highlighting key historical information and enhancing the contextual awareness of current decisions. The weight Wi of the i-th historical vehicle-road fusion feature can be obtained using equation (3):

[0092] Wi =e -Δt1 (3)

[0093] In the formula, Δt1 represents the time difference between the i-th historical vehicle-road fusion feature and the current vehicle-road fusion feature, in seconds (s). This weight value can also be set as a fixed value.

[0094] The historical vehicle-road fusion feature network submodule is used to fuse weighted historical vehicle-road fusion features with current vehicle-road fusion features. This submodule concatenates or weights the weighted historical vehicle-road fusion features output by the attention submodule with the features at the current time step to generate enhanced vehicle-road fusion features, thereby improving the model's adaptability and stability to dynamic environments.

[0095] The map network is used to perceive and encode high-precision map elements such as lane lines and road signs. This network outputs map-related features, supporting online map construction and perception. The map network includes a map perception submodule, a map perception header, and a map attention network. The map perception submodule extracts map-related features from vehicle-to-infrastructure (V2I) features and identifies the spatial distribution of map elements through convolutional or encoder structures. The map perception header decodes map features during model training, outputs real-time map perception results (such as lane line boundaries), and calculates the map construction loss Lmap to supervise model optimization; this header does not participate in calculations during the inference phase. The map attention network applies an attention mechanism to map features, filters key map information, and outputs weighted features to the prediction submodule to improve prediction accuracy.

[0096] The Traffic Participant Network (TPN) is used to perceive and encode the state of traffic participants (such as vehicles and pedestrians). This network outputs traffic participant-related features, supporting object detection and tracking. The TPN consists of a traffic participant perception submodule and a traffic participant perception header. The traffic participant perception submodule extracts traffic participant features from vehicle-to-infrastructure (V2I) features and identifies the participant's position, speed, and category through a detection network. The traffic participant perception header decodes traffic participant features during the training phase, outputting real-time perception results such as 3D detection boxes, and calculates the dynamic element detection loss Ldet to supervise the model; this header is not activated during inference.

[0097] The prediction network is used to predict the future trajectories of other vehicles or pedestrians based on map and traffic participant features. The network outputs motion prediction features to support risk assessment. The prediction network consists of a prediction submodule and a prediction header. The prediction submodule fuses weighted map features output from a map attention network and traffic participant perception features, predicting the future states of traffic participants using a temporal model (such as LSTM or GNN). The prediction header decodes the prediction features during training, outputs motion prediction results (such as trajectory points), and calculates the motion prediction loss Lmotion for supervision; during inference, only the feature stream is passed to the planning submodule.

[0098] The planning submodule is used to ultimately generate the planned trajectory of the autonomous vehicle. This submodule integrates features from the prediction network output, map features, and traffic participant features, and directly outputs the safe trajectory of the autonomous vehicle through planning algorithms such as reinforcement learning or optimization models, and calculates the planning loss Lplan to optimize driving smoothness and safety.

[0099] The LOSS function in this solution is shown below:

[0100] L = Ldet + Lmap + Lmotion + Lplan + Lpre

[0101] Where Ldet represents the dynamic element detection loss calculated by the traffic participant perception header, Ldet corresponds to Figure 3 The perception results of traffic participants; Lmotion represents the motion prediction loss calculated by the prediction header, and Lmotion corresponds to Figure 3 The prediction result in the image; Lmap represents the map building loss calculated by the map-aware header, Lmap corresponds to Figure 3 The map perception result in the middle; Lplan represents the planning loss calculated by the planning submodule, Lplan corresponds to Figure 3 The vehicle planning trajectory in the network. Lpre represents the loss value of the roadside feature temporal alignment network, which is specifically expressed as Equation (1):

[0102] Lpre = sum ( Wm * | PRE[ L(m,t),Δt) ]- L(m,t+Δt) | 2 (1)

[0103] Where sum represents the weighted summation function; * represents the multiplication sign; | | 2 Represents the L2 norm of a vector, and PRE is a roadside feature spatiotemporal alignment network (e.g., Figure 2(The network structure shown in the dashed box) represents the roadside features at time t from the m-th roadside device, and L(m,t+Δt) represents the roadside features of the m-th roadside device at time t+Δt; Wm represents the weight of the m-th roadside device, where m=1,2,3,4. The model is trained using offline data. During training, observation data at time t+Δt can be obtained for self-supervised training of the module parameters. In this embodiment, the data acquisition interval is 100ms, and Δt is represented by 1 to 5, where these five values ​​represent transmission delays of 100ms, 200ms, 300ms, 400ms, and 500ms, respectively.

[0104] The weights of each device are different. In this embodiment, Wm adopts equation (2):

[0105] Wm =(e -Δs_m ) / sum(e -Δs_m (2)

[0106] Here, Δs_m is the distance between the vehicle and the m-th roadside device; the greater the distance, the lower the weight. Alternatively, Wm can be implemented using existing methods that set a fixed value based on the distance between the vehicle and the m-th roadside device.

[0107] This invention also provides an end-to-end autonomous driving method based on spatiotemporal alignment and integrating vehicle, road, and cloud technologies. The method includes:

[0108] Roadside processing steps: Acquire roadside sensor data and extract roadside features using a roadside perception network. Vehicle-side processing steps: Extract vehicle-side features based on vehicle sensor data, and transform the vehicle-side features into a vehicle-centric BEV space using a vehicle-side BEV feature encoder to obtain spatiotemporally aligned vehicle-side features.

[0109] Spatiotemporal alignment steps: The spatiotemporal alignment network module takes the roadside features n times before the current roadside feature a as the historical roadside feature h1, and predicts the roadside features at the future time (t+Δt) based on the current roadside feature a, the transmission delay display value Δt and the historical roadside feature h1 to compensate for the transmission delay within a preset range. At the same time, the roadside features at the future time (t+Δt) are projected onto the BEV space consistent with the vehicle-side features through the vehicle-side BEV encoder to obtain the spatiotemporally aligned roadside features; the loss value Lpre of the roadside feature temporal alignment network is expressed as Equation (1).

[0110] Feature fusion step: Dynamically fuse spatiotemporally aligned roadside features and spatiotemporally aligned vehicle-side features to output vehicle-road fusion features.

[0111] End-to-end planning steps: Based on vehicle-road fusion features, the prediction planning network module predicts the motion of other vehicles and outputs the planned trajectory of the vehicle itself. During training, the prediction planning network module adopts a multi-task loss function L: L = Ldet + Lmap + Lmotion + Lplan + Lpre; where Ldet represents the dynamic element detection loss calculated by the traffic participant perception header; Lmotion represents the motion prediction loss calculated by the prediction header; Lmap represents the map building loss calculated by the map perception header; and Lplan represents the planning loss calculated by the planning submodule.

[0112] In the above spatiotemporal alignment steps, the historical roadside feature pool is used to store historical roadside features h1 in the form of a queue. The roadside feature temporal alignment network uses a deep learning model F(a, Δt, h1) as input to predict the time-aligned (t+Δt) roadside features.

[0113] In the above feature fusion steps, the vehicle-road feature fusion module is also used to support handling the situation where the number of roadside sensors is inconsistent by using zero-value padding. When the number of roadside sensors in the actual scenario is less than the preset number, zero-value padding is used during fusion.

[0114] To address the issues of current single-vehicle intelligent end-to-end algorithms, this paper utilizes the vehicle-road-cloud integrated end-to-end autonomous driving solution described in the above embodiment, fusing vehicle-side and roadside information within an end-to-end paradigm. Based on the vehicle-road-cloud integrated scenario data from Suzhou High-speed Rail New City, we constructed a model training and testing dataset. The dataset includes 300 scenario data points, with 240 scenarios in the training set and 60 scenarios in the test set, each scenario lasting 20 seconds. Experiments were conducted on this dataset, comparing the single-vehicle intelligent end-to-end solution with the solution of this invention. The end-to-end module uses the same backbone network, and this solution can significantly reduce the perception error rate and improve trajectory planning accuracy. Specific quantitative results are shown below:

[0115]

[0116] Among them, 3DObjectDetection is the accuracy of traffic participant detection, and the higher the accuracy value, the better; boundarydetection is the accuracy of traffic sign detection, and the higher the accuracy value, the better; minADE is the error of the predicted trajectory, and the lower the error value, the better.

[0117] Analysis of the experimental results shows that, compared with the single-vehicle intelligent end-to-end solution, the solution of this invention can effectively solve problems such as long distance, small objects, and severe occlusion, and improve the stability and safety of the autonomous driving system in long-tail scenarios.

[0118] Vehicle-to-everything (V2X) systems can integrate perception information from both the vehicle and the roadside, providing vehicles with a wider perception range and less obstruction, effectively addressing perception challenges in scenarios involving long distances, small objects, and severe occlusion. However, current V2X systems generally employ a post-perception fusion approach, where the vehicle and roadside systems perform perception separately. The roadside system sends its perception results to the vehicle, which then fuses these results with its own perception modules to output the final perception result. This post-fusion approach has several problems. Firstly, it is limited by the accuracy of perception confidence calculation. Secondly, in an end-to-end paradigm, the vehicle no longer has explicit perception results, making post-fusion impossible. Therefore, this approach is fundamentally unsuitable for the end-to-end era.

[0119] This invention not only solves the problem of end-to-end fusion of roadside and vehicle-side information, but also addresses the challenges encountered during its practical implementation:

[0120] 1. Roadside feature transmission delay: There is a 100-500ms time delay in transmitting roadside features to the vehicle. By the time the vehicle acquires the roadside features, they are already 100-500ms behind the vehicle's data. Data fusion needs to consider predicting the current roadside features based on the received roadside features to eliminate the impact of this time delay. Historical information is essential for accurate prediction and needs to be input into the model.

[0121] 2. Inconsistent number of roadside sensors: Some intersections have only one set of sensing equipment, while others have one set of sensing equipment in each of the four directions;

[0122] 3. Vehicle-side computing load issue: More roadside sensors introduce more computing power. In order to achieve real-time inference, it is necessary to make full use of roadside computing equipment and reduce the computing power on the vehicle side.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Those skilled in the art should understand that modifications can be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vehicle-road-cloud integrated end-to-end autonomous driving device based on spatiotemporal alignment, characterized in that, include: Roadside equipment, used to extract roadside features based on a roadside sensing network; The vehicle-side component includes a vehicle-road-cloud integrated end-to-end network, which receives vehicle sensor data and roadside features. The integrated vehicle-road-cloud end-to-end network includes a roadside feature acquisition unit, a vehicle-side feature acquisition unit, a vehicle-road feature fusion module, and a prediction and planning network module. The roadside feature acquisition unit includes a spatiotemporal alignment network module corresponding to the roadside equipment. The spatiotemporal alignment network module includes a historical roadside feature update submodule, a historical roadside feature pool, a roadside feature time alignment network, and a vehicle-side BEV encoder; wherein: the historical roadside feature update submodule is used to take the roadside features of the current roadside feature a before the previous n times as historical roadside features h1, input them into the historical roadside feature pool, and store them in the historical roadside feature pool; the roadside feature time alignment network predicts the roadside features at the future time (t+Δt) based on the current roadside feature a, the transmission delay Δt and the historical roadside features h1, in order to compensate for the transmission delay within a preset range, and at the same time, projects the roadside features at the future time (t+Δt) onto the BEV space consistent with the vehicle-side features through the vehicle-side BEV encoder to obtain spatiotemporally aligned roadside features; the loss value Lpre of the roadside feature time alignment network is expressed as Equation (1): Lpre = sum ( Wm * | PRE[ L(m,t),Δt) ]- L(m,t+Δt) | 2 ) (1) Where, sum represents the weighted summation function, L(m,t) represents the feature prediction value from the m-th roadside device at time t, L(m,t+Δt) represents the feature value of the m-th roadside device at time (t+Δt), PRE is the roadside feature spatiotemporal alignment network, Wm represents the weight of the m-th roadside device; * represents multiplication, | | 2 Represents the L2 norm of a vector; The vehicle-side feature acquisition unit is used to extract vehicle-side features based on vehicle sensor data. The vehicle-side BEV feature encoder transforms the vehicle-side features into a vehicle-centric BEV space to obtain spatiotemporally aligned vehicle-side features. The vehicle-road feature fusion module is used to dynamically fuse spatiotemporally aligned roadside features and spatiotemporally aligned vehicle-end features to output vehicle-road fused features; The predictive planning network module predicts the motion of other vehicles based on vehicle-road fusion features and outputs the vehicle's planned trajectory. During training, the predictive planning network module uses a multi-task loss function L: L = Ldet + Lmap + Lmotion + Lplan + Lpre; where Ldet represents the dynamic element detection loss calculated by the traffic participant perception header of the predictive planning network module; Lmotion represents the motion prediction loss calculated by the prediction header of the predictive planning network module; Lmap represents the map building loss calculated by the map perception header of the predictive planning network module; and Lplan represents the planning loss calculated by the planning submodule of the predictive planning network module.

2. The vehicle-road-cloud integrated end-to-end autonomous driving device based on spatiotemporal alignment as described in claim 1, characterized in that, Wm is obtained using equation (2): Wm =(e -Δs_m ) / sum(e -Δs_m ) (2) Where Δs_m is the distance between the vehicle and the m-th roadside device.

3. The vehicle-road-cloud integrated end-to-end autonomous driving device based on spatiotemporal alignment as described in claim 1 or 2, characterized in that, The roadside feature time alignment network uses a deep learning model F(a, Δt, h1) as input, taking the current roadside feature a, the transmission delay Δt, and the historical roadside feature h1 as input, and outputs the predicted roadside feature value for the future time (t+Δt) after time alignment.

4. The vehicle-road-cloud integrated end-to-end autonomous driving device based on spatiotemporal alignment as described in claim 3, characterized in that, The vehicle-road feature fusion module is also used to support handling situations where the number of roadside sensors is inconsistent through zero-value padding. When the number of roadside sensors in the actual scenario is less than the preset number, zero-value padding is used during fusion.

5. The vehicle-road-cloud integrated end-to-end autonomous driving device based on spatiotemporal alignment as described in claim 3, characterized in that, The predictive planning network module includes: The historical vehicle-road fusion feature update submodule is used to manage the time series of historical vehicle-road fusion features; The historical vehicle-road fusion feature attention submodule, based on multi-frame data in the historical vehicle-road fusion feature pool, analyzes temporal dependencies through an attention network, assigns weights to vehicle-road fusion features at different times, and calculates weighted historical vehicle-road fusion features; The historical vehicle-road fusion feature network submodule is used to fuse weighted historical vehicle-road fusion features with current vehicle-road fusion features to generate enhanced vehicle-road fusion features; The map network includes a map perception submodule, a map perception header, and a map attention network. The map perception submodule is used to extract map-related features from the enhanced vehicle-road fusion features and identify the spatial distribution of map elements. The map perception header is used to decode map features during the model training phase, output real-time map perception results, and calculate Lmap to supervise the model. The map attention network is used to apply an attention mechanism to map features, filter key map information, and output the results. The Traffic Participant Network is used to perceive and encode the state of traffic participants. It includes a Traffic Participant Perception Submodule and a Traffic Participant Perception Header. The Traffic Participant Perception Submodule is used to extract traffic participant features from vehicle-road fusion features and identify the position, speed and category of participants through a detection network. The Traffic Participant Perception Header is used to decode traffic participant features during the training phase, output real-time perception results, and calculate Ldet to supervise the model. The prediction network is used to predict the future trajectories of other vehicles or pedestrians based on map and traffic participant features. It includes a prediction submodule and a prediction header. The prediction submodule is used to fuse the weighted map features output by the map attention network and the traffic participant perception features, and predict the future state of the traffic participants. The prediction header is used to decode the prediction features during the training phase, output the motion prediction results, and calculate Lmotion for supervision. The planning submodule integrates features from the prediction network output, map features, and traffic participant features. It then uses a planning algorithm to generate the vehicle's planned trajectory and calculates Lplan to optimize driving smoothness and safety.

6. A vehicle-road-cloud integrated end-to-end autonomous driving method based on spatiotemporal alignment, characterized in that, include: Roadside processing steps: Acquire roadside sensor data and extract roadside features using a roadside sensing network; Vehicle-side processing steps: Extract vehicle-side features based on vehicle sensor data, and transform the vehicle-side features into a vehicle-centric BEV space through a vehicle-side BEV feature encoder to obtain spatiotemporally aligned vehicle-side features; Spatiotemporal alignment steps: The spatiotemporal alignment network module takes the roadside features n times before the current roadside feature a as the historical roadside feature h1, and predicts the roadside features at the future time (t+Δt) based on the current roadside feature a, the transmission delay Δt and the historical roadside feature h1 to compensate for the transmission delay within a preset range. At the same time, the roadside features at the future time (t+Δt) are projected onto the BEV space consistent with the vehicle-side features through the vehicle-side BEV encoder to obtain the spatiotemporally aligned roadside features. The loss value Lpre of the roadside feature time-aligned network is expressed as Equation (1): sum ( Wm * | PRE[ L(m,t),Δt) ]- L(m,t+Δt) | 2 ) (1) Where, sum represents the weighted summation function, L(m,t) represents the feature prediction value from the m-th roadside device at time t, L(m,t+Δt) represents the feature value of the m-th roadside device at time (t+Δt), PRE is the roadside feature spatiotemporal alignment network, Wm represents the weight of the m-th roadside device; * represents multiplication, | | 2 Represents the L2 norm of a vector; Feature fusion steps: Dynamically fuse spatiotemporally aligned roadside features and spatiotemporally aligned vehicle-side features to output vehicle-road fusion features; End-to-end planning steps: Based on vehicle-road fusion features, the prediction planning network module predicts the motion of other vehicles and outputs the planned trajectory of the vehicle itself. During training, the prediction planning network module adopts a multi-task loss function L: L = Ldet + Lmap + Lmotion + Lplan + Lpre; where Ldet represents the dynamic element detection loss calculated by the traffic participant perception header; Lmotion represents the motion prediction loss calculated by the prediction header; Lmap represents the map building loss calculated by the map perception header; and Lplan represents the planning loss calculated by the planning submodule.

7. The end-to-end autonomous driving method based on spatiotemporal alignment for vehicle-road-cloud integration as described in claim 6, characterized in that, Wm is obtained using equation (2): Wm =(e -Δs_m ) / sum(e -Δs_m ) (2) Where Δs_m is the distance between the vehicle and the m-th roadside device.

8. The end-to-end autonomous driving method based on spatiotemporal alignment for vehicle-road-cloud integration as described in claim 6 or 7, characterized in that, In the spatiotemporal alignment step, the historical roadside feature pool is used to store historical roadside features h1 in the form of a queue. The roadside feature temporal alignment network uses the deep learning model F(a, Δt, h1) as input, the current roadside feature a, the transmission delay Δt and the historical roadside feature h1, and outputs the predicted value of the roadside feature at the future time (t+Δt) after temporal alignment.

9. The end-to-end autonomous driving method based on spatiotemporal alignment for vehicle-road-cloud integration as described in claim 8, characterized in that, In the feature fusion step, the vehicle-road feature fusion module is also used to support handling the situation where the number of roadside sensors is inconsistent through zero-value padding. When the number of roadside sensors in the actual scenario is less than the preset number, zero-value padding is used during fusion.

10. The end-to-end autonomous driving method based on spatiotemporal alignment for vehicle-road-cloud integration as described in claim 8, characterized in that, In the end-to-end planning process, the predictive planning network module includes: The historical vehicle-road fusion feature update submodule is used to manage the time series of historical vehicle-road fusion features; The historical vehicle-road fusion feature attention submodule, based on multi-frame data in the historical vehicle-road fusion feature pool, analyzes temporal dependencies through an attention network, assigns weights to vehicle-road fusion features at different times, and calculates weighted historical vehicle-road fusion features; The historical vehicle-road fusion feature network submodule is used to fuse weighted historical vehicle-road fusion features with current vehicle-road fusion features to generate enhanced vehicle-road fusion features; The map network includes a map perception submodule, a map perception header, and a map attention network. The map perception submodule is used to extract map-related features from the enhanced vehicle-road fusion features and identify the spatial distribution of map elements. The map perception header is used to decode map features during the model training phase, output real-time map perception results, and calculate Lmap to supervise the model. The map attention network is used to apply an attention mechanism to map features, filter key map information, and output the results. The Traffic Participant Network is used to perceive and encode the state of traffic participants. It includes a Traffic Participant Perception Submodule and a Traffic Participant Perception Header. The Traffic Participant Perception Submodule is used to extract traffic participant features from vehicle-road fusion features and identify the position, speed and category of participants through a detection network. The Traffic Participant Perception Header is used to decode traffic participant features during the training phase, output real-time perception results, and calculate Ldet to supervise the model. The prediction network is used to predict the future trajectories of other vehicles or pedestrians based on map and traffic participant features. It includes a prediction submodule and a prediction header. The prediction submodule is used to fuse the weighted map features output by the map attention network and the traffic participant perception features, and predict the future state of the traffic participants. The prediction header is used to decode the prediction features during the training phase, output the motion prediction results, and calculate Lmotion for supervision. The planning submodule integrates features from the prediction network output, map features, and traffic participant features. It then uses a planning algorithm to generate the vehicle's planned trajectory and calculates Lplan to optimize driving smoothness and safety.

Citation Information

Patent Citations

  • Feature-result level fusion vehicle infrastructure cooperative sensing method, medium and electronic equipment

    CN116958763A

  • Autonomous vehicle road cloud fusion sensing method

    CN117111085A