Automatic driving method based on modal fusion and Bessel optimization
By adopting modal fusion and Bessel optimization methods in the autonomous driving system, the limitations of the existing system in multimodal perception fusion and motion planning are solved, and a safer and more comfortable autonomous driving effect in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510525457.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing end-to-end autonomous driving system has limitations in multimodal perception fusion and motion planning, resulting in insufficient global three-dimensional context modeling, high real-time computing load, high error propagation risk, and surge in violation rate in complex scenarios, making it difficult to achieve joint space-time optimization.
The autonomous driving method based on modal fusion and Bezier optimization is adopted to synchronize environmental information through a multi-sensor system, image and point cloud features are extracted using forward-view cameras and lidar systems, and trajectory prediction is realized through cross-modal feature fusion, feature refinement and Bezier curve decoder to ensure the smoothness and kinematic feasibility of the trajectory.
The vehicle collision rate is reduced by 52% and pedestrian collisions in dense urban market scenarios, ensuring the C2-continuity and smoothness of the trajectory, improving riding comfort, and enhancing the environmental perception ability and interpretability of the neural network architecture.
Smart Images

Figure CN120071303A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to an autonomous driving method based on modality fusion and Bezier optimization. Background Art
[0002] The end-to-end autonomous driving system provides a paradigm-shifting alternative to traditional modular architectures by establishing a direct mapping from raw sensor inputs to planned trajectories or control signals. By integrating the traditional perception-planning-control process into a unified neural framework, it fundamentally eliminates the error-prone intermediate representations and cascading failures inherent in sequential processing chains - especially coordinate system misalignment artifacts and cumulative delay effects. Additionally, the simplicity of this end-to-end architecture conceptually bypasses the need for handcrafted rules while enabling joint optimization of environmental understanding and motion generation.
[0003] Existing end-to-end autonomous driving systems have significant limitations in multi-modal perception fusion. Although recent research has verified the complementarity of multi-source data through early, mid-term, and late fusion of camera and depth modalities, and has also demonstrated the effectiveness of semantic and depth information as intermediate representations, mainstream methods are still limited to one-way cross-modal projections of multi-view lidar representations such as bird's-eye view (BEV) and front view (RV), resulting in insufficient global three-dimensional context modeling and making it difficult to ensure safe decision-making in dense traffic flows. Even with the improvement of interpretability through safety-constrained action generation, its complex cross-modal attention mechanism and multi-stage feature transformation still face real-time computational load and error propagation risks, and recent experiments have further revealed the defect of the sharp increase in violation rates of existing frameworks in complex scenarios, highlighting the spatio-temporal mismatch problem between perceptual representation and dynamic scene understanding.
[0004] In the field of motion planning, although early end-to-end methods have verified the feasibility of latent space trajectory generation, they have caused instability in optimization and an interpretability crisis due to the opacity of the decision-making process. Subsequent hybrid architectures based on explicit cost maps have improved interpretability through handcrafted rules, but their construction of high-dimensional spatial grids has led to a sharp increase in computational overhead and their generalization ability is weak in unknown traffic configurations. Some methods, such as those that improve planning constraint satisfaction through vectorized scene encoding and object-centered representation, still split prediction and planning into sequential tasks, failing to achieve spatio-temporal joint optimization and ignoring the inherent geometric dynamics coupling between the two, resulting in limited trajectory smoothness and obstacle avoidance ability in dense traffic flows.
[0005] Deeper system-level defects stem from the architectural fragmentation of the perception-planning link. In existing multimodal fusion frameworks, although perception is enhanced through cross-sensor feature mapping, their hierarchical processing flow leads to the accumulation of modal alignment errors, and discrete planning representations make it difficult to ensure kinematic feasibility and comfort. Even if perspective-invariant representations can be constructed through self-supervised learning, there are still efficiency bottlenecks in their multi-stage feature propagation. Although semantic point cloud fusion schemes optimize the representation of driving tasks, they do not solve the curvature continuity constraint problem of trajectory parameterization. These limitations together result in an essential bottleneck in the existing systems' ability to handle environmental uncertainties, maintain spatio-temporal consistency, and ensure the compliance of motion planning. There is an urgent need for a tightly coupled perception-planning co-optimization framework. Summary of the Invention
[0006] An object of an embodiment of the present invention is to provide an autonomous driving method based on modal fusion and Bezier optimization, aiming to solve the problems raised in the above-mentioned background technology.
[0007] The embodiment of the present invention is implemented as follows. An autonomous driving method based on modal fusion and Bezier optimization includes the following steps: Step 1: Data input and modal feature extraction; Use the in-vehicle multi-sensor system to synchronously collect environmental information. Specifically, use a front-view camera to capture high-resolution two-dimensional image data. At the same time, use a lidar system to obtain three-dimensional point cloud data of the environment. After obtaining the original data, parallelly send these two heterogeneous data streams into their respective dedicated feature extraction backbone networks. For image data, gradually abstract image feature maps with hierarchical semantics from the original pixels, and finally output preliminary image features ; for point cloud data, capture the local geometric details and global spatial distribution of the point cloud data to generate preliminary lidar features ; Step 2: Cross-modal feature fusion; After obtaining the preliminary single-modal features, take and as inputs and send them to the cross-modal feature fusion module. Through the multi-head cross-attention mechanism, interact and aggregate the information of the two to output cross-modal fusion features ; Step 3: Feature refinement and dedicated feature generation; To adapt to the requirements of different downstream tasks, process and differentiate the features: on the one hand, perform upsampling and interpolation operations on the preliminary image features to obtain restored image features with more optimized resolution or channel dimensions ; on the other hand, also perform upsampling interpolation processing on the cross-modal fusion features to generate general fusion features .
[0008] Based on the general fusion features , two key feature representations are derived; firstly, through multi-layer convolution, downsampling, and upsampling operations, is transformed into the bird's-eye view (BEV) space to obtain the bird's-eye view features ; secondly, in order to capture the spatio-temporal dynamics of the scene and combine vehicle intentions, the general fusion features are integrated with dynamic path queries, real-time ego-vehicle state information (such as speed, acceleration), and navigation information, and enhanced using the spatio-temporal attention mechanism, finally generating spatio-temporal features containing temporal and context information , which are specifically used for the trajectory prediction task; Step 4: Multi-task decoding and loss function; Finally, using multi-level features, namely , and , the model drives multiple decoders in parallel to complete different perception and prediction tasks; specifically including: firstly, using for depth map decoding to estimate the scene pixel-level depth, and supervised using the L1 loss; secondly, using for image semantic segmentation decoding to assign semantic classes to each pixel in the image, and supervised using the cross-entropy loss; thirdly, using input to the Bezier decoder to predict the control points of the Bezier curve, and then construct a smooth Bezier curve, and obtain the final trajectory prediction result through curve sampling, and this task is optimized using a composite loss function, including reconstruction loss, regularization constraint, smoothness constraint, and C² continuity constraint, to ensure the accuracy and feasibility of the trajectory; fourthly, using for semantic segmentation decoding in the BEV perspective to identify road structures and drivable areas in the bird's-eye view, etc., and supervised using the cross-entropy loss; fifthly, using for object 3D bounding box decoding to detect and locate objects in three-dimensional space, and jointly supervised using the L1 loss, cross-entropy loss, and custom loss.
[0009] The beneficial effects of an autonomous driving method based on modality fusion and Bezier optimization provided by an embodiment of the present invention are as follows: (1) Unify cross-modal attention and Bezier-based trajectory optimization, enabling the end-to-end framework to achieve a 52% reduction in vehicle collision rate and no pedestrian collisions in dense urban scenarios; (2) The Bezier decoder ensures C 2-Continuity, with inherent smoothness and kinematic feasibility, greatly reduces acceleration during driving and improves ride comfort; (3) Established a spatio-temporal attention mechanism to enhance trajectory planning in an end-to-end autonomous driving system, enhancing the performance of trajectory planning in both the time and space dimensions and demonstrating top-notch results in simulations; (4) Achieved unified multi-task perception and supervision, significantly enhancing the environmental perception ability and the interpretability of the neural network architecture, while bridging the semantic-geometry gap in scene cognition. Description of the Drawings
[0010] Figure 1 Flowchart of an autonomous driving method based on modal fusion and Bezier optimization provided by an embodiment of the present invention; Figure 2 Schematic diagram of the autonomous driving architecture in an autonomous driving method based on modal fusion and Bezier optimization provided by an embodiment of the present invention; Figure 3 Visualization of the Bezier-based trajectory generation framework; Figure 4 Comparison of trajectory point distributions of this method in the Longest6 benchmark test (top row) and TF++ (bottom row); where (a) and (b) are right-angle turns at a double-lane T-junction, (c) is a right-angle turn at a multi-lane T-junction, (d) is a lane change within a roundabout, and (e) is a lane change on a straight road. Detailed Implementation Modes
[0011] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0012] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.
[0013] As Figure 1 shown, an autonomous driving method based on modal fusion and Bezier optimization provided by an embodiment of the present invention aims to comprehensively understand and accurately predict the driving environment by fusing multi-modal perception information from cameras and lidar. The method includes the following steps: Step 1: Data input and modal feature extraction; The starting point of this method is to synchronously collect environmental information using an in-vehicle multi-sensor system. Specifically, a forward-facing camera with a 110° field of view is used to capture high-resolution two-dimensional image data. At the same time, a 16-line Light Detection and Ranging (LiDAR) system is utilized to obtain three-dimensional point cloud data of the environment, providing precise geometric structures and depth information. After obtaining the raw data, the system feeds these two heterogeneous data streams into their respective dedicated feature extraction backbone networks in parallel. For the image data, a pre-trained RegnetY-3.2GF neural network with excellent performance is loaded as the image feature extraction module. This module abstracts hierarchical semantic image feature maps layer by layer from the original pixels through a series of convolutional, pooling, and non-linear activation operations, and finally outputs preliminary image features , which is a feature vector. For the point cloud data, the LiDAR feature extraction module uses an initialized RegnetY-3.2GF neural network. These networks are designed to effectively capture the local geometric details and global spatial distributions of the point cloud data, generating preliminary LiDAR features , which is in the form of a feature vector.
[0014] Step 2: Cross-modal feature fusion; After obtaining the preliminary single-modal features, in order to effectively integrate complementary information, a cross-modal feature fusion module is designed; this module receives and as inputs, and interacts and aggregates the information of the two through a multi-head cross-attention mechanism, outputting cross-modal fusion features containing rich information .
[0015] Step 3: Feature refinement and dedicated feature generation; To adapt to the requirements of different downstream tasks, the features are further processed and differentiated: on the one hand, the preliminary image features are subjected to upsampling and interpolation operations to obtain restored image features with more optimized resolution or channel dimensions , which are mainly used for perception tasks based on the image perspective; on the other hand, the cross-modal fusion features are also subjected to similar upsampling and interpolation processes to generate general fusion features .
[0016] Based on the general fusion features , two key feature representations are derived. Firstly, through operations such as multi-layer convolution, downsampling, and upsampling (similar to the common Bird's Eye View (BEV) feature generation network structure), is transformed into the Bird's Eye View (BEV) space to obtain BEV features , this feature provides a basis for subsequent perception tasks from the BEV perspective. Second, to capture the spatio-temporal dynamics of the scene and incorporate vehicle intentions, the general fusion feature is integrated with dynamic path queries, real-time ego-vehicle state information (such as speed, acceleration), and navigation information, and enhanced using spatio-temporal attention mechanisms, finally generating spatio-temporal features containing temporal and context information , which are specifically used for the trajectory prediction task.
[0017] Step 4: Multi-task decoding and loss function; Finally, using multi-level features ( , and ), the model drives multiple decoders in parallel to complete different perception and prediction tasks. Specifically, first, use for depth map decoding to estimate the scene pixel-level depth, and use L1 loss for supervision; second, use for image semantic segmentation decoding to assign semantic categories to each pixel in the image, and use cross-entropy loss for supervision; third, use input to the Bezier decoder to predict the control points of the Bezier curve, and then construct a smooth Bezier curve, and obtain the final trajectory prediction result through curve sampling. This task is optimized using a composite loss function, including reconstruction loss, regularization constraint, smoothness constraint, and C² continuity constraint, to ensure the accuracy and feasibility of the trajectory; fourth, use for semantic segmentation decoding from the BEV perspective to identify road structures, drivable areas, etc. in the bird's-eye view, and use cross-entropy loss for supervision; fifth, use for 3D object bounding box decoding to detect and locate objects in three-dimensional space, and use a combination of L1 loss, cross-entropy loss, and custom loss for joint supervision. Through this multi-task, end-to-end learning framework, this method can make full use of multi-modal data to achieve comprehensive perception and dynamic prediction of complex traffic scenarios.
[0018] In the embodiment of the present invention, the proposed autonomous driving architecture is specifically as Figure 2 shown, with three-dimensional perception integrating a front-view RGB camera and lidar, including: (1) Dual-branch feature embedding with cross-modal attention fusion; (2) Auxiliary task branch and spatio-temporal enhanced trajectory decoding; (3) Joint optimization through multi-constraint supervision to achieve safe perception trajectory generation.
[0019] As Figure 2As shown, as a preferred embodiment of the present invention, in the step 1, first, the RegnetY-3.2GF neural network backbone is used to extract preliminary features from the images and lidar point cloud data obtained from vehicle-mounted sensors respectively and . After channel alignment, these features ( and ) are connected and enhanced through position encoding to form embedded features :
[0020] Among them, is the preliminary image feature, is the preliminary lidar feature, the position encoding is , is the embedded feature
[0021] As a preferred embodiment of the present invention, in the step 2, the fusion process is decomposed into a dual complementary path, where linear projection is used to calculate a set of queries, keys, and values:
[0022]
[0023]
[0024]
[0025] Among them, represents the position function, projects the LiDAR feature into the BEV grid, realizes the perspective-to-BEV conversion through depth perception feature enhancement, is the query of the preliminary image feature weight, is the query of the preliminary lidar feature weight, is key and value weight, is key and value weight; this dual-query formula realizes two collaborative attention flows, namely and : First, focuses on the lidar-derived key / value , injecting precise spatial priors into visual features; second, interacts with the RGB-based key / value to enrich the point cloud features with contextual semantics:
[0026]
[0027] where is the activation function, is the feature dimension of each attention head, is the number of attention heads; Then, two attention streams are combined through gated fusion:
[0028] where represents the learnable gating parameter, is the sigmoid activation, represents element-wise matrix multiplication.
[0029] As Figure 2 shown, as a preferred embodiment of the present invention, in step 3, the is reduced in image dimension through BilinearUp interpolation to obtain the reduced image feature and the is reduced in BEV dimension through TrilinearUp interpolation to obtain the general fusion feature :
[0030]
[0031] where represents the parameterized deformation grid, represents the interpolation weight, represents the sampling coordinates, encapsulates the time offset, time mixing weight and the relative timestamp , is the weight parameter of, is for at the t-th time step.
[0032] Then, 3D reconstruction is performed through convolution (Conv) and upsampling (Upsample) to obtain the bird's-eye view feature :
[0033] Among them, represents depth convolution, represents pointwise convolution, represents the index of the network layer in this structure, is the output of the -th layer network, is the convolution function, is the upsampling function. Before using transposed convolution for upsampling and skipping connections from the corresponding encoder stage, the architecture gradually downsamples the features through strided convolution.
[0034] Restore image features Decode and bird's-eye view features After parameter adjustment through multimodal prediction error in the subsequent decoding, encapsulates implicit and precise 3D scene perception features. Innovatively define a set of learnable parameters , among which represents the number of historical prediction trajectories, represents the number of time steps of the current trajectory planning, represents the dimension of the GRU cell. First, apply temporal self-attention to a set of learnable query parameters (parameters that can be automatically adjusted and optimized according to the reverse gradient during model training, similar to the weight parameters in the model network):
[0035] Among them, is the historical feature query. This parameter is still a learnable query parameter, but due to the self-attention effect, a superscript is added for easy distinction from the previous learnable parameters; is the multi-head self-attention function; Next, incorporate the ego vehicle speed into the observation state. The fused feature undergoes channel transformation, position encoding, and feature flattening operations, and then is concatenated with the encoded ego state to form the enhanced feature :
[0036] Among them, is an unfolding function (or data unfolding function) used to unfold multi-dimensional matrix data into a one-dimensional vector; is a multi-layer perceptron composed of multiple linear layers and non-linear activation functions; represents the concatenated vector; is a position vector obtained by position encoding and used to add position information to the features.
[0037] Subsequently, spatio-cross attention is performed between and to generate spatio-temporal features :
[0038] wherein, is the multi-head cross-attention function.
[0039] As a preferred embodiment of the present invention, in step 4, and are decoded into corresponding multi-modal outputs. These outputs supplement the core trajectory prediction branch, utilize for spatio-temporal feature interaction, and achieve trajectory generation based on the Bézier curve; specifically including the following steps: Step 4.1: Restore image features Decoding: The LiDAR interaction information is merged into , and its spatial features are enhanced through feature complementarity. The depth and semantic branches are respectively derived from for decoding the depth map and the semantic segmentation map. Through this process, the perception of the ego vehicle for the surrounding distance and object categories is enhanced, thereby improving the trajectory planning ability of the core branch. Specifically, it is transformed through a cascaded structure of multi-level deconvolution and interpolation upsampling . The N-level bilinear interpolation doubles the spatial resolution, facilitating the reconstruction of spatial information from low-dimensional features to high-resolution output:
[0040]
[0041]
[0042]
[0043] wherein, , respectively represent the sampling parameters of the th layer in the corresponding formula. As Figure 1 shows, the depth estimation branch and the semantic segmentation branch adopt the same multi-level architecture, but maintain independent transposed convolution and interpolation sampling parameters and . This design enables domain-adaptive feature learning for these two tasks, while keeping the output images and at the same spatial resolution; Step 4.2: Aerial view features Decode; Compared with the aerial view features go through multiple convolutional and upsampling layers, thus generating more distinct spatial and semantic context features. Using to perform BEV perspective BevSemantic and BoundingBox predictions, it supplements the front view depth and semantic segmentation maps derived from to enhance the model's 3D information and semantic perception of the global environment. For the BevSemantic branch, given that has gone through multiple layers of convolution and upsampling, a lightweight feature transformation architecture is adopted to map the spatial features to semantic categories:
[0044] wherein, represents the convolution function; The BoundingBox branch, as a pure 3D detection task, adopts a multi-branch parallel independent task head design, including the heatmap of the target center point , the size of the bounding box , the center offset and for the joint prediction of the yaw angle :
[0045] where represents each prediction head, (set to 15) is the number of discrete direction classes for rough direction estimation.
[0046] Step 4.3: Bézier curve decoding; The Bézier branch is the core component of the present invention. It decodes the spatio-temporal features to obtain the Bézier curve control points, thereby obtaining the Bézier curve, and then samples to obtain the planned trajectory.
[0047] The scenario where this method is used is a navigation scenario with discrete target points , and these discrete target points are usually separated by at least 50 m. The target point information is embedded into the spatio-temporal features to guide the trajectory planning. Multiple GRU units sequentially predict the control points of the Bézier curve :
[0048] wherein, , represents the offset of the th layer GRU based on , The offset output serves as the control point of the final output trajectory. is the total number of GRUcell layers and the number of curve control points. This work uses cubic Bézier curves for trajectory planning, balancing lightweight calculations with smooth dynamics. The continuous curve B is derived from the control points , and trajectory points are obtained through sampling :
[0049] Among them, represents the equidistant sampling function, represents the position parameter on the Bézier curve.
[0050] This method ensures a smooth and dynamically feasible trajectory, guaranteeing continuity while maintaining computational efficiency. After decoding to obtain the output, a dedicated loss function is tailored for each perception and planning component. The supervision mechanism includes two categories: First, auxiliary task supervision , which is derived from the BEV feature decoder and the image feature decoder; Second, trajectory planning constraints , which are physical constraints applied to the Bézier branch. Specifically: (1) Auxiliary task supervision , which specifically includes: depth estimation supervision of the depth map in the depth estimation branch , semantic category supervision of the semantic segmentation branch , semantic category supervision of the BEV semantic branch segmentation and supervision of various heads in the BoundingBox branch .
[0051] Depth estimation supervision : The depth estimation branch is supervised using the L1 loss to ensure the accuracy of the depth information metric for 3D scene reconstruction:
[0052] Among them, and represent the predicted depth map and the ground truth depth map respectively.
[0053] Semantic segmentation supervision: The front view (Iisem) and BEV (Ibsem) semantic segmentation tasks use cross-entropy loss on categories (aligned with the CARLA simulator categories, including pedestrians, vehicles, roads, and traffic lights):
[0054] Among them, is the ground truth in the form of one-hot vectors; BoundingBox Branch Supervision: The detection branch combines multiple supervision objectives: 1) the height / width / center point / center point offset / yaw angle residual of the object: supervised by the L1 loss of continuous parameters; 2) yaw angle classification: cross-entropy loss of discrete yaw bins; 3) Heatmap prediction: improved focal loss supervised by Gaussian encoding;
[0055] where the hyperparameters and are used to adjust the balance between positive / negative samples and hard / easy samples respectively, while is used to exclude the supervision calculation of traffic participants outside the predefined boundary range; corresponds to the supervision signal of each prediction head; the Heatmap ground truth is generated by a Gaussian kernel procedure centered on the annotated target position.
[0056] The final auxiliary task loss is expressed as a weighted sum of these component loss terms:
[0057] (2) Trajectory Planning Constraints , the Bézier trajectory prediction branch is optimized by a composite loss as Figure 3 shown. This loss implicitly obtains the target point guidance and safety constraints through the L1 reconstruction loss, and strengthens the kinematic feasibility of the vehicle through regularization constraints, curvature continuity constraints, and smooth control point constraints to enhance the trajectory smoothness and feasibility.
[0058] Reconstruction Loss : Ensures geometric alignment with the ground truth trajectory to ensure lower performance limitations;
[0059] where represents the number of prediction time intervals, represents the final predicted trajectory points, represents the ground truth trajectory points.
[0060] Control Point Regularization : Guarantees the kinematic prior by constraining the alignment of the initial control points with the ego vehicle speed :
[0061] Smoothness Constraint : Constrain the consistency of the spacing between adjacent control points to avoid sudden trajectory jitter to ensure driving comfort:
[0062] Among them, The threshold determined from vehicle power.
[0063] Curvature continuity : Maintain the continuity of the second derivative through control point constraints to ensure the continuity of the trajectory curvature change rate, so that the predicted trajectory conforms to the dynamic characteristics of the vehicle:
[0064] Trajectory planning constraints Combine these components with empirical weights:
[0065] Set weight parameters , , and , to ensure the vehicle power feasibility and ride comfort under the appropriate performance lower limit.
[0066] As a preferred embodiment of the present invention, the learnable parameter path Queries of the spatio-temporal attention of the present invention retains 6 historical time steps, and at the same time predicts a 2-second future trajectory sampled at 250 millisecond intervals. The architecture specifications include: 1) Image encoding of RegNetY-3.2GF through pre-trained weights; 2) LiDAR flow processing through RegNetY-3.2GF initialized from scratch; 3) Cross-modal attention with 4 heads and a 0.1 dropout probability; 4) Multi-scale BEV reconstruction using sequential 1×1→3×3→3×3 convolutional kernels; 5) Bézier decoder implementing an 8-head self / cross attention mechanism with 6 layers of stacked transformations; 6) GRU units with 256-dimensional inputs and 64-dimensional hidden states, and the output dimension is aligned with the Bézier control point count; 7) Unified 3×3→1×1 convolutional pattern for BEV semantics and bounding box branches. The training scheme uses AdamW for three-phase optimization for all branches: Phase 1, 30 epochs of pre-training (lr = 3e-4); Phase 2, 30 epochs of core branch refinement (lr = 1e-4); Phase 3, 15 epochs of joint fine-tuning (lr = 3e-5) to enhance cross-branch collaboration. The implementation occurs on a dual RTX 3090 GPU with a batch size of 16 and takes approximately 180 training hours.
[0067] The evaluation framework of the present invention includes two benchmarks: the CARLA benchmark, which is used to evaluate driving performance indicators such as driving scores, violation scores and collision rates in various CARLA leaderboard route scenarios; and the curve benchmark, which is used to evaluate the trajectory smoothness, ride comfort and kinematic feasibility of the vehicle on the CARLA leaderboard route.
[0068] CARLA Benchmarks: The Town05 long / short benchmarks and the Longest6 benchmarks were thoroughly tested using the industry-standard CARLA leaderboard evaluation protocol, with a focus on the more comprehensive Longest6 benchmark proposed by TransFuser. Longest6 is derived from the CARLA leaderboard 1.0 routes, aggregating the 6 longest trajectories in Towns 01-006 (average length: 1.5 km). Each route has a unique environmental configuration, combining 6 weather conditions (cloudy; wet; moderate rain; wet cloudy; heavy rain; light rain) and 6 lighting states (night; twilight, the time before sunrise and after sunset when the sun illuminates the sky but is not visible); dawn; morning; noon; sunset), enhanced with adversarial traffic scenarios. The Town05 benchmark contains two configurations: 1) Town05Short: 10 routes (100-500 m) with 3 each; 2) Town05 Long: 10 extended routes (1000-2000m) with 10 intersections. The benchmark challenges the system with different road geometries (multi-lane arteries, single-lane roads, bridges, highways) and dynamic agent interactions.
[0069] Curved Benchmarks: For trajectory planning evaluation, Routes 13, 23, and 25 were specifically selected from the Longest6 benchmark due to their challenging geometric configurations—each containing multiple curved sections, roundabouts, and S-ends, rigorously testing trajectory smoothness, dynamic feasibility, and passenger comfort.
[0070] The CARLA benchmark uses ten key indicators for evaluation: 1) Driving score, which is the product of route completion and violation penalty; 2) Route completion, which is the percentage of route distance completed; 3) Violation score, which is a composite score aggregated from multiple violation types; 4) Pedestrian collision, which is the collision rate with pedestrians; 5) Vehicle collision, which is the collision rate with vehicles; 6) Static object collision, which is the collision rate of static elements; 7) Red light violation, which is the probability of traffic light violation; 8) Road deviation, which is the percentage of deviation from the specified navigation route; 9) Route timeout, which is the percentage of routes that exceed the maximum allowed time; 10) Vehicle stuck, which is the probability of extended agent inactivity during driving. The specific evaluation results of the Longest6 benchmark are shown in Table 1.
[0071] Table 1 Longest6 benchmark results
[0072] In the Longest6 benchmark, this method achieved a state-of-the-art driving score (DS) of 70, exceeding TF++ by 1%, while reducing the vehicle collision rate by 52% (0.31 vs. 0.83). Notably, it maintained zero pedestrian collisions and eliminated the infraction scores (IS) for red-light violations and blocking scenarios, demonstrating robust safety awareness in dense urban environments.
[0073] The specific evaluation results for the Town05 benchmark are shown in Table 2. For the Town05 benchmark, this method established new performance records: 83.75 DS (+9.05% higher than the previous state-of-the-art) and 98.15% route completion on long routes, with a completion rate of 98.50% on short routes. This demonstrates the special ability of this method in handling complex road geometries and extended navigation challenges. The results particularly emphasize the effectiveness of the spatio-temporal memory bank of this method in maintaining the consistency of extended trajectories.
[0074] Table 2 Town05 Benchmark Test Results
[0075] To quantify the trajectory quality, four evaluation criteria were established: 1) Lateral deviation and standard deviation from the expert trajectory to evaluate path tracking accuracy; 2) Average curvature and standard deviation of the ego-vehicle trajectory to measure smoothness; 3) Kinematic feasibility analysis based on the curvature estimation model proposed in "Vehicle Dynamics and Control":
[0076] where, represents the yaw rate, the vehicle speed, the front-wheel angle, is the wheelbase; 4) Jerk (time derivative of acceleration) statistic, which is directly related to passenger comfort.
[0077] Table 3 Route 13 Test Results of the Curve Benchmark
[0078] Table 4 Route 23 Test Results of the Curve Benchmark
[0079] Table 5 Route 25 Test Results of the Curve Benchmark
[0080] The quantitative analysis of Tables 3, 4, and 5 for the curve benchmark reveals a fundamental improvement in trajectory quality.Figure 4 Visual comparison of the trajectory distribution between the representative scenarios of the present invention and TF++ is provided to support the tabular results. Compared with the waypoint-based (without Bezier constraint wp (ours)) and TF++ baselines, the Bezier decoder of the present invention reduces the average curvature (AvgC) by 0.1183 (0.0592 vs. 0.175) and 0.1502 (0.0592 vs. 0.20994) respectively on Route 13, while achieving a lower jitter of 0.1427 (AvgJ: 0.1522 vs. 0.249). The standard deviation of curvature (StdC) for all routes is reduced by 0.05 - 0.12, confirming the 2- inherent smoothness of continuous trajectories. These metrics are directly related to the observed 52% reduction in the collision rate, as kinematically compliant trajectories enable precise vehicle control under dynamic constraints.
[0081] The experimental results of the three benchmarks comprehensively verify the advantages of Fused-ST Bezier in perception-planning integration and trajectory quality. The consistent performance improvement between the metrics and scenarios validates the core innovations of the present invention: 1) The cross-modal attention mechanism effectively bridges the geometric-semantic gap between LiDAR and camera inputs; 2) The Bezier-based trajectory parameterization ensures physical feasibility while reducing the complexity of learning; 3) The spatio-temporal memory propagation enables coherent long-term reasoning. These advancements jointly push the boundaries of end-to-end autonomous driving systems in terms of safety, comfort, and reliability.
[0082] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An autonomous driving method based on modal fusion and Bessel optimization, characterized in that: The following steps are involved: Step 1: Data input and modal feature extraction; The vehicle-mounted multi-sensor system is used to synchronously collect environmental information; obtain two-dimensional image data and three-dimensional point cloud data; after obtaining the original data, the two heterogeneous data streams are sent to their own dedicated feature extraction backbone network in parallel; for image data, image feature maps with hierarchical semantics are abstracted layer by layer from the original pixels, and finally the preliminary image features are output. ; For point cloud data, capture the local geometric details and global spatial distribution of point cloud data and generate preliminary lidar features ; Step 2: Cross-modal feature fusion; Will and As input, it is sent to the cross-modal feature fusion module, which interacts and aggregates the information of the two through the multi-head cross attention mechanism and outputs the cross-modal fusion feature ; Step 3: Feature refinement and special feature generation; Process and differentiate features: On the one hand, the preliminary image features Perform upsampling and interpolation operations to obtain restored image features with more optimized resolution or channel dimensions On the other hand, cross-modal fusion features Upsampling interpolation is also performed to generate universal fusion features ; Based on universal fusion features , derives two key feature representations; first, through multi-layer convolution, downsampling and upsampling operations, Transform to bird's-eye view space to obtain bird's-eye view features ; Second, in order to capture the spatiotemporal dynamics of the scene and combine the vehicle’s intention, the general fusion feature It is integrated with dynamic path query, real-time vehicle status information and navigation information, and enhanced using the spatiotemporal attention mechanism to generate spatiotemporal features containing time sequence and context information. , specifically for trajectory prediction tasks; Step 4: Multi-task decoding and loss function; Finally, using multi-level features, i.e. , and , the model drives multiple decoders in parallel to complete different perception and prediction tasks.
2. The automatic driving method based on modal fusion and Bessel optimization according to claim 1, characterized in that: In step 1, the RegnetY-3.2GF neural network backbone is first used to extract preliminary features from the images obtained from the front-view camera and the lidar point cloud data. and ; After the channels are aligned, and By position encoding Connect and enhance to form embedded features : ; in, is the preliminary image feature, is the preliminary lidar feature, is the position code, is the embedded feature.
3. The automatic driving method based on modal fusion and Bessel optimization according to claim 2, characterized in that: In step 2, the fusion process is decomposed into a dual complementary path, where linear projection is used to compute a set of queries, keys, and values: ; ; ; ; in, represents the position function, Project the LiDAR features into the BEV grid, Transformation from perspective to BEV is achieved through depth perception feature enhancement. is the preliminary image feature Query The weight of is the preliminary lidar feature Query The weight of yes The weights of the keys and values, yes The weights of the keys and values of ; this dual query formulation implements two coordinated attention flows, namely and :First, Focus on lidar-derived keys / value , injecting precise spatial priors into visual features; second, with rgb based keys / value Interactively enrich point cloud features with contextual semantics: ; ; in, is the activation function, is the feature dimension of each attention head, is the number of attention heads; Then combine the two attention streams through gated fusion: ; in, represents the learnable gating parameters, is the sigmoid activation, Represents element-wise matrix multiplication.
4. The automatic driving method based on modal fusion and Bessel optimization according to claim 3, characterized in that: In step 3, the BilinearUp interpolation is used to Restore the image dimension and obtain the restored image features and through TrilinearUp interpolation Perform BEV dimension reduction to obtain universal fusion features : ; ; in, represents a parameterized deformable mesh, represents the interpolation weight, represents the sampling coordinates, Encapsulates time offset and time mixing weights and relative timestamps , for The weight parameter, represents the time step t. ; Then perform 3D reconstruction through convolution and upsampling to obtain the bird's-eye view features. : ; in, represents the depthwise convolution, represents point-wise convolution, represents the index of the network layer in the structure, For the The output of the layer network , is the convolution function, is the upsampling function; the architecture progressively downsamples features using strided convolutions before upsampling using transposed convolutions and skipping connections from the corresponding encoder stages; Restore image features Decoding and Bird's Eye View Features In the subsequent decoding, after adjusting the parameters through the multimodal prediction error, Encapsulates implicit and precise 3D scene perception features; defines a set of learnable parameters ,in represents the number of historical prediction trajectories, Indicates the number of time steps for the current trajectory planning, represents the dimension of the GRU unit; first, temporal self-attention is applied to a set of learnable query parameters middle: ; in, For historical feature queries, is the multi-head self-attention function; Next, the ego vehicle speed Included in the observation state; common fusion features After channel transformation, position encoding and feature flattening operations, it is then connected with the encoded self-state to form an enhanced feature : ; in, It is an expansion function used to expand multi-dimensional matrix data into a one-dimensional vector; It is a multi-layer perceptron, which consists of multiple linear layers and non-linear activation functions; represents the concatenation vector; is a position vector obtained by position encoding and used to add position information to the feature; Later, in and Perform spatial cross attention between them to generate spatiotemporal features : ; in, is the multi-head cross attention function.
5. The automatic driving method based on modal fusion and Bessel optimization according to claim 4, characterized in that: The step 4 specifically comprises the following steps: Step 4.1: Restoring image features decoding: LiDAR interaction information is incorporated into In the above example, the spatial features are enhanced by feature complementarity; the depth and semantic branches are respectively derived from Export, respectively used to decode the depth map and semantic segmentation map; specifically, through the cascade structure conversion of multi-level deconvolution and interpolation upsampling ; N-level bilinear interpolation doubles the spatial resolution, making it easier to reconstruct spatial information from low-dimensional features to high-resolution output: ; ; ; ; in, ; Respectively represent the corresponding formula Layer sampling parameters; depth estimation branch and semantic segmentation branch Use the same multi-stage architecture but keep independent transposed convolution and interpolation sampling parameters and ; The domains of these two tasks can be adaptively learned and the output image and Maintain consistent spatial resolution; Step 4.2: Bird's Eye View Features decoding; use Perform BEV perspective BevSemantic and BoundingBox predictions, supplemented by The derived front view depth and semantic segmentation map are used to enhance the model's 3D information and semantic perception of the global environment; for the BevSemantic branch, a lightweight feature transformation architecture is used to map spatial features to semantic categories: ; in, represents the convolution function; As a pure 3D detection task, the BoundingBox branch adopts a multi-branch parallel independent task head design, including the target center point heat map , bounding box size , Center Offset and The joint prediction of : ; in represents each prediction head, is the number of discrete orientation classes used for coarse orientation estimation, set to 15; Step 4.3: B́ezier curve decoding; The B́ezier branch obtains the Bezier curve control points by decoding the spatiotemporal features, and then obtains the planned trajectory through sampling; The scene used is a discrete target point In the navigation scenario, the target point information is embedded into the spatiotemporal features to guide trajectory planning; the multi-layer GRU units predict the control points of the Bézier curve in turn. : ; in, , Indicates Layer GRU is based on The offset of The offset output is the final output trajectory control point. is the total number of GRU cell layers and the number of curve control points; this work uses cubic B́ezier curves for trajectory planning, balancing lightweight computation with smooth dynamics; the continuous curve B comes from the control points , sample the trajectory points : ; in, represents the equally spaced sampling function, Represents the position parameters on the Bezier curve; After decoding and output, a dedicated loss function is tailored for each perception and planning component; the supervision mechanism includes two categories: first, auxiliary task supervision , auxiliary task supervision derived from the BEV feature decoder and the image feature decoder; second, trajectory planning constraints , a physical constraint applied to the B́ezier branch.
6. The automatic driving method based on modal fusion and Bessel optimization according to claim 5, characterized in that: In step 4.3, the auxiliary task supervision Specifically include: Depth estimation supervision of depth estimation branch depth map , semantic category supervision of semantic segmentation branch , semantic category supervision for BEV semantic branch segmentation and supervision of various heads in the BoundingBox branch ; Depth Estimation Supervision : Use L1 loss to supervise the depth estimation branch to ensure the accuracy of depth information measurement for 3D scene reconstruction: ; Among them, and Represent the predicted depth map and the real depth map respectively; Semantic Segmentation Supervision: and The semantic segmentation task is Using cross entropy loss on the categories: ; in, is the true value in the form of a one-hot vector; BoundingBox branch supervision: The detection branch combines multiple supervision targets: first, the height / width / center point / center point offset / yaw angle residual of the object: supervised by L1 loss with continuous parameters; second, yaw angle classification: cross entropy loss of discrete yaw boxes; third, Heatmap prediction: Gaussian coding supervision improved focal loss; ; Among them, the hyperparameters and are used to adjust the balance between positive / negative samples and hard / easy samples, respectively, and Supervision calculations for excluding traffic participants outside predefined boundaries; Corresponding to the supervision signal of each prediction head; Heatmap true value is generated by a Gaussian kernel procedure centered at the annotated target location; Final auxiliary task loss is expressed as a weighted sum of these constituent loss components: ; Trajectory planning constraints The target point guidance and safety constraints are implicitly obtained through the L1 reconstruction loss, and the kinematic feasibility of the vehicle is strengthened through regularization constraints, curvature continuity constraints, and smooth control point constraints to enhance trajectory smoothness and feasibility; Reconstruction losses : Ensure geometric alignment with the ground truth trajectory; ; in, represents the number of time intervals for prediction, represents the final predicted trajectory point, represents the real trajectory point; Control Point Regularization :By constraining the initial control point and the self-vehicle speed Alignment to ensure kinematic priors: ; Smoothness Constraint : Constrain the consistency of the distance between adjacent control points to avoid sudden trajectory jitter to ensure driving comfort: ; in, Thresholds determined from vehicle dynamics; Curvature Continuity : The continuity of the second-order derivative is maintained through control point constraints to ensure the continuity of the trajectory curvature change rate, so that the predicted trajectory conforms to the dynamic characteristics of the vehicle: ; Trajectory planning constraints Combining these components with empirical weights: ; Setting weight parameters , , and to ensure that vehicle power feasibility and ride comfort are considered with an appropriate performance lower limit.
Citation Information
Patent Citations
Automatic driving vehicle track prediction method and device and electronic equipment
CN113705636A
Multi-mode-based automatic driving perception method and device, equipment and medium
CN115879060A
Transform-based time sequence point cloud three-dimensional target detection
CN116740424A
Automatic driving decision planning method and system based on visual language large model
CN118238848A
Double-attention network model training method and vehicle control method
CN118485995A
Cited By
End-to-end automatic driving method and system based on multi-modal attention fusion
CN121224768A
An end-to-end automatic driving method and system based on multi-modal attention fusion
CN121224768B
Multi-mode unmanned aerial vehicle intelligent navigation method and system
CN121297870A
A multi-modal unmanned aerial vehicle intelligent navigation method and system
CN121297870B
Feature decoupling and reconstruction-based automatic driving heterogeneous collaborative domain adaptation method
CN122220822A