An autonomous driving method based on modal fusion and Bessel optimization

Through modal fusion and Bessel-optimized autonomous driving methods, the shortcomings of end-to-end autonomous driving systems in multimodal perception fusion and motion planning are solved, and efficient vehicle trajectory planning and environmental perception are achieved, reducing collision rates and improving riding comfort.

CN120071303BActive Publication Date: 2025-07-22JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510525457.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-22
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing end-to-end autonomous driving system has problems such as insufficient global three-dimensional context modeling, accumulation of cross-modal projection errors, high real-time computing load, high violation rate in complex scenarios, opaque trajectory generation in motion planning, separation of planning and perception, large calculation overhead, and failure to achieve joint space-time optimization.

Method used

Using a method based on modal fusion and Bezier optimization, two-dimensional images and three-dimensional point cloud data are obtained through a multi-sensor system, cross-modal feature fusion is used to generate bird's-eye and space-time features, combined with Bezier curve decoder for trajectory prediction, and multi-task decoding and loss functions are used for supervision to ensure the accuracy and feasibility of the trajectory.

Benefits of technology

The vehicle collision rate is reduced by 52% in dense urban market scenarios, improving riding comfort, enhancing environmental perception capabilities and interpretability of neural network architecture, and ensuring the spatiotemporal consistency and safety of trajectory planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071303B_ABST
    Figure CN120071303B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of autonomous driving technology, and provides an autonomous driving method based on modality fusion and Bezier optimization. It collaboratively processes camera images and lidar point clouds through a bidirectional cross-modal attention mechanism, realizes dynamic feature recalibration in the perspective space and the bird's-eye view space, enhances visual semantics by using lidar geometric priors, and simultaneously optimizes point cloud interpretation through visual features, significantly improving 3D perception accuracy. In motion planning, a differentiable Bezier curve decoder is innovatively adopted, combined with a spatio-temporal attention mechanism to generate a C² continuous parametric trajectory generation scheme, ensuring kinematic feasibility and dynamic scene adaptability. The present invention achieves breakthrough performance in the CARLA benchmark test, establishing a new paradigm for collaborative optimization of perception and planning for urban autonomous driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to an autonomous driving method based on modality fusion and Bezier optimization. Background Art

[0002] The end-to-end autonomous driving system provides a paradigm-shifting alternative to traditional modular architectures by establishing a direct mapping from raw sensor inputs to planned trajectories or control signals. By integrating the traditional perception-planning-control process into a unified neural framework, it fundamentally eliminates the error-prone intermediate representations and cascading failures inherent in sequential processing chains - particularly coordinate system misalignment artifacts and cumulative delay effects. Additionally, the simplicity of this end-to-end architecture conceptually bypasses the need for manual rules while enabling joint optimization of environmental understanding and motion generation.

[0003] Existing end-to-end autonomous driving systems have significant limitations in multi-modal perception fusion. Although recent research has verified the complementarity of multi-source data through early, mid-term, and late fusion of camera and depth modalities, and also demonstrated the effectiveness of semantic and depth information as intermediate representations, mainstream methods are still limited to one-way cross-modal projections of multi-view lidar representations such as bird's-eye view (BEV) and front view (RV), resulting in insufficient global three-dimensional context modeling and making it difficult to ensure safe decision-making in dense traffic flows. Even with the improvement of interpretability through safety-constrained action generation, its complex cross-modal attention mechanism and multi-stage feature transformation still face risks of real-time computational load and error propagation, and recent experiments have further revealed the defect of the sharp increase in violation rates of existing frameworks in complex scenarios, highlighting the spatio-temporal mismatch problem between perceptual representation and dynamic scene understanding.

[0004] In the field of motion planning, although early end-to-end methods verified the feasibility of latent space trajectory generation, they suffered from optimization instability and an interpretability crisis due to the opaque decision-making process. Subsequent hybrid architectures based on explicit cost maps improved interpretability through manual rules, but their high-dimensional spatial grid construction led to a sharp increase in computational overhead and poor generalization ability in unknown traffic configurations. Some methods, such as those that improved planning constraint satisfaction through vectorized scene encoding and object-centered representations, still separated prediction and planning into sequential tasks, failing to achieve spatio-temporal joint optimization and ignoring their inherent geometric-dynamic coupling, resulting in limited trajectory smoothness and obstacle avoidance capabilities in dense traffic flows.

[0005] Deeper system-level defects stem from the architectural fragmentation of the perception-planning link. In existing multi-modal fusion frameworks, although perception is enhanced through cross-sensor feature mapping, their hierarchical processing flow leads to the accumulation of modal alignment errors, and discrete planning representations make it difficult to ensure kinematic feasibility and comfort. Even if perspective-invariant representations can be constructed through self-supervised learning, there are still efficiency bottlenecks in their multi-stage feature propagation. Although semantic point cloud fusion schemes optimize the representation of driving tasks, they do not solve the curvature continuity constraint problem of trajectory parameterization. These limitations together result in an essential bottleneck in the existing system's response to environmental uncertainty, maintenance of spatio-temporal consistency, and compliance of motion planning, and there is an urgent need for a tightly coupled perception-planning collaborative optimization framework. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide an autonomous driving method based on modal fusion and Bezier optimization, aiming to solve the problems proposed in the above background technology.

[0007] The embodiments of the present invention are implemented as follows. An autonomous driving method based on modal fusion and Bezier optimization includes the following steps:

[0008] Step 1: Data input and modal feature extraction;

[0009] Use the in-vehicle multi-sensor system to synchronously collect environmental information. Specifically, use a front-view camera to capture high-resolution two-dimensional image data. At the same time, use a lidar system to obtain three-dimensional point cloud data of the environment. After obtaining the original data, send these two heterogeneous data streams in parallel to their respective dedicated feature extraction backbone networks. For image data, layer by layer abstract image feature maps with hierarchical semantics from the original pixels, and finally output the preliminary image feature X img ; for point cloud data, capture the local geometric details and global spatial distribution of the point cloud data to generate the preliminary lidar feature X lidar ;

[0010] Step 2: Cross-modal feature fusion;

[0011] After obtaining the preliminary single-modal features, take X img and X lidar as inputs and send them to the cross-modal feature fusion module. Through the multi-head cross-attention mechanism, interact and aggregate the information of the two to output the cross-modal fusion feature X fout ;

[0012] Step 3: Feature refinement and dedicated feature generation;

[0013] To adapt to the requirements of different downstream tasks, process and differentiate the features: on the one hand, through BilinearUp interpolation for X foutPerform image dimension reduction to obtain a restored image feature F with more optimized resolution or channel dimensions img , which is mainly used for perception tasks based on the image perspective; on the other hand, perform BEV dimension reduction on X fout through TrilinearUp interpolation to obtain the fused feature F fused .

[0014] Based on the general fused feature F fused , two key feature representations are derived; firstly, through multi-layer convolution, downsampling, and upsampling operations, transform F fused to the bird's-eye view (BEV) space to obtain the bird's-eye view feature F bev ; secondly, in order to capture the spatio-temporal dynamics of the scene and combine vehicle intentions, integrate the general fused feature F fused with dynamic path queries, real-time ego-vehicle state information (such as speed, acceleration), and navigation information, and use spatio-temporal attention mechanisms for enhancement processing, finally generating the spatio-temporal feature F st containing temporal and context information, which is specifically used for trajectory prediction tasks;

[0015] Step 4: Multi-task decoding and loss function;

[0016] Finally, using multi-level features, namely F img , F st and F bev , the model drives multiple decoders in parallel to complete different perception and prediction tasks; specifically including: firstly, using F img for depth map decoding to estimate the scene pixel-level depth, and using L1 loss for supervision; secondly, using F img for image semantic segmentation decoding to assign semantic categories to each pixel in the image, and using cross-entropy loss for supervision; thirdly, inputting F st into the Bezier decoder to predict and generate the control points of the Bezier curve, and then constructing a smooth Bezier curve, and obtaining the final trajectory prediction result through curve sampling. This task is optimized using a composite loss function, including reconstruction loss, regularization constraint, smoothness constraint, and C 2 continuity constraint to ensure the accuracy and feasibility of the trajectory; fourthly, using F bev for semantic segmentation decoding in the BEV perspective to identify road structures and drivable areas in the bird's-eye view, etc., and using cross-entropy loss for supervision; fifthly, using F bev for object 3D bounding box decoding to detect and locate objects in three-dimensional space, and using L1 loss, cross-entropy loss, and custom loss for joint supervision.

[0017] An autonomous driving method based on modal fusion and Bezier optimization provided by an embodiment of the present invention has the following beneficial effects:

[0018] (1) Unify cross-modal attention and trajectory optimization based on to reduce the vehicle collision rate by 52% and achieve no pedestrian collisions in the end-to-end framework in dense urban scenarios;

[0019] (2) The Bezier decoder ensures C 2- continuity, has inherent smoothness and kinematic feasibility, greatly reduces the acceleration during driving, and improves the riding comfort;

[0020] (3) Establish a spatio-temporal attention mechanism to enhance trajectory planning in the end-to-end autonomous driving system, enhance the performance of trajectory planning in the time and space dimensions, and show top results in the simulation;

[0021] (4) Achieve unified multi-task perception and supervision, significantly enhance the environmental perception ability and the interpretability of the neural network architecture, and at the same time bridge the semantic-geometry gap in scene cognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of an autonomous driving method based on modal fusion and Bezier optimization provided by an embodiment of the present invention;

[0023] Figure 2 is a schematic diagram of an autonomous driving architecture in an autonomous driving method based on modal fusion and Bezier optimization provided by an embodiment of the present invention;

[0024] Figure 3 is a visualization of a Bezier-based trajectory generation framework;

[0025] Figure 4 is a comparison of trajectory point distributions of this method in the Longest6 benchmark (top row) and TF++ (bottom row); where (a) and (b) are right-angle turns at a two-lane T-junction, (c) is a right-angle turn at a multi-lane T-junction, (d) is a lane change within a roundabout, and (e) is a lane change on a straight road. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0027] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.

[0028] As Figure 1As shown in the figure, a self-driving method based on modal fusion and Bessel optimization provided by an embodiment of the present invention aims to comprehensively understand and accurately predict the driving environment by fusing multi-modal perception information of cameras and lidar. The method includes the following steps:

[0029] Step 1: Data input and modal feature extraction;

[0030] The starting point of this method is to synchronously collect environmental information using the in-vehicle multi-sensor system. Specifically, a front-view camera with a 110° viewing angle is used to capture high-resolution two-dimensional image data; at the same time, a 16-line lidar (LiDAR) system is used to obtain the three-dimensional point cloud data of the environment, providing accurate geometric structures and depth information. After obtaining the original data, the system parallelly sends these two heterogeneous data streams into their respective dedicated feature extraction backbone networks. For the image data, a pre-trained RegnetY-3.2GF neural network with excellent performance is loaded as the image feature extraction module. Through a series of convolutional, pooling, and non-linear activation operations, this module gradually abstracts image feature maps with hierarchical semantics from the original pixels and finally outputs the preliminary image feature X img , which is a feature vector. For the point cloud data, the lidar feature extraction module uses the initialized RegnetY-3.2GF neural network. These networks are designed to effectively capture the local geometric details and global spatial distributions of the point cloud data and generate the preliminary lidar feature X lidar , which is in the form of a feature vector.

[0031] Step 2: Cross-modal feature fusion;

[0032] After obtaining the preliminary single-modal features, in order to effectively integrate complementary information, a cross-modal feature fusion module is designed; this module receives X img and X lidar as inputs, and interacts and aggregates the information of the two through the multi-head cross-attention mechanism, and outputs the cross-modal fusion feature X fout .

[0033] Step 3: Feature refinement and dedicated feature generation;

[0034] In order to adapt to the requirements of different downstream tasks, the features are further processed and differentiated: on the one hand, the image dimension of X fout is restored through BilinearUp interpolation to obtain the restored image feature F img with more optimized resolution or channel dimensions, which is mainly used for perception tasks from the image perspective; on the other hand, the BEV dimension of X fout is restored through TrilinearUp interpolation to obtain the fusion feature F fused。

[0035] Based on the general fusion feature F fused , two key feature representations are derived. First, through operations such as multi-layer convolution, downsampling, and upsampling (similar to common BEV feature generation network structures), F fused is transformed into the bird's-eye view (BEV) space to obtain the bird's-eye view feature F bev , which provides a basis for subsequent perception tasks from the BEV perspective. Second, in order to capture the spatio-temporal dynamics of the scene and combine vehicle intentions, the general fusion feature F fused is integrated with dynamic path queries, real-time ego-vehicle state information (such as speed, acceleration), and navigation information, and enhanced processing is performed using a spatio-temporal attention mechanism. Finally, the spatio-temporal feature F st containing temporal and context information is generated, which is specifically used for the trajectory prediction task.

[0036] Step 4: Multi-task decoding and loss function;

[0037] Finally, using multi-level features (F img , F st and F bev ), the model drives multiple decoders in parallel to complete different perception and prediction tasks. Specifically, first, F img is used for depth map decoding to estimate the scene pixel-level depth, and the L1 loss is used for supervision; second, F img is used for image semantic segmentation decoding to assign semantic classes to each pixel in the image, and the cross-entropy loss is used for supervision; third, F st is input into the Bezier decoder to predict the control points of the Bezier curve, and then a smooth Bezier curve is constructed. The final trajectory prediction result is obtained through curve sampling. This task is optimized using a composite loss function, including reconstruction loss, regularization constraint, smoothness constraint, and C 2 continuity constraint to ensure the accuracy and feasibility of the trajectory; fourth, F bev is used for semantic segmentation decoding from the BEV perspective to identify road structures, drivable areas, etc. in the bird's-eye view, and the cross-entropy loss is used for supervision; fifth, F bev is used for object 3D bounding box decoding to detect and locate objects in three-dimensional space, and joint supervision is performed using L1 loss, cross-entropy loss, and custom loss. Through this multi-task, end-to-end learning framework, this method can make full use of multi-modal data to achieve comprehensive perception and dynamic prediction of complex traffic scenes.

[0038] In the embodiment of the present invention, the proposed autonomous driving architecture is specifically as Figure 2As shown, the three-dimensional perception with the integration of a front-view RGB camera and lidar includes: (1) Dual-branch feature embedding with cross-modal attention fusion; (2) Auxiliary task branch and spatio-temporal enhanced trajectory decoding; (3) Joint optimization through multi-constraint supervision to achieve safe perception trajectory generation.

[0039] As Figure 2 shown, as a preferred embodiment of the present invention, in step 1, first, the RegnetY-3.2GF neural network backbone is used to extract the preliminary features X img and X lidar from the images and lidar point cloud data obtained from vehicle-mounted sensors respectively. After channel alignment, these features (X img and X lidar ) are connected and enhanced through the position encoding P pos to form the embedded feature X emb :

[0040]

[0041] Among them, X img is the preliminary image feature, X lidar is the preliminary lidar feature, the position encoding is P pos , and X emb is the embedded feature.

[0042] As a preferred embodiment of the present invention, in step 2, the fusion process is decomposed into two complementary paths, where linear projection is used to calculate a set of queries, keys, and values:

[0043]

[0044] Among them, PE(·) represents the position function, BEVGrid(·) projects the LiDAR feature into the BEV grid, and P(·) realizes the perspective-to-BEV conversion through depth perception feature elevation. is the weight of the query Q img of the preliminary image feature X rgb , is the weight of the query Q lidar of the preliminary lidar feature X pc , is the weight of the key K img and the value V rgb of X rgb , is the weight of the key K lidar and the value V lidar of X lidar ; this dual-query formula realizes two collaborative attention flows, namely O img2pc and O pc2img : First, Qrgb Focus on the lidar-derived key K lidar / value V lidar , inject precise spatial priors into visual features; secondly, Q pc interacts with the rgb-based key K rgb / value V rgb to enrich the point cloud features with contextual semantics:

[0045]

[0046]

[0047] where Softmax(·) is the activation function, D h is the feature dimension of each attention head, and H is the number of attention heads;

[0048] Then, two attention streams are combined through gated fusion:

[0049]

[0050] where W g represents the learnable gating parameter, σ(·) is the sigmoid activation, and ⊙ represents element-wise matrix multiplication.

[0051] As Figure 2 shown, as a preferred embodiment of the present invention, in step 3, the image dimension of X fout is reduced through BilinearUp interpolation to obtain the reduced image feature F img and the BEV dimension of X fout is reduced through TrilinearUp interpolation to obtain the general fusion feature F fused :

[0052]

[0053] where Θ grid ∈R H×W×2 represents the parameterized deformation grid, α ij represents the interpolation weight, p ij represents the sampling coordinates, encapsulates the time offset, the time mixing weight β t and the relative timestamp Δτ, θ ij is the weight parameter of Θ grid , represents at the t-th time step

[0054] Then, 3D reconstruction is performed through convolution (Conv) and upsampling (Upsample) to obtain the bird's-eye view feature F bev:

[0055]

[0056] Among them, DWConv(·) represents depth convolution, PWConv(·) represents pointwise convolution, n represents the index of the network layer in this structure. is the F output by the nth layer network bev , is the convolution function, and Upsample(·) is the upsampling function. Before using transposed convolution for upsampling and skip connection from the corresponding encoder stage, the architecture gradually downsamples the features through strided convolution.

[0057] Restore the image feature F img Decode and the bird's-eye view feature F bev After parameter adjustment through multimodal prediction error in the subsequent decoding, F fused encapsulates implicit and precise 3D scene perception features. Innovatively define a set of learnable parameters where H t = 6 represents the number of historical prediction trajectories, T step = 8 represents the number of time steps for the current trajectory planning, and Z gru = 256 represents the dimension of the GRU unit. First, apply temporal self-attention to a set of learnable query parameters Q wp (parameters that can be automatically adjusted and optimized according to the reverse gradient during model training, similar to the weight parameters in the model network):

[0058] Q′ wp = MHA self (Q wp , Q wp , Q wp )

[0059] Among them, Q′ wp is the historical feature query. This parameter is still a learnable query parameter, but with the effect of self-attention. To facilitate differentiation from the previous learnable parameters, a superscript is added; MHA self (·) is the multi-head self-attention function;

[0060] Next, incorporate the ego vehicle speed S ego = v ego into the observation state. The fused feature F fusion undergoes channel transformation, position encoding, and feature flattening operations, and then is concatenated with the encoded ego state to form the enhanced feature F enh :

[0061] F flat = Flatten(F fused + Ppos )

[0062] F ego = MLP(S ego ) + P pos

[0063]

[0064] where Flatten(·) is an unfolding function (or data unfolding function) for unfolding multi-dimensional matrix data into a one-dimensional vector; MLP(·) is a multi-layer perceptron composed of multiple linear layers and non-linear activation functions; denotes the concatenated vector; P pos is a position vector obtained by position encoding and is used to add position information to the features.

[0065] Subsequently, spatial cross-attention is performed between Q′ wp and F enh to generate spatio-temporal feature F st :

[0066] F st = MHA cross (Q′ wp , F enh , F enh )

[0067] where MHA cross (·) is the multi-head cross-attention function.

[0068] As a preferred embodiment of the present invention, in step 4, F img and F bev are decoded into corresponding multi-modal outputs. These outputs supplement the core trajectory prediction branch, utilize F fused for spatio-temporal feature interaction, and achieve trajectory generation based on the curve; specifically, it includes the following steps:

[0069] Step 4.1: Restore the image feature F img Decode:

[0070] The LiDAR interaction information is merged into F img , and its spatial features are enhanced through feature complementarity. The depth and semantic branches are respectively derived from F img and are respectively used to decode the depth map and the semantic segmentation map. Through this process, the perception of the ego vehicle regarding the surrounding distance and object categories is enhanced, thereby improving the trajectory planning ability of the core branch. Specifically, through a cascaded structure of multi-level deconvolution and interpolation upsampling N-level bilinear interpolation doubles the spatial resolution, facilitating the reconstruction of spatial information from low-dimensional features to high-resolution output:

[0071]

[0072]

[0073] Among them, and respectively represent the sampling parameters of the n-th layer in the corresponding formula. As Figure 1 shown, the depth estimation branch and the semantic segmentation branch adopt the same multi-level architecture, but maintain independent transposed convolution and interpolation sampling parameters and This design enables domain adaptive feature learning for these two tasks, while keeping the output images I depth and I isem at a consistent spatial resolution;

[0074] Step 4.2: Decode the bird's-eye view feature F bev ;

[0075] Compared with F img the bird's-eye view feature F bev goes through multiple convolutional and upsampling layers, thus generating more distinct spatial and semantic context features. Using F bev for BEV perspective BevSemantic and BoundingBox predictions complements the front view depth and semantic segmentation maps derived from F img to enhance the model's 3D information and semantic perception of the global environment. For the BevSemantic branch, given that F bev has already gone through multiple layers of convolution and upsampling, a lightweight feature transformation architecture is adopted to map the spatial features to semantic classes:

[0076] I bsem = Upsample(Conv(Conv(F bev )))

[0077] where Conv(·) represents the convolution function;

[0078] The BoundingBox branch, as a pure 3D detection task, adopts a multi-branch parallel independent task head design, including the joint prediction of the object center point heatmap O hp , the bounding box size O wh , the center offset O yc and O yr for synthesizing the yaw angle O y :

[0079] O x= MLP(Conv(Conv(F bev )))

[0080]

[0081] where O x ∈ {O hp , O wh , O os , O yr , O yc} represents each prediction head, and N rc (set to 15) is the number of discrete direction classes for rough direction estimation.

[0082] Step 4.3: Curve decoding;

[0083] The branch is the core component of the present invention. It decodes the spatio-temporal features to obtain the Bezier curve control points, thereby obtaining the Bezier curve, and then samples to obtain the planned trajectory.

[0084] The scenario of this method is a navigation scenario with discrete target points P tg . These discrete target points are usually separated by at least 50m. The target point information is embedded in the spatio-temporal features to guide the trajectory planning. The multi-layer GRU units sequentially predict the control points P b of the curve:

[0085] F tg = MLP(P tg )

[0086]

[0087] x i = MLP(h i )

[0088] where h0 = Concat(F tg , F st ), and x i represents the offset of the i-th layer GRU based on x i-1 . The offset of is the final output trajectory control point. t c = 4 is the total number of GRUcell layers and the number of curve control points. This work uses a cubic curve for trajectory planning to balance lightweight calculations with smooth dynamics. The continuous curve B is derived from the control points P b , and the trajectory points are obtained by sampling

[0089]

[0090] P trj = Sample(B)

[0091] where Sample(·) represents an equidistant sampling function, and s represents the position parameter on the Bezier curve.

[0092] This method ensures a smooth and dynamically feasible trajectory, while maintaining computational efficiency and ensuring C 2- continuity. After decoding to obtain the output, dedicated loss functions are tailored for each perception and planning component. The supervision mechanism includes two categories: one is the auxiliary task supervision L A , which is the auxiliary task supervision derived from the BEV feature decoder and the image feature decoder; the other is the trajectory planning constraint L T , which is applied to the physical constraints of the

[0093] (1) The auxiliary task supervision L A , specifically including: the depth estimation supervision L depth of the depth map in the depth estimation branch, the semantic category supervision of the semantic segmentation branch the semantic category supervision of the BEV semantic branch segmentation and the supervision L bbox of various heads in the BoundingBox branch.

[0094] The depth estimation supervision L depth : Use the L1 loss to supervise the depth estimation branch to ensure the accuracy of the depth information measurement for 3D scene reconstruction:

[0095]

[0096] where and D represent the predicted depth map and the ground truth depth map respectively.

[0097] Semantic segmentation supervision: The front view (Iisem) and BEV (Ibsem) semantic segmentation tasks use the cross-entropy loss on C = 22 categories (aligned with the CARLA simulator categories, including pedestrians, vehicles, roads, and traffic lights):

[0098]

[0099] where y c is the ground truth in the form of a one-hot vector;

[0100] BoundingBox branch supervision: The detection branch combines multiple supervision objectives: 1) The height / width / center point / center point offset / yaw angle residual of the object: supervised using the L1 loss of continuous parameters; 2) Yaw angle classification: cross-entropy loss for discrete yaw bins; 3) Heatmap prediction: improved focal loss supervised by Gaussian encoding;

[0101]

[0102] Among them, the hyperparameters α and γ are used to adjust the balance between positive / negative samples and hard / easy samples, respectively, while is used to exclude the supervision calculation of traffic participants outside the predefined boundary range; y x corresponds to the supervision signal of each prediction head; the Heatmap ground truth y p is generated by a Gaussian kernel procedure centered on the annotated target position.

[0103] The final auxiliary task loss L A is expressed as the weighted sum of these component loss terms:

[0104]

[0105] (2) Trajectory planning constraint L T , The trajectory prediction branch is optimized by a composite loss, as Figure 3 shown. This loss implicitly obtains target point guidance and safety constraints through the L1 reconstruction loss, and strengthens the kinematic feasibility of the vehicle through regularization constraints, curvature continuity constraints, and smooth control point constraints to enhance trajectory smoothness and feasibility.

[0106] The reconstruction loss L rs : Ensures geometric alignment with the ground truth trajectory to ensure lower performance limitations;

[0107]

[0108] Among them, T represents the number of predicted time intervals, represents the final predicted trajectory point, p trj represents the ground truth trajectory point.

[0109] The control point regularization L cp : Guarantees the kinematic prior by constraining the alignment of the initial control point with the ego vehicle speed :

[0110]

[0111] The smoothness constraint L sm:Constrain the consistency of the spacing between adjacent control points to avoid sudden trajectory jitter and ensure driving comfort:

[0112]

[0113] where d ref is the threshold determined from vehicle dynamics.

[0114] Curvature continuity L cc : Maintain second-order derivative continuity through control point constraints to ensure the continuity of the trajectory curvature change rate and make the predicted trajectory conform to the dynamic characteristics of the vehicle:

[0115] L cc = ||P3 - 2P2 + P1||2

[0116] Trajectory planning constraint L T Combine these components with empirical weights:

[0117] L T = w rs ·L rs + w cp ·L cp + w sm ·L sm + w cc ·L cc

[0118] Set the weight parameters w rs 、w cp 、w sm and w cc to ensure vehicle power feasibility and ride comfort while considering the appropriate performance lower limit.

[0119] As a preferred embodiment of the present invention, the learnable parameter path Queries of the spatio-temporal attention of the present invention retain 6 historical time steps and simultaneously predict a 2-second future trajectory sampled at 250-millisecond intervals. The architecture specifications include: 1) Image encoding of RegNetY-3.2GF through pre-trained weights; 2) LiDAR stream processing through RegNetY-3.2GF initialized from scratch; 3) Cross-modal attention with 4 heads and a 0.1 dropout probability; 4) Multi-scale BEV reconstruction using sequential 1×1→3×3→3×3 convolutional kernels; 5) The decoder implements an 8-head self / cross-attention mechanism with 6 stacked transformations; 6) A GRU cell with a 256-dimensional input and a 64-dimensional hidden state, and the output dimension is the same as Control point count alignment; 7) Unified 3×3→1×1 convolution pattern for BEV semantics and bounding box branches. The training scheme uses AdamW for three-phase optimization for all branches: Phase 1, 30 epochs of pre-training (lr = 3e-4); Phase 2, 30 epochs of core branch refinement (lr = 1e-4); Phase 3, 15 epochs of joint fine-tuning (lr = 3e-5) to enhance cross-branch collaboration. The implementation occurs on dual RTX 3090 GPUs with a batch size of 16 and takes approximately 180 training hours.

[0120] The evaluation framework of the present invention includes two benchmarks: the CARLA benchmark, which is used to evaluate driving performance indicators such as driving scores, violation scores, and collision rates in various CARLA leaderboard route scenarios; and the curve benchmark, which is used to evaluate the trajectory smoothness, ride comfort, and kinematic feasibility of vehicles on the CARLA leaderboard routes.

[0121] CARLA benchmark: The Town05 long / short benchmark and the Longest6 benchmark were comprehensively tested using the industry-standard CARLA leaderboard evaluation protocol, with a focus on the more comprehensive Longest6 benchmark proposed by TransFuser. Longest6 is derived from the CARLA leaderboard 1.0 route and aggregates 6 of the longest trajectories (average length: 1.5 km) in Towns 01 - 006. Each route has a unique environmental configuration, combining 6 weather conditions (cloudy; humid; moderate rain; humid cloudy; heavy rain; light rain) and 6 lighting states (night; twilight, i.e., the time before sunrise and after sunset when the sun illuminates the sky but the sun is not visible; dawn; morning; noon; sunset), enhancing adversarial traffic scenarios. The Town05 benchmark test contains two configurations: 1) Town05 Short: 10 routes (100 - 500 m), 3 for each route; 2) Town05 Long: 10 extended routes (1000 - 2000 m) with 10 intersections. This benchmark challenges systems with different road geometries (multi-lane arterials, single-lane roads, bridges, highways) and dynamic agent interactions.

[0122] Curve benchmark: For trajectory planning evaluation, Routes 13, 23, and 25 were specifically selected from the Longest6 benchmark test because of their challenging geometric configurations - each configuration contains multiple curved sections, roundabouts, and S-turns, strictly testing trajectory smoothness, dynamic feasibility, and passenger comfort.

[0123] The CARLA benchmark evaluation uses ten key metrics: 1) Driving Score, which is the product of route completion and violation penalties; 2) Route Completion, which is the percentage of the route distance completed; 3) Violation Score, which is a comprehensive score aggregated from various violation types; 4) Pedestrian Collision, which is the collision rate with pedestrians; 5) Vehicle Collision, which is the collision rate with vehicles; 6) Static Object Collision, which is the collision rate with static elements; 7) Red Light Violation, which is the probability of traffic light violations; 8) Road Deviation, which is the percentage deviation from the designated navigation route; 9) Route Timeout, which is the percentage of routes that exceed the maximum allowed time; 10) Vehicle Stuck, which is the probability of extended agent inactivity during driving. The specific evaluation results of the Longest6 benchmark are shown in Table 1.

[0124] Table 1 Longest6 Benchmark Test Results

[0125]

[0126] In the Longest6 benchmark test, this method achieved a SOTA driving score (DS) of 70, exceeding TF++ by 1%, while reducing the vehicle collision rate by 52% (0.31 vs. 0.83). Notably, it maintained zero pedestrian collisions and eliminated the violation score (IS) for traffic light violations and blocking scenarios, demonstrating robust safety awareness in a dense urban environment.

[0127] The specific evaluation results of the Town05 benchmark are shown in Table 2. For the Town05 benchmark, this method set a new performance record: 83.75 DS (+9.05% higher than the previous SOTA) and 98.15% route completion on long routes, with a completion rate of 98.50% on short routes. This demonstrates the special ability of this method in handling complex road geometries and extended navigation challenges. The results particularly emphasize the effectiveness of the spatio-temporal memory bank of this method in maintaining the consistency of extended trajectories.

[0128] Table 2 Town05 Benchmark Test Results

[0129]

[0130] To quantify the trajectory quality, four evaluation criteria were established: 1) Lateral deviation and standard deviation from the expert trajectory to evaluate path tracking accuracy; 2) Average curvature and standard deviation of the ego-vehicle trajectory to measure smoothness; 3) Kinematic feasibility analysis based on the curvature estimation model proposed in "Vehicle Dynamics and Control":

[0131]

[0132] Among them, represents the yaw rate, v ego vehicle speed, δ fThe front wheel angle, and L is the wheelbase between the front and rear axles;

[0133] 4) The Jerk (time derivative of acceleration) statistic is directly related to passenger comfort.

[0134] Test results of Route 13 with curve benchmark in Table 3

[0135]

[0136] Test results of Route 23 with curve benchmark in Table 4

[0137]

[0138] Test results of Route 25 with curve benchmark in Table 5

[0139]

[0140] The quantitative analysis of Tables 3, 4, and 5 for the curve benchmark reveals a fundamental improvement in trajectory quality. Figure 4 It provides a visual comparison of the trajectory distribution between the present invention and TF++ in a representative scenario of the present invention to support the tabular results. Compared with the waypoint-based (without Bezier constraint wp(ours)) and the TF++ baseline, the decoder of the present invention reduces the average curvature (AvgC) by 0.1183 (0.0592 vs. 0.175) and 0.1502 (0.0592 vs. 0.20994) respectively on Route 13, while achieving a lower jerk of 0.1427 (AvgJ: 0.1522 vs. 0.249). The standard deviation of curvature (StdC) for all routes is reduced by 0.05 - 0.12, confirming the 2- inherent smoothness of the continuous trajectory. These metrics are directly related to the observed 52% reduction in the collision rate because kinematically compliant trajectories enable precise vehicle control under dynamic constraints.

[0141] The experimental results of the three benchmarks comprehensively verify the advantages in perception planning integration and trajectory quality. The consistent performance improvement between the metrics and the scenarios validates the core innovations of the present invention: 1) The cross-modal attention mechanism effectively bridges the geometric-semantic gap between LiDAR and camera inputs; 2) The trajectory parameterization based on... ensures physical feasibility while reducing the complexity of learning; 3) Spatiotemporal memory propagation enables coherent long-term reasoning. These advancements together push the boundaries of end-to-end autonomous driving systems in terms of safety, comfort, and reliability.

[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An autonomous driving method based on modal fusion and Bessel optimization, characterized in that, The following steps are involved: Step 1: Data input and modal feature extraction; Synchronously collect environmental information using an in-vehicle multi-sensor system; obtain two-dimensional image data and three-dimensional point cloud data; after obtaining the original data, feed these two heterogeneous data streams into their respective dedicated feature extraction backbone networks in parallel; for the image data, layer by layer abstract image feature maps with hierarchical semantics from the original pixels, and finally output the preliminary image feature X img ; for the point cloud data, capture the local geometric details and global spatial distribution of the point cloud data to generate the preliminary lidar feature X lidar ; Step 2: Cross-modal feature fusion; Take X img and X lidar as inputs and send them to the cross-modal feature fusion module. Through the multi-head cross-attention mechanism, interact and aggregate the information of the two to output the cross-modal fusion feature X fout ; Step 3: Feature refinement and special feature generation; Process and differentiate the features: On the one hand, perform image dimension reduction on X through BilinearUp interpolation fout to obtain a restored image feature F with more optimized resolution or channel dimensions img ; on the other hand, perform BEV dimension reduction on X through TrilinearUp interpolation fout to obtain a fused feature F fused ; Based on the general fusion feature F fused , two key feature representations are derived; firstly, through multi-layer convolution, downsampling and upsampling operations, F fused is transformed into the bird's-eye view space to obtain the bird's-eye view feature F bev ; secondly, in order to capture the spatio-temporal dynamics of the scene and combine the vehicle intention, the general fusion feature F fused is integrated with the dynamic path query, the real-time ego-vehicle state information and the navigation information, and enhanced by the spatio-temporal attention mechanism, and finally the spatio-temporal feature F st containing temporal and context information is generated, which is specifically used for the trajectory prediction task; Step 4: Multi-task decoding and loss function; Finally, using multi-level features, namely F img , F st and F bev , the model drives multiple decoders in parallel to complete different perception and prediction tasks.

2. The automatic driving method based on modal fusion and Bessel optimization according to claim 1, characterized in that In the step 1, initially, the preliminary features X img and X lidar are respectively extracted from the images obtained by the front camera and the lidar point cloud data using the RegnetY-3.2GF neural network backbone; after channel alignment, X img and X lidar are connected and enhanced through the position encoding P pos to form the embedded feature X emb : Among them, X img is the preliminary image feature, X lidar is the preliminary lidar feature, P pos is the position encoding, X emb is the embedded feature.

3. The autonomous driving method based on modal fusion and Bessel optimization according to claim 2, characterized in that, In step 2, the fusion process is decomposed into a dual complementary path, where linear projection is used to compute a set of queries, keys, and values: Among them, PE(·) represents the position function, BEVGrid(·) projects LiDAR features into the BEV grid, and P(·) realizes the perspective-to-BEV conversion through depth-aware feature enhancement. is the query Q of the preliminary image feature X img ; rgb is the weight of is the query Q of the preliminary LiDAR feature X lidar ; pc is the weight of is the weight of the key and value of X img ; is the weight of the key and value of X lidar ; This dual-query formula realizes two collaborative attention flows, namely O img2pc and O pc2img : First, Q rgb focuses on the lidar-derived key K lidar / value V lidar , injecting an accurate spatial prior into the visual features; Second, Q pc interacts with the rgb-based key K rgb / value V rgb to enrich the point cloud features with contextual semantics: Among them, Softmax(·) is the activation function, D h is the feature dimension of each attention head, and H is the number of attention heads; Then combine the two attention streams through gated fusion: where, W g represents learnable gating parameters, σ(·) is the sigmoid activation, and ⊙ represents element-wise matrix multiplication.

4. The automatic driving method based on modal fusion and Bessel optimization according to claim 3, characterized in that In the said step 3, perform image dimension reduction on X through BilinearUp interpolation fout to obtain the reduced image feature F img and perform BEV dimension reduction on X through TrilinearUp interpolation fout to obtain the general fusion feature F fused : where, Θ grid ∈R H×W×2 represents a parameterized deformed mesh, α ij represents an interpolation weight, p ij represents a sampling coordinate, encapsulates the time offset, the time blending weight β t and the relative timestamp Δτ, θ ij is the weight parameter of Θ grid and is for representing at the t-th time step Then, three-dimensional reconstruction is performed through convolution and upsampling to obtain the bird's-eye view feature F bev : Among them, DWConv(·) represents depthwise convolution, PWConv(·) represents pointwise convolution, n represents the index of the network layer in this structure, is the F output by the nth layer network bev , is the convolution function, and Upsample(·) is the upsampling function; before the architecture uses transposed convolution for upsampling and skips connections from the corresponding encoder stage, it gradually downsamples the features through strided convolution; Restore image feature F img Decode and bird's-eye view feature F bev After parameter adjustment by multi-modal prediction error in subsequent decoding, F fused Encapsulates implicit and precise 3D scene perception features; defines a set of learnable parameters Where H t = 6 represents the number of historical prediction trajectories, T step = 8 represents the number of time steps for the current trajectory planning, Z gru = 256 represents the dimension of the GRU unit; first apply temporal self-attention to a set of learnable query parameters Q wp as follows: Q′ wp = MHA self (Q wp , Q wp , Q wp ) Among them, Q′ wp is the historical feature query, and MHA self (·) is the multi-head self-attention function; Next, the ego vehicle speed S ego = v ego is brought into the observation state; the general fusion feature F fused undergoes channel transformation, position encoding, and feature flattening operations, and then is connected to the encoded ego state to form the enhanced feature F enh : F flat = Flatten(F fused + P pos ) F ego = MLP(S ego ) + P pos Among them, Flatten(·) is an unfolding function used to unfold multi-dimensional matrix data into a one-dimensional vector; MLP(·) is a multi-layer perceptron composed of multiple linear layers and non-linear activation functions; ⊕ represents concatenating vectors; P pos is a position vector obtained from position encoding and used to add position information to features; Subsequently, at Q′ wp and F enh perform spatial cross-attention to generate spatio-temporal feature F st : F st = MHA cross (O′ wp , F enh , F enh ) Among them, MHA cross (·) is the multi-head cross-attention function.

5. The autonomous driving method based on modal fusion and Bessel optimization according to claim 4, characterized in that, The step 4 specifically comprises the following steps: Step 4.1: Restore the image feature F img Decoding: LiDAR interaction information is merged into F img where its spatial features are enhanced through feature complementarity; the depth and semantic branches are respectively derived from F img and are used to decode the depth map and the semantic segmentation map respectively; specifically, through a cascaded structure of multi-level deconvolution and interpolation upsampling N-level bilinear interpolation doubles the spatial resolution, facilitating the reconstruction of spatial information from low-dimensional features to high-resolution output: Among them, and respectively represent the sampling parameters of the n-th layer in the corresponding formula; the depth estimation branch and the semantic segmentation branch adopt the same multi-level architecture, but maintain independent transposed convolution and interpolation sampling parameters and enable adaptive feature learning for the domains of these two tasks, and make the output images I depth and I isem maintain a consistent spatial resolution; Step 4.2: Aerial view feature F bev Decode; Use F bev Perform BEV perspective BevSemantic and BoundingBox predictions, supplementing the front view depth and semantic segmentation maps derived from F img to enhance the model's 3D information and semantic perception of the global environment; for the BevSemantic branch, a lightweight feature transformation architecture is adopted to map spatial features to semantic classes: I bsem = Upsample(Conv(Conv(F bev ))) Where Conv(·) represents the convolution function; The BoundingBox branch, as a pure 3D detection task, adopts a multi-branch parallel independent task head design, including the object center point heat map O hp , the bounding box size O wh , the center offset O yc and O yr joint prediction for synthesizing the yaw angle O y : O x = MLP(Conv(Conv(F bev ))) where O x ∈{O hp , O wh , O os , O yr , O yc} represents each prediction head, and N rc is the number of discrete direction classes for rough direction estimation, set to 15; Step 4.3: Curve decoding; The branch decodes the spatio-temporal features to obtain the control points of the Bezier curve, thereby obtaining the Bezier curve, and then samples to obtain the planned trajectory; The scenario used is a navigation scenario with discrete target points P tg ; the target point information is embedded into spatio-temporal features to guide trajectory planning; multiple layers of GRU cells sequentially predict the control points P b : F tg = MLP(P tg ) x i = MLP(h i ) where h0 = Concat(F tg , F st ), x i represents the offset of the i-th layer GRU based on x i-1 . The offset of outputs the final output trajectory control point. t c = 4 is the total number of GRUcell layers and the number of curve control points. This work uses a cubic curve for trajectory planning to balance lightweight calculations with smooth dynamics. The continuous curve B is derived from the control points P b , and the trajectory points are obtained by sampling P trj = Sample(B) Where Sample(·) represents an equally spaced sampling function, and s represents the position parameter on the Bezier curve; After decoding to obtain the output, a dedicated loss function is tailored for each perception and planning component; the supervision mechanism includes two categories: one is the auxiliary task supervision L A , which is the auxiliary task supervision derived from the BEV feature decoder and the image feature decoder; the other is the trajectory planning constraint L T , which is applied to the physical constraints of the branch.

6. The autonomous driving method based on modal fusion and Bessel optimization according to claim 5, wherein In step 4.3, the auxiliary task supervision L A specifically includes: the depth estimation supervision L of the depth map in the depth estimation branch depth , the semantic class supervision of the semantic segmentation branch the semantic class supervision of the BEV semantic branch segmentation and the supervision L of various heads in the BoundingBox branch bbox ; Depth Estimation Supervision L depth : Use the L1 loss to supervise the depth estimation branch to ensure the metric accuracy of the depth information for 3D scene reconstruction: Among them, among them D represents the predicted depth map and the ground truth depth map respectively; Semantic segmentation supervision: I isem and I bsem The semantic segmentation task uses cross-entropy loss on C = 22 classes: where y c is the ground truth in one-hot vector form; BoundingBox branch supervision: The detection branch combines multiple supervision targets: first, the height / width / center point / center point offset / yaw angle residual of the object: supervised by L1 loss with continuous parameters; second, yaw angle classification: cross entropy loss of discrete yaw boxes; third, Heatmap prediction: Gaussian coding supervision improved focal loss; Among them, the hyperparameters α and γ are used to adjust the balance between positive / negative samples and hard / easy samples, respectively, while is used for supervised calculation to exclude traffic participants outside the predefined boundary range; y x corresponds to the supervised signal for each prediction head; Heatmap ground truth y p is generated through a Gaussian kernel procedure centered on the annotated target location; Final auxiliary task loss L A is expressed as a weighted sum of these component loss terms: Trajectory planning constraint L T The target point guidance and safety constraints are implicitly obtained through the L1 reconstruction loss, and the kinematic feasibility of the vehicle is strengthened through regularization constraints, curvature continuity constraints, and smooth control point constraints to enhance the trajectory smoothness and feasibility; Reconstruction loss L rs : Ensure geometric alignment with the ground truth trajectory; where T represents the number of predicted time intervals, represents the finally predicted trajectory point, P trj represents the true trajectory point; Control Point Regularization L cp : Ensure kinematic priors by constraining the initial control points to align with the ego vehicle speed : Smoothness Constraint L sm : To constrain the consistency of the distances between adjacent control points and avoid sudden trajectory jitters to ensure driving comfort: where d ref is a threshold determined from vehicle power; Curvature Continuity L cc : By maintaining the continuity of the second derivative through control point constraints, the continuity of the trajectory curvature change rate is ensured, so that the predicted trajectory conforms to the dynamic characteristics of the vehicle: L cc = ||P3 - 2P2 + P1||² Trajectory planning constraint L T Combine these components with empirical weights: L T = w rs ·L rs + w cp ·L cp + w sm ·l sm + w cc ·L cc Set the weight parameter w rs , w cp , w sm and w cc , to ensure vehicle power feasibility and ride comfort under the appropriate performance lower limit.

Citation Information

Patent Citations

  • Automatic driving vehicle track prediction method and device and electronic equipment

    CN113705636A

  • Transform-based time sequence point cloud three-dimensional target detection

    CN116740424A